A kitchen-sink collection of experimental auxiliary-head architectures built on DeepSeek Flash, exploring anything that might extend MOE architectures
Nicholai Mitchko
nmitchko
AI & ML interests
Fine-tuning, Scaling, Enablement, Activation Engineering
Recent Activity
upvoted a paper about 24 hours ago
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning
Chains updated a model 10 days ago
nmitchko/deepseek-v4-Flash-CoLaR published a model 14 days ago
nmitchko/deepseek-v4-Flash-CoLaROrganizations
None yet