A kitchen-sink collection of experimental auxiliary-head architectures built on DeepSeek Flash, exploring anything that might extend MOE architectures
Nicholai Mitchko
nmitchko
AI & ML interests
Fine-tuning, Scaling, Enablement, Activation Engineering
Recent Activity
upvoted a paper 2 days ago
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning
Chains updated a model 11 days ago
nmitchko/deepseek-v4-Flash-CoLaR published a model 15 days ago
nmitchko/deepseek-v4-Flash-CoLaROrganizations
None yet