deep layers
diverse data
compact size
overfitting
grokking
I know the next best small model is somewhere in the intersection of these features.
Need no worries, once I get enough money I will start funding small language model devs so we can get rid of the hands of big companies once and for all
Yes, I got the script for ArithMark-3 and I can re-run it with no issues.
Mostly across harness details rather than a completed five-seed sweep. We have checked the checkpoint under repeated runs and small evaluation-pipeline variations, but I do not want to represent that as a formal variance estimate.
Your distinction is valid though, prompt formatting and answer parsing measure harness sensitivity, while multiple seeds under one frozen configuration might establish the actual noise band. I don't know if it will be published at all. Since is a little bit chaotic keeping base skills while improving them so our iteration method is barely reproducible.
Thatโs a fair point. At this scale, instruction tuning is definitely a capacity tradeoff rather than a free capability layer, so preserving the base modelโs strengths will be one of the main acceptance criteria.
Regarding the post-merge result, we have done additional validation runs and the performance appears directionally consistent, although I would prefer to publish the full repeated-evaluation results once the methodology and comparison conditions are finalized. At 90M, even small evaluation details can meaningfully affect the reported ranking.
The instruct version will be treated as a separate checkpoint rather than a replacement for the base model.
Can't wait to see what the community ๐ชdo with this! ๐๐๐
palmer-006 (90M)Cool stuff right there! Keep it up