Harvard-DCML/ADAPT-Qwen3-8.5B
Text Generation • 0.5B • Updated • 294
Data-Centric ML
A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)
Boomerang Distillation Enables Zero-Shot Model Size Interpolation