Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective
• 82
None defined yet.
Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
Smaller Models, Better Rejects: Preference Distillation Scaling