mlx-community/Qwen3.8-27B-OptiQ-4bit Image-Text-to-Text • 27B • Updated about 7 hours ago • 12.1k • 19
Semi-Supervised Reward Modeling via Iterative Self-Training Paper • 2409.06903 • Published Sep 10, 2024 • 1 • 1