DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 5 days ago • 146
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 8 days ago • 210
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 15 days ago • 367
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 14 days ago • 172
Online Learning with LLM Experts from Limited Feedback Paper • 2609.05820 • Published 17 days ago • 6
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 22 days ago • 63
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 19 days ago • 242
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 21 days ago • 119
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 22 days ago • 312
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Paper • 2608.20958 • Published Aug 21 • 58
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 182
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations Paper • 2608.15930 • Published Aug 16 • 47
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries Paper • 2608.05604 • Published Aug 6 • 81
MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks Paper • 2506.05982 • Published Jun 6, 2025 • 3
SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published Aug 10 • 84
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published Jul 8 • 95