Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 12 days ago • 150
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published 17 days ago • 289
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 19 days ago • 342
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published about 1 month ago • 309
alphaedge-ai/siglip2-giant-opt-patch16-384-aze-32768 Zero-Shot Image Classification • 2B • Updated Jul 23 • 5 • 1
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents Paper • 2606.19047 • Published Jun 17 • 4
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification Paper • 2606.01476 • Published May 31 • 9
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 243
PANDO: Efficient Multimodal AI Agents via Online Skill Distillation Paper • 2605.24785 • Published May 26 • 11