Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Paper • 2607.28611 • Published 24 days ago • 22
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published Jul 19 • 167
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published Jul 18 • 141
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising Paper • 2607.00407 • Published Jul 1 • 10
fpadovani/dan-latn-10mb-ppt-shuff-dyck-100mb_seed10 Text Generation • 39.1M • Updated Jul 6 • 147 • 1
jovaldivieso/double_integrator_casadi_diffusion_policy Robotics • 251k • Updated 23 days ago • 323 • 2
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration Paper • 2605.20025 • Published May 19 • 191
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Paper • 2605.20266 • Published May 18 • 56