DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 6 days ago • 169
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training Paper • 2609.14306 • Published 10 days ago • 19
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve Paper • 2608.16884 • Published Aug 17 • 19
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published Aug 10 • 787
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration Paper • 2605.17423 • Published May 17 • 31
Seedance 2.0: Advancing Video Generation for World Complexity Paper • 2604.14148 • Published Apr 15 • 171
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning Paper • 2604.12374 • Published Apr 14 • 39
Large Language Models Align with the Human Brain during Creative Thinking Paper • 2604.03480 • Published Apr 3 • 7
Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models Paper • 2604.02315 • Published Apr 3 • 5
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens Paper • 2603.23516 • Published Mar 6 • 53
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality Paper • 2602.14080 • Published Feb 15 • 23
On the Mechanism and Dynamics of Modular Addition: Fourier Features, Lottery Ticket, and Grokking Paper • 2602.16849 • Published Feb 18 • 8
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy Paper • 2602.17363 • Published May 15 • 8
Preliminary sonification of ENSO using traditional Javanese gamelan scales Paper • 2602.14560 • Published Feb 16 • 1
On Surprising Effectiveness of Masking Updates in Adaptive Optimizers Paper • 2602.15322 • Published Feb 17 • 11