Multilingual GSM-Symbolic: What determines capability transfer across languages? Paper • 2610.03367 • Published 4 days ago • 25
Multilingual GSM-Symbolic: What determines capability transfer across languages? Paper • 2610.03367 • Published 4 days ago • 25
Selecting The Most Informative Tokens in Natural Language Autoencoders Paper • 2609.37040 • Published 7 days ago • 17
Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic LLM Architectures Paper • 2608.19889 • Published Aug 20
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data Paper • 2608.13517 • Published Aug 13 • 33
PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models Paper • 2606.09697 • Published Jun 8 • 7
Isolating Culture Neurons in Multilingual Large Language Models Paper • 2508.02241 • Published Nov 11, 2025
Selecting The Most Informative Tokens in Natural Language Autoencoders Paper • 2609.37040 • Published 7 days ago • 17
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data Paper • 2608.13517 • Published Aug 13 • 33
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment Paper • 2606.10747 • Published Jun 9 • 13
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment Paper • 2606.10747 • Published Jun 9 • 13
BrainSurgery: Reproducible and Reliable Declarative Weight Manipulations for Model Editing and Upcycling Paper • 2606.09707 • Published Jun 8 • 8
BrainSurgery: Reproducible and Reliable Declarative Weight Manipulations for Model Editing and Upcycling Paper • 2606.09707 • Published Jun 8 • 8
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes Paper • 2502.06728 • Published Feb 10, 2025
Are We Really Making Much Progress in Text Classification? A Comparative Review Paper • 2204.03954 • Published Apr 8, 2022
Efficient Continual Learning for Small Language Models with a Discrete Key-Value Bottleneck Paper • 2412.08528 • Published Dec 11, 2024
FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Federatedly-trained Large Language Models Paper • 2602.08818 • Published Feb 9 • 3