Disaggregated Quantization: Specializing LLM Prefill and Decode Paper • 2609.26333 • Published 9 days ago • 89
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs Paper • 2609.29845 • Published 7 days ago • 99
Tiny llm ablation Collection Small language-model ablations trained on FineWeb-Edu, with evaluation confidence intervals and TensorBoard logs. • 5 items • Updated 8 days ago
Tiny llm ablation Collection Small language-model ablations trained on FineWeb-Edu, with evaluation confidence intervals and TensorBoard logs. • 5 items • Updated 8 days ago