Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Paper • 2607.07508 • Published Jul 8 • 30
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning Paper • 2012.13255 • Published Dec 22, 2020 • 6
Compacter: Efficient Low-Rank Hypercomplex Adapter Layers Paper • 2106.04647 • Published Jun 8, 2021 • 2
Training language models to follow instructions with human feedback Paper • 2203.02155 • Published Mar 4, 2022 • 26
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Paper • 2205.14135 • Published May 27, 2022 • 16
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Paper • 2210.09261 • Published Oct 17, 2022 • 2
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers Paper • 2210.17323 • Published Oct 31, 2022 • 12
Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning Paper • 2303.10512 • Published Mar 18, 2023 • 5
Direct Preference Optimization: Your Language Model is Secretly a Reward Model Paper • 2305.18290 • Published May 29, 2023 • 71