SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published 1 day ago • 70
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 116
PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus Objectives Paper • 2505.19558 • Published May 26, 2025 • 1
ICE Collection In-Context Editing: Learning Knowledge from Self-Induced Distributions • 2 items • Updated Jun 8
PoliCon Collection PoliCon: Evaluating LLMs on Achieving Diverse Political Consensus Objectives • 2 items • Updated Jun 8
UAPO Collection Adaptive Preference Optimization with Uncertainty-aware Utility Anchor • 4 items • Updated Jun 8
SAVE Collection The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement • 4 items • Updated Jun 8
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement Paper • 2605.30888 • Published May 29 • 10
ICE Collection In-Context Editing: Learning Knowledge from Self-Induced Distributions • 2 items • Updated Jun 8
SAVE Collection The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement • 4 items • Updated Jun 8
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement Paper • 2605.30888 • Published May 29 • 10
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement Paper • 2605.30888 • Published May 29 • 10