CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published 16 days ago • 83
A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning Paper • 2510.15444 • Published Oct 17, 2025 • 151
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning Paper • 2502.00511 • Published Feb 1, 2025 • 12