Models and data for On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models