Qwen3-8B-CalibSFT-RLCR

Paper Code Dataset

Introduction

Qwen3-8B-CalibSFT-RLCR is Qwen3-8B-CalibSFT further trained with RLCR, a confidence-aware reinforcement learning method. It is released with our paper On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models. Compared with RLCR trained from Qwen3-8B, it improves accuracy, discrimination, and calibration on both in-distribution and out-of-distribution benchmarks. For training details, please refer to our GitHub repository.

Usage

Use the system prompt and user format from training:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "SUSTech/Qwen3-8B-CalibSFT-RLCR"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")

SYSTEM_PROMPT = (
    "A conversation between User and Assistant. The user asks a question, and the Assistant solves it. "
    "The assistant first thinks about the reasoning process in the mind and analyzes its confidence about "
    "the solution and then provides the user with the final answer as well as its confidence level. "
    "The confidence level indicates how certain the Assistant is about its answer, expressed as a decimal "
    "between 0 and 1 with exactly two decimal places, enclosed within <confidence> </confidence> tags. "
    "The response must strictly follow this format: <think> reasoning process here </think> "
    "<answer> final short answer only </answer> <confidence> 0.xx </confidence>. "
    "The <answer> tag must contain only the final answer string needed for exact-match evaluation, "
    "not a full sentence, explanation, or reasoning."
)
question = "What is the smallest positive integer n such that n^2 + n is divisible by 12?"
messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": f"\n\nPROBLEM: {question}\n\n"},
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=16384, do_sample=True, temperature=0.6, top_p=0.95, top_k=20)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
# <think> ... </think> <answer> ... </answer> <confidence> 0.xx </confidence>

Evaluation

Results are averaged over 8 in-distribution math benchmarks (DeepScaleR-Eval, MATH-500, MinervaMath, OlympiadBench, GSM8K, AIME 2024–2026) and 8 out-of-distribution benchmarks (HotpotQA, TriviaQA, DROP, MuSiQue, LiveBenchReasoning, NQOpen, PopQA, WebQuestions), sampling at temperature 0.6 with 4 responses per question (32 on AIME). DeepScaleR-Eval consists of 2,000 DeepScaleR questions held out from training.

Model In-distribution Out-of-distribution
Pass@1 ↑AUROC ↑Brier ↓ECE ↓ Pass@1 ↑AUROC ↑Brier ↓ECE ↓
Qwen3-8B59.3271.9033.2034.3645.3665.7145.1546.94
Qwen3-8B-RLCR65.3382.3614.4314.4643.9669.3831.1431.17
Qwen3-8B-CalibSFT-RLCR67.2086.3212.778.5945.7171.3123.2418.20

Citation

@article{wang2026pitfalls,
  title={On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models},
  author={Wang, Shuoyuan and Luo, Beier and Zeng, Hao and Yu, Chengyao and Zhang, Songxin and Xie, Zejian and Jing, Bingyi and Wei, Hongxin},
  journal={arXiv preprint arXiv:2609.32470},
  year={2026}
}
Downloads last month
14
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SUSTech/Qwen3-8B-CalibSFT-RLCR

Finetuned
Qwen/Qwen3-8B
Finetuned
(1)
this model

Dataset used to train SUSTech/Qwen3-8B-CalibSFT-RLCR

Collection including SUSTech/Qwen3-8B-CalibSFT-RLCR

Paper for SUSTech/Qwen3-8B-CalibSFT-RLCR