Predict human preference to LLM responses.
Binfeng Xu
billxbf
AI & ML interests
evolving back to apes
Organizations
None yet
models 21
billxbf/qwen3.5-4b-pi-polar
4B • Updated • 5
billxbf/qwen3.5-4b-opencode-polar
4B • Updated • 5
billxbf/qwen3.5-4b-qwencode-polar
4B • Updated • 21
billxbf/qwen3.5-4b-claudecode-polar
4B • Updated • 5
billxbf/qwen3.5-4b-codex-polar-step72
Reinforcement Learning • 5B • Updated • 5
billxbf/zephyr-7b-dpo-iter1
Text Generation • 274k • Updated • 8
billxbf/zephyr-7b-dpo-iter3
Text Generation • 266k • Updated • 7
billxbf/zephyr-7b-dpo-iter2
Text Generation • 266k • Updated • 6
billxbf/Nano-Raccoon-Preview-1104
425k • Updated • 5
billxbf/zephyr-7b-sft-iter3
Text Generation • 266k • Updated • 7
datasets 20
billxbf/math_pile_v3
Viewer • Updated • 1.52M • 28
billxbf/ultrafeedback-dpo-iter3
Viewer • Updated • 20.4k • 13
billxbf/ultrafeedback-dpo-iter1
Viewer • Updated • 20.4k • 10
billxbf/ultrafeedback-dpo-iter2
Viewer • Updated • 20.4k • 12
billxbf/ultrafeedback-sft-iter3
Viewer • Updated • 20.4k • 22
billxbf/ultrafeedback-sft-iter2
Viewer • Updated • 20.4k • 11
billxbf/ultrafeedback-sft-iter1
Viewer • Updated • 20.4k • 15
billxbf/verified100-chitchat
Viewer • Updated • 100 • 13
billxbf/verified100-lite
Viewer • Updated • 100 • 22
billxbf/verified100
Viewer • Updated • 100 • 19