Video HopChain Collection Multi-hop video questions for RLVR, and the 8B model trained on them with Second-Wave Exploration. • 2 items • Updated 5 days ago
Running on CPU Upgrade Featured 3.3k The Smol Training Playbook 📚 3.3k The secrets to building world-class LLMs
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Paper • 2509.23661 • Published Sep 28, 2025 • 52
Video RLVR — final training data Collection The datasets behind our Qwen3-VL-8B video RLVR runs: the base 24f/100k mixture and the HopChain v5 multi-hop corpora. • 2 items • Updated 15 days ago
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 24 days ago • 272
Qwen3-VL-8B video RLVR - GRPO checkpoint ladder Collection Every published checkpoint of the 8B video-QA GRPO campaign, ranked by core-3 mean_accuracy (5,645 rows). Untrained base = 0.4426. • 9 items • Updated 23 days ago