SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models Paper • 2608.29974 • Published 5 days ago • 5
Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection Paper • 2607.25310 • Published Jul 28 • 5
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization Paper • 2607.25659 • Published Jul 28 • 84
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published Jul 23 • 26
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published Jul 20 • 198
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published Jul 16 • 212
Metacognition in LLMs: Foundations, Progress, and Opportunities Paper • 2607.11881 • Published Jul 13 • 30
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published Jul 14 • 108
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 86
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published Jul 4 • 77
Mind the Heads: Topological Representation Alignment for Multimodal LLMs Paper • 2606.23885 • Published Jun 22 • 5
penfever/terminal_bench_2_exp_psu_swesmith_1K_glm_4_7_traces_jupiter__4_0__Qwen3_8B_20264b24e003 Viewer • Updated Jul 3 • 2.73k • 19 • 1
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Paper • 2606.24428 • Published Jun 23 • 52