What Accuracy and Gradient Cosine Miss: Evaluating Feedback Alignment via Scale Stability, Reference Validity, and Depth Utility Paper • 2606.21126 • Published Jun 19
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 5 days ago • 171
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction Paper • 2603.20939 • Published Mar 21
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems Paper • 2508.08833 • Published Aug 12, 2025
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 6 days ago • 140
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published 5 days ago • 168