MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs Paper • 2608.29286 • Published Aug 29 • 1
MMPCBench: Benchmarking Multimodal Large Language Models on Proactive Critique of Flawed Inputs Paper • 2608.29286 • Published Aug 29 • 1
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models Paper • 2609.38827 • Published 8 days ago • 60
More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models Paper • 2609.38827 • Published 8 days ago • 60
ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World Paper • 2505.19095 • Published May 25, 2025 • 2
ARE: Scaling Up Agent Environments and Evaluations Paper • 2509.17158 • Published Sep 21, 2025 • 36
Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability Paper • 2508.04017 • Published Aug 6, 2025 • 12
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models Paper • 2505.23715 • Published May 29, 2025 • 2