Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching Paper • 2609.01404 • Published 1 day ago • 18
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge Paper • 2604.18164 • Published Apr 20 • 4
naver-hyperclovax/HyperCLOVAX-SEED-Vision-Instruct-3B Text Generation • 4B • Updated Sep 16, 2025 • 28.3k • 221
Evaluating Multimodal Generative AI with Korean Educational Standards Paper • 2502.15422 • Published Feb 21, 2025 • 10
Evaluating Multimodal Generative AI with Korean Educational Standards Paper • 2502.15422 • Published Feb 21, 2025 • 10 • 3
Cream: Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models Paper • 2305.15080 • Published May 24, 2023
Evaluating Multimodal Generative AI with Korean Educational Standards Paper • 2502.15422 • Published Feb 21, 2025 • 10