One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing Paper • 2609.04190 • Published 15 days ago • 8
Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching Paper • 2602.12221 • Published Feb 12 • 6
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation Paper • 2601.16210 • Published Jan 22
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation Paper • 2412.15209 • Published Dec 19, 2024 • 1
Uncertainty in Action: Confidence Elicitation in Embodied Agents Paper • 2503.10628 • Published Mar 13, 2025
One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing Paper • 2609.04190 • Published 15 days ago • 8
One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing Paper • 2609.04190 • Published 15 days ago • 8
Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching Paper • 2602.12221 • Published Feb 12 • 6
HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation Paper • 2506.21546 • Published Jun 26, 2025 • 2
HalluSegBench: Counterfactual Visual Reasoning for Segmentation Hallucination Evaluation Paper • 2506.21546 • Published Jun 26, 2025 • 2