VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification Paper • 2609.06245 • Published 16 days ago • 28
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models Paper • 2608.27550 • Published 25 days ago • 82
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Paper • 2604.08539 • Published Apr 9 • 49
bryanzhou008/vit-mae-base-finetuned-eurosat Image Classification • 85.8M • Updated Oct 21, 2024 • 9 • 1
bryanzhou008/swin-tiny-patch4-window7-224-finetuned-eurosat Image Classification • 27.6M • Updated Oct 30, 2024 • 21 • 1
bryanzhou008/vit-base-patch16-224-in21k-finetuned-eurosat Image Classification • 85.8M • Updated Oct 30, 2024 • 12 • 1
bryanzhou008/vit-base-patch16-224-in21k-finetuned-inaturalist Image Classification • 85.8M • Updated Aug 19, 2025 • 94 • 2
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction Paper • 2511.20937 • Published Nov 26, 2025 • 16
Running on Zero Agents Featured 115 SAM3 Video Segmentation 🐠 115 Track and label objects in videos using text prompts or clicks