Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
Abstract
A hierarchical taxonomy and dense supervision strategy improve diffusion-based image editing through fine-grained concepts, large-scale paired data, and granular evaluation.
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.
Community
๐ ConceptEdit: Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- VicEdit: Learning to Edit Videos from Visual In-Context Examples (2026)
- InnoText: A Unified Model for Visual Text Generation and Editing (2026)
- Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing (2026)
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation (2026)
- Evaluation-Verification Reward for Consistent Multi-Reference Image Editing (2026)
- EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing (2026)
- CPI-Bench: A Comprehensive, Practical and Intelligent Benchmark for Real-World Image Editing (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.16812 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 2
inclusionAI/ConceptEdit-Bench
Spaces citing this paper 0
No Space linking this paper