Papers
arxiv:2610.02967

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

Published on Oct 2
· Submitted by
Yuanhao Ban
on Oct 9
Authors:
,
,
,
,
,

Abstract

Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals. Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking. A key challenge is how to combine these heterogeneous reward signals. We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction. In the Arena text-to-image leaderboard (https://arena.ai/), our RL-trained Flux2dev achieves an Elo rating 69 points above the base model, and our post-trained Ideogram-4 surpasses every open-source model on the leaderboard, reaching an Elo of 1223.5. (Claims of state-of-the-art performance are based on the Arena leaderboard snapshot as of September 4, 2026.) Our results suggest that effective rewards for frontier generative-model training require broad coverage of user intent and robustness to exploitation under optimization. To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.

Community

Paper submitter

Human preference alone isn’t enough for text-to-image post-training: an image can look appealing while missing prompt details or violating the requested style.
This work combines a preference reward model trained on ~5 million human votes with rubric rewards for prompt faithfulness and constraint satisfaction. It also introduces anti-reward-hacking gates and combines complementary policies through weight averaging.
The recipe improves already strong models in live, blind human evaluation: FLUX.2-dev gains 69 Arena points. Our post-trained Ideogram 4 surpasses all publicly listed open-source models on the Text-to-Image Arena leaderboard as of the September 4, 2026 snapshot.
We’d love to hear your thoughts!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.02967
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.02967 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.02967 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.