TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
Abstract
Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often become longer, accumulate narrow sample-specific rules, and generalize poorly beyond the training distribution. We study this failure mode as prompt distributional overfitting and argue that it reflects a lack of representation control in discrete text-space optimization. We formalize this view through representational inefficiency, a dual-factor measure that decomposes prompt inefficiency into capacity cost and scope narrowness, attributing distributional prompt overfitting to their coupled growth during optimization. We propose TextReg, a regularization framework that realizes a soft-penalty objective through regularized textual gradients, combining Dual-Evidence Gradient Purification, Semantic Edit Regularization, and Regularization-Guided Prompt Update. Across multiple reasoning benchmarks, TextReg substantially improves out-of-distribution (OOD) generalization, with accuracy gains of up to +11.8% over TextGrad and +16.5% over REVOLVE.
Community
Can prompts overfit just like machine learning models do?
Prompt optimization has emerged as a powerful paradigm for improving LLM performance. However, we observed a recurring phenomenon: as optimization progresses, prompts often become longer, accumulate increasingly specific instructions, and generalize worse beyond the training distribution.
In this work, we study this failure mode as prompt distributional overfitting.
Rather than viewing prompts as arbitrary text, we treat them as structured representations of task knowledge and ask a fundamental question:
What makes an optimized prompt generalize?
To answer this question, we introduce representational inefficiency, a perspective that attributes prompt overfitting to the coupled growth of:
๐น Capacity Cost โ prompts becoming unnecessarily long
๐น Scope Narrowness โ rules becoming increasingly specific and less reusable
Building on this insight, we propose TextReg, a regularization framework for text-space prompt optimization that combines:
โ
Dual-Evidence Gradient Purification
โ
Semantic Edit Regularization
โ
Regularization-Guided Prompt Update
Across multiple reasoning benchmarks and model families, TextReg consistently improves out-of-distribution generalization, achieving gains of up to +11.8% over TextGrad and +16.5% over REVOLVE.
More broadly, we hope this work contributes to a better understanding of how prompts evolve during optimization and how regularization principles can be extended from parameter space to text space.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification (2026)
- Prompt-Robust Language Models: Which Training Strategies Work? (2026)
- SEPO: Evidence-Grounded Prompt Optimization via Structural Editing (2026)
- CAPO: Constraint-Aware Prompt Optimization for LLM Agents (2026)
- Ready2Blend: From Natural-Language Instructions to Composable Alignment Prompts (2026)
- DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal Reasoning (2026)
- Is Human-Readable Text Necessary for Effective LLM Fine-Tuning? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2605.21318 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper