TextReg Curbs Prompt-Optimizer Overfit, Beats TextGrad by 11.8%
TL;DR
- TextReg lifts out-of-distribution reasoning accuracy by up to +11.8% over TextGrad and +16.5% over REVOLVE across multiple benchmarks.
- On Tracking Shuffled Objects, TextReg beats TextGrad by +10.0 (5obj) and +9.9 (7obj) on Llama-3.1-8B-Instruct.
- Baselines can regress: on Phi-3.5-Mini-Instruct, REVOLVE underperforms zero-shot CoT on all six datasets tested.
Prompt optimizers that iteratively rewrite instructions can leave the prompt worse than the plain Chain-of-Thought baseline they started from. That is the headline finding of TextReg, a paper from Georgia Tech and the University of Illinois Urbana-Champaign that names the failure mode "prompt distributional overfitting" and proposes a fix.
The authors (Lucheng Fu, Ye Yu, Yiyang Wang, Yiqiao Jin, Haibo Jin, B. Aditya Prakash and Haohan Wang) report that TextReg lifts out-of-distribution accuracy by "up to +11.8% over TextGrad and +16.5% over REVOLVE" across multiple reasoning benchmarks. The paper argues that iterative rewriters tend to produce prompts that "become longer, accumulate narrow sample-specific rules, and generalize poorly beyond the training distribution."
The framework has three stages. Dual-Evidence Gradient Purification filters raw textual gradients using both the current mini-batch and a global RuleBank of recurring directives. Semantic Edit Regularization watches capacity and scope changes after each rewrite. Regularization-Guided Prompt Update then rewrites the prompt while preserving task-faithful corrections.
Evaluation ran on six Big Bench Hard tasks (Logical Deduction and Tracking Shuffled Objects at three, five and seven objects) plus GSM8K, SVAMP and MultiArith, with four open-source test engines (Qwen2-7B-Instruct, Phi-3.5-Mini-Instruct, Llama-3-8B-Instruct and Llama-3.1-8B-Instruct), Qwen2.5-7B-Instruct as the forward engine, and GPT-4o as the shared backward engine driving all optimization operations. On the hardest cells, TextReg beats TextGrad by +10.0 (5obj) and +9.9 (7obj) on Llama-3.1-8B-Instruct for Tracking Shuffled Objects, and by +8.4 (5obj) and +10.3 (7obj) on Llama-3-8B-Instruct.
The sharper point in the paper is about the baselines themselves. On Phi-3.5-Mini-Instruct, the authors write, "REVOLVE underperforms CoT on all six datasets," which they say exposes the prompt overfitting problem that motivates the work. It lands into AI Weekly's prompt-engineering coverage, which has been thin lately and more oriented to vendor guides than methods papers.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper TextReg Regularizes Prompt Optimization, Lifts Out-of-Distribution Reasoning Accuracy