GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
Abstract
The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.
Community
👋 We’re excited to share GUI-HARVEST, which enables GUI agents to improve automatically by evolving their execution harness while keeping model weights frozen.
The idea: learn from what actually happens on screen. GUI-HARVEST compares screenshots and action traces across repeated runs, identifies recurring failures across tasks, and translates them into reusable code changes. Each change must pass checks on both task performance and its predicted behavioral effects.
📊 The figure highlights results on OSWorld-Verified:
- Six backbones, three step budgets: GUI-HARVEST achieves the highest score in every model–budget comparison shown.
- Optimize at 15 steps, evaluate at 15/50/100: the optimized harnesses remain frozen across evaluation budgets.
- Strong gains without weight updates: Qwen3-VL-32B improves by 12.33 percentage points over its initial harness at 15 steps, while Gemini 3.1 Pro reaches 79.14% at 100 steps.
We hope this helps make GUI agents more capable through systematic learning from execution experience. Feedback and discussion are welcome!
💻 Code: https://github.com/GaryYang12345/GUI-HARVEST
📄 Paper: https://arxiv.org/abs/2610.00948
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents (2026)
- VERSE: Verified Self-Evolving Optimizer for Agent Harnesses (2026)
- Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory (2026)
- StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments (2026)
- Causal Improvement Graph for Agentic Harness Optimization (2026)
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? (2026)
- Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.00948 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
