LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
Abstract
EvoMap externalizes verified execution experience into reusable structured Gene, improving long-workflow task completion and reducing token costs across diverse models.
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.
Community
New work: LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
Skills tell agents what to do. Genes preserve what actually worked.
Across 7 models, EvoMap Genes from verified execution experience outperform Skills by 8.7–15.5 points on long-workflow tasks.
Don’t rediscover experience. Reuse it.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows (2026)
- DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness (2026)
- StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows (2026)
- SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows (2026)
- DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations (2026)
- BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services (2026)
- StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.23200 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper