Abstract
General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist models and tools leaves unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions should be coordinated. We present Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs, which schedule independent operations concurrently and register successful outputs for downstream and cross-turn reuse across replaceable execution backends. Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval. On UniM-90, the harness raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, while increasing relative Semantic--Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, respectively; Strict Structure Score reaches 100.00 and 99.78. These results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.
Community
Hi everyone! Sharing our new work: Omni-IO Skills: Harnessing Your Agent Omni-Native π
Your agent is only one harness away from being omni-native!
General-purpose agents like Claude and Codex are strong at reasoning and planning, but audio, video, 3D, and coordinated media workflows still depend on external tools. Extending a foundation model to new modalities ties capability growth to costly retraining. Real tasks also never stay within one modality. For example, "watch this video, compose a BGM, and turn it into a slide deck with a cover" chains understanding, generation, and conversion, and reuses intermediate assets across turns.
Research question: Can a plug-and-play harness make an existing general-purpose agent omni-native while preserving its reasoning and planning core?
Our answer: Omni-IO Skills, a multimodal harness that wraps around the agent instead of replacing it:
- π§© Hierarchical Skills: 27 Skills covering 38 representative tasks across 7 artifact modalities and 4 capability families (understanding, generation, reasoning, retrieval)
- π Declarative execution graphs: multi-step workflows are planned as dependency graphs, and independent tasks run in parallel
- ποΈ Asset Registry: intermediate outputs persist and are reused across turns
- π Pluggable backends: switching providers only requires a config change
On UniM, with two host agents (GPT-5.6 Sol and Claude Sonnet 5), support rates rise from under 40% to 100%, and the SQCS rises from about 27% to 75β78%.
Code, along with 20 example workflows, is open-sourced. Feel free to read, share, and star! π
π Paper: https://arxiv.org/abs/2609.31847
π» Code: https://github.com/any2any-mllm/Omni-IO-Skill
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production (2026)
- OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning (2026)
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work (2026)
- SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness (2026)
- Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities (2026)
- openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents (2026)
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.31847 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper