Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
Abstract
DisCo is a research agent that distills operational knowledge into reusable skills, significantly improving autonomous ML research performance across benchmarks.
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent. We call this missing layer operational knowledge, the know-how that separates knowing a method from making it work. That knowledge is not absent from the field. It appears in repositories and papers, but in forms written for human readers and too large to load during a task. Once distilled into compact, verified skills, this knowledge can be reused across tasks rather than rediscovered during each run. We present DisCo, a skill-powered research agent that creates skills and uses them during research. Its distillation runs in two complementary forms: task-agnostic, condensing the field's widely used repositories into reusable skills, and task-oriented, producing the skills a concrete task calls for. The former, applied across the open ecosystem, yields the AREX-Skill Library, with 5,000+ verified skills distilled from 1,000 widely used ML repositories and organized into 20 areas and 178 capability families. With the GPT-5.5 backbone, research harness, and downstream execution budget held fixed, the skill-equipped research agent scores 134.3% higher on MLE-bench, 34.4% higher on PaperBench, 9.2% higher on FrontierCS, and 14.0% higher on PassNet than the same agent without skills. These gains come from adding distilled operating context under that fixed setup.
Community
📄 Daily Papers:https://huggingface.co/papers/2609.02749
📄 Paper:https://arxiv.org/abs/2609.02749
💻 GitHub:https://github.com/VectorSpaceLab/AREX-Skill
🚀 DisCo CLI:https://www.npmjs.com/package/@arex-skill/disco
https://www.tensorbrife.site/podcast/1ffb05d6-c35c-4dc1-b36c-0b1a8627bfc0
finding hard to read this paper, dont worry we gat you
tensorbrife analyzes and turn huggingface papers daily into podcast that can be listined to even while walking
to listin to this paper just click the link above and join today
or register at https://tensorbrife.site
Get this paper in your agent:
hf papers read 2609.02749 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper