Papers
arxiv:2609.29474

CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding

Published on Sep 24
· Submitted by
Andrea Gurioli
on Sep 28
Authors:
,
,
,
,

Abstract

Public software repositories, like GitHub and Software Heritage Archive, store billions of files, yet extracting their implicit engineering knowledge ---i.e., the algorithms they implement, the paradigms they follow, the patterns they instantiate, and the application domains they serve--- remains challenging, as current tools are constrained to syntactic and token-level analysis. We present a pipeline for building an open-taxonomy semantic annotation of source code using a code-specialised Large Language Model. The extracted entities are grounded in Wikidata through a three-stage linking procedure: a deterministic SPARQL stage handles unambiguous entities, a Deep Research Agent resolves the residual long tail, and a hierarchy-rollup stage imports the parent-of closure of each resolved Wikidata identifier. The resulting annotations are materialised as a source-code-specific open-taxonomy knowledge graph. We further introduce a calibrated quality-assurance protocol that quantifies annotation precision by combining a small human gold set with an LLM-as-a-judge filter. We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code. Our graph, named CodeGraph, contains approximately 158 million nodes, which include around 145 million files, about 63,000 extracted concept entities (such as algorithms, paradigms, design patterns, and application domains), and roughly 19,800 grounded Wikidata entities. Furthermore, CodeGraph features approximately 1 billion typed edges that connect files to their respective concepts, link these concepts to their grounded Wikidata identifiers, and relate them to their parent categories, covering 14 programming languages.

Community

CodeGraph is a pipeline that uses an LLM (Qwen3-Coder-30B) to label about 145M source files from Stack-Edu, across 14 programming languages, with four kinds of concepts: domains, algorithms, paradigms and design patterns.
The extracted labels are cleaned into a shared vocabulary of about 63K concepts, then linked to Wikidata in three stages: a SPARQL lookup, a Deep Research Agent for ambiguous cases, and an import of each linked concept's Wikidata parent categories.
The result is a typed property graph with about 158M nodes and about 1.02B edges that ties files to local concepts and to about 19.8K Wikidata entities, and it is much larger than the closest prior resource, GraphGen4Code.
Label quality was checked with a small human-annotated gold set and an LLM verifier (Devstral-2) run over the whole corpus, aimed at uses such as semantic code search and automated algorithm selection.

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.29474 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.29474 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.