CodeGraph: First Open-Taxonomy Code Knowledge Graph Spans 158M Nodes and 1B Edges Across 167M Files
Developed by researchers from Software Heritage and European research universities, CodeGraph introduces the first large-scale open-taxonomy knowledge graph for source code. Scaling across 167 million files in the Stack-Edu corpus, it extracts implicit engineering abstractions—algorithms, architectural paradigms, and design patterns—via a code-specialized LLM and a three-stage Wikidata entity-linking pipeline, materializing 158 million nodes and approximately 1 billion typed edges across 14 languages.
- •Beyond Syntactic AST Parsing: Moves past token and syntax tree representations by explicitly mapping source code implementations to high-level engineering paradigms, design patterns, and algorithmic taxonomy.
- •Massive 158M Node and 1B Edge Scale: Built across 167 million files in Stack-Edu across 14 languages, capturing 145 million file nodes, 63,000 extracted concept entities, and 19,800 grounded Wikidata identifiers.
- •Three-Stage Linking Pipeline: Combines deterministic SPARQL entity linking, a Deep Research Agent for residual long-tail disambiguation, and hierarchy rollup for parent closures, calibrated via human gold sets and LLM judges.


