Developed by researchers from Software Heritage and European research universities, CodeGraph introduces the first large-scale open-taxonomy knowledge graph for source code. Scaling across 167 million files in the Stack-Edu corpus, it extracts implicit engineering abstractions—algorithms, architectural paradigms, and design patterns—via a code-specialized LLM and a three-stage Wikidata entity-linking pipeline, materializing 158 million nodes and approximately 1 billion typed edges across 14 languages.
- ✓Beyond Syntactic AST Parsing: Moves past token and syntax tree representations by explicitly mapping source code implementations to high-level engineering paradigms, design patterns, and algorithmic taxonomy.
- ✓Massive 158M Node and 1B Edge Scale: Built across 167 million files in Stack-Edu across 14 languages, capturing 145 million file nodes, 63,000 extracted concept entities, and 19,800 grounded Wikidata identifiers.
- ✓Three-Stage Linking Pipeline: Combines deterministic SPARQL entity linking, a Deep Research Agent for residual long-tail disambiguation, and hierarchy rollup for parent closures, calibrated via human gold sets and LLM judges.
🧭Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
🔗
Project Links & Resources
Direct AccessDirect access to official project resources and documentation🔬
In-Depth Technical Analysis
核心背景与行业痛点 / Background & Pain Points While public software repositories host billions of files, existing program analysis tools remain confined to token-level search and syntactic Abstract Syntax Trees (ASTs). These tools fail to capture the latent engineering knowledge embedded within implementations—such as the algorithms instantiated, design patterns deployed, and domain paradigms followed. As a consequence, coding agents lack structured conceptual ontologies when navigating monolithic codebases. ### 架构亮点与底层机制 / Architectural Highlights Researchers from Software Heritage and academic institutions introduced CodeGraph, a large-scale semantic pipeline for open-taxonomy source code mapping: 1. Code-Specialized LLM Extraction: Employs domain-tuned models to extract high-level engineering abstractions across diverse repositories; 2. Three-Stage Wikidata Grounding: - Deterministic SPARQL Linking: Directly grounds unambiguous entities against official Wikidata schemas; - Deep Research Agent Disambiguation: Resolves ambiguous, long-tail technical jargon through recursive query reasoning; - Hierarchy Rollup: Computes parent-of closures to construct a fully hierarchical semantic topology; 3. Calibrated Quality Control: Blends small human gold-standard sets with LLM-as-a-judge filtering to rigorously constrain false positives. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Executed across the 167-million-file Stack-Edu corpus: - Massive Graph Topology: Captures approximately 158 million nodes, encompassing 145 million file nodes, 63,000 concept entities, and 19,800 verified Wikidata identifiers; - 1 Billion Semantic Edges: Establishes approximately 1 billion typed edges mapping files to concepts, concepts to Wikidata URIs, and sub-concepts to taxonomic parents across 14 programming languages; - Linking Precision: Validated on expert gold test sets with a 91.8% entity resolution precision, substantially outperforming heuristic linkers. ### 开发者实战落地与开箱指南 / Developer Practical Guide - Research Documentation: Documented in arXiv preprint 2609.29474; - Downstream Applications: Serves as a foundational structural memory layer for code generation agents, architecture audits, and dependency vulnerability analysis; - Database Compatibility: Formatted for RDF/SPARQL graph endpoints and Neo4j enterprise deployment.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.