Introduction
How a documentation agent built on HydraDB's graph-backed retrieval scored 90.7 on the nlohmann/json repository in CodeWikiBench, compared with 66.1 for DeepWiki and 61.3 for CodeWiki.
A coding repository is a network of files, modules, types and functions connected by calls, includes, inheritance and data flow. Most coding tools still explore it like a folder. They list directories, grep for strings and read one file at a time, then try to reconstruct the relationships in the model's context window. We built a documentation agent on a different premise. First, the repository is ingested into HydraDB, which turns the code into a queryable knowledge graph. Every step of the agent then starts from graph-aware retrieval instead of a blank directory listing.
We tested this harness on CodeWikiBench [1, 2], a benchmark of repository-level documentation. Each repository has a hierarchical rubric derived from the maintainers' own documentation, and a panel of three LLM judges scores every generated wiki. We followed the benchmark's methodology and compared harness against harness: the DeepWiki harness [3], the CodeWiki harness [1], and our HydraDB-backed harness. All three are scored on the same repository commit, rubric, judge models and scoring formula. On the C++ library nlohmann/json, our harness scores 90.73 ± 3.29 / 100, and all three judges agree that 47 of 57 criteria are covered. The published results on the same rubric are 66.06 (33/57) for DeepWiki and 61.28 (30/57) for CodeWiki [1, Table 4]. We will publish the open-source harness repo when we will run full benchmark across 22 repos for reproduction of results.
Why a repository should be treated as a graph
Ask an engineer how nlohmann::json parses a string, and the answer crosses many files. basic_json::parse builds an input adapter and hands it to parser. The parser pulls tokens from lexer and emits events to a SAX handler, and json_sax_dom_parser turns those events back into a basic_json tree. No single file contains that explanation. It exists only in the relationships between files.
Coding tools approach those relationships in roughly three ways:
| Approach | How relationships are found | Limitation |
|---|---|---|
| File-system agents | ls, grep, and reading files one at a time. The model infers connections in its own context. | Every connection must be rediscovered in every session, and the context window fills with raw file contents before any structure appears. |
| Partial graphs | A static AST or dependency graph is built once, then used to split the repository into modules. CodeWiki takes this approach [1]. | The graph shapes the plan but is not something the writing agent can query. It is computed once per run, with no retrieval over it and no memory across versions. |
| Graph-backed memory (HydraDB) | The repository is ingested into a persistent knowledge graph with relations between entities. Retrieval combines semantic, keyword and graph signals. | This is the approach tested here. |
HydraDB is built as a memory layer rather than a search index. Ingested content is organized around an ontology of entities and relations, and time is part of the data model, so knowledge can be versioned and reasoned about over time. Retrieval is hybrid: vector similarity, keyword matching and graph context work together, and the relations connected to a query can be forced into the ranking. For code, this means an agent can ask "how does the parser hand values to the DOM builder?" and get evidence ranked with the code's relationships taken into account, instead of a list of files that happen to contain the word "parser".
Example of a graph of single source file on HydraDB:

The question we wanted to answer is whether this makes a measurable difference on a real benchmark.
2. Building the Harness
We built a complete harness around HydraDB: a controller, an indexing layer, a tool-restricted documentation agent and an evaluation layer. The agent's task is to read a repository snapshot and write a multi-page technical wiki with source citations and architecture diagrams, similar to what DeepWiki or CodeWiki produce.
Overview of Harness Pipeline:

2.1 Agent architecture
The harness is a controller and one documentation agent. The controller prepares the snapshot, waits until HydraDB has indexed every eligible file, and then runs the agent through a fixed sequence of sessions. Session count, page count, and the start of evaluation are set by the controller. Inside a session, the agent's job is to produce that session's artifact.

Each session starts with an empty message history. The survey note and the outline are the only state that later sessions inherit. The same model, the same tools, and the same step budget run the survey, the outline, and every page.
Inside a session.
Before the model speaks, the controller sends the session task to HydraDB and places the top 8 chunks in the context. The model then loops. On each step it may call tools, or it ends the session by calling finish alone with the artifact (a JSON survey, a JSON outline, or one Markdown page). A response may batch up to 12 tool calls. finish cannot be mixed with other calls. The session allows 20 model steps and 15 minutes. On the last step the controller tells the model to finish from the evidence it already has.

Tools.
The agent has four ways to look at the repository, plus finish. It has no shell, no network, and no path that can leave the indexed allowlist.
| Tool | What it returns | Bound |
|---|---|---|
| memory_search | HydraDB chunks for one natural-language question | 8 chunks; same query settings as §2.2 |
| list_files | Indexed paths containing a substring | 150 paths per call, paginated |
| read_file | Numbered source lines and the commit-pinned URL | 160 lines from one allowlisted file |
| module_notes | The saved survey: summary, named modules, evidence ranges, open questions | Available after the survey is saved |
| finish | The session artifact | Must be the only call in that response |
read_file is a lookup in the in-memory allowlist. A path outside that list, a bad line range, or any other tool error comes back as an error object in the same turn, and the model can try again. Survey evidence is checked before it is saved: every cited range must be lines read_file actually returned in that session, and each range is 1–20 lines. Dependency labels (calls, uses, imports, and so on) are stored as inferences from those reads. Later sessions are told to treat them as a reading list and to confirm behavior with HydraDB and read_file before writing it down.
What each session produces.
| Session | Input besides the mandatory search | Output |
|---|---|---|
| Survey | Directory inventory of the allowlisted corpus | One note: summary, up to 6 named modules, evidence ranges, open questions |
| Outline | That note, via module_notes | JSON plan of at most 6 pages: slug, title, description |
| Page | The note, plus that page's title and description | One Markdown page, with commit-pinned source links and Mermaid where the code supports it |
Pages are written one after another. A finished page is checkpointed by content hash, so a resumed run reuses it. After the sixth page, the controller exports the wiki into CodeWikiBench's navigation format and hands it to the judge panel in §3. The judges are a separate system: they read the finished wiki through docs_navigator and never call the documentation agent's tools.
2.2 Stage by stage
The controller downloads the benchmark record and passes only the repository URL and commit hash to the agent. It fetches that exact commit and applies a fixed input policy. For nlohmann/json, 263 files (3.6 MB) are eligible: 205 test files, 46 library headers under include/, and build and tooling files. 911 files are excluded, including all 564 documentation files. The agent never sees the maintainers' documentation or the rubric it will be graded against, so it has to learn the library from the code alone.
Build the HydraDB graph.
Each eligible file is uploaded to a dedicated HydraDB collection for the run. Its source ID is derived from the repository, commit, path and content hash. HydraDB extracts entities and relations and builds the graph and embedding index on its own servers. The harness does not supply any parsing rules for C++. Generation starts only after every source reports completed, and this run reached 263 of 263.
Every agent query is sent to HydraDB's /query endpoint with these settings:
| Setting | Value | Effect |
|---|---|---|
| query_by | hybrid | Combines semantic and keyword retrieval |
| mode | thinking | HydraDB's deeper retrieval mode |
| graph_context | true | Uses graph context in retrieval |
| query_forceful_relations | true | Forces relations connected to the query into consideration |
| ids | the run's allowlisted source IDs | Keeps retrieval inside this snapshot |
| max_results | 8 | Top 8 chunks per query |
HydraDB accepts at most 200 source IDs per request. With 263 files, each logical query is therefore sent as two requests with disjoint allowlists, and the results are merged by score. Results are filtered back to the allowlist and the run's collection. Identical queries are answered from a cache.
One generator model (GPT-6 Astra, served through OpenRouter) works in three kinds of sessions, and each one begins with a required HydraDB query:
- Survey. Reads entry points and saves module notes: named modules, evidence line ranges, relationships inferred from the code, and open questions. Every piece of evidence must point to source lines the agent actually read during the session.
- Outline. Plans six pages using the notes and retrieved evidence.
- Pages. Six sessions of up to 20 model steps each. Each page ends with finish and includes commit-pinned citations and Mermaid diagrams.
The agent has no shell and no arbitrary filesystem access. read_file returns at most 160 lines from an allowlisted file, and memory_search goes to HydraDB.
Evaluate. The generated wiki is exported in CodeWikiBench's navigation format. It is then scored by the benchmark's judge code, pinned to the upstream commit [6].
2.3 What the Agent run looked like
| Metric | Value |
|---|---|
| Sources indexed in HydraDB | 263 / 263 |
| Logical HydraDB queries | 19 (8 mandatory, 11 requested by the agent) |
| HTTP search requests | 38 (two shards per query) |
| Evidence chunks delivered | 152 from 41 distinct files |
| Tool calls | 206 read_file · 11 memory_search · 8 list_files · 7 module_notes |
| Generator tokens | 1.47 M across 49 model calls |
| Wall time (survey to final page) | 19 min 47 s |
| Output | 6 pages · 8,445 words · 6 Mermaid diagrams |
| Source citations | 190 / 190 valid path and line ranges at the pinned commit |
Retrieval and file reading worked together. HydraDB returned relevant chunks from 41 files at the start of each session. The agent then spent its steps on targeted read_file calls to confirm exact lines before citing them.
3. Benchmark
CodeWikiBench was introduced with the CodeWiki paper by FPT Software AI Center and the University of Melbourne [1]. It targets a gap that metrics like BLEU and ROUGE cannot fill: judging whether repository-level documentation actually covers how a system works.
Dataset. The public release [2] contains 22 open-source repositories in seven languages: Python, Java, JavaScript, TypeScript, C, C++ and C#. Each record contains:
- the repository URL and the exact commit evaluated;
- docs_tree and structured_docs, the maintainers' official documentation parsed into a tree;
- rubrics, a hierarchical rubric derived from that documentation.
Rubric-generator agents built on Claude Sonnet 4, Gemini 2.5 Pro and Kimi K2 each drafted a rubric from the official documentation, and the drafts were merged into one [1, §3.1]. Each node has a requirement and a weight from 1 to 3. The nlohmann/json rubric has 8 top-level areas and 57 leaf criteria:
| Area | Weight | Criteria |
|---|---|---|
| Core JSON object model and type system | 3 | 7 |
| STL-compatible container interface | 3 | 7 |
| Serialization and deserialization engine | 3 | 13 |
| Type conversion and serialization framework | 3 | 9 |
| JSON Pointer and path navigation | 2 | 6 |
| JSON Patch and document modification | 2 | 5 |
| Exception handling and error management | 2 | 4 |
| Configuration and customization | 2 | 6 |
Three judge models from different model families score every leaf criterion: Gemini 2.5 Flash [7], GPT-OSS 120B [8] and Kimi K2 Instruct [9]. Each judge opens wiki pages with a docs_navigator tool and returns 1 if the criterion is explained, described or mentioned, and 0 otherwise, with reasoning and evidence.
Scoring. Every result in this post is written as a score and a spread, for example 90.73 ± 3.29. The scoring method comes from the CodeWikiBench paper [1, §3.3].
Step 1: score each criterion. The three judges each vote 1 (documented) or 0 (not documented) on a leaf criterion. The leaf's score is the average of the votes, so a 1/1/0 vote scores 0.67. The spread \sigma is the standard deviation of the three votes. It measures how much the judges disagreed:
| Votes | Leaf score | \sigma | Meaning |
|---|---|---|---|
| 1 / 1 / 1 | 1.00 | 0.00 | All judges agree it is documented |
| 1 / 1 / 0 | 0.67 | 0.58 | Judges disagree |
| 1 / 0 / 0 | 0.33 | 0.58 | Judges disagree |
| 0 / 0 / 0 | 0.00 | 0.00 | All judges agree it is missing |
Step 2: combine up the tree. Every node above the leaves takes the weighted average of its children's scores. A child with weight 3 counts three times as much as a child with weight 1. The spreads are combined with the same weights:
In these formulas, n is a node, are its children, their weights, their scores and their spreads. The spread formula squares the weights and takes a square root. This is the standard way to combine independent uncertainties, so one disputed criterion doesn't dominate the total.
Reading the result. In "90.73 ± 3.29":
- 90.73 is the weighted share of the rubric that the judges consider documented.
- ± 3.29 measures how much the judges disagreed on the way to that score. Lower means stronger agreement.
The spread is not a confidence interval. It doesn't estimate how much the score would change if the wiki were generated again. The paper also reports coverage, the number of criteria that are satisfied. Here, a criterion counts as satisfied only when all three judges give it a 1.
4. Comparing Harnesses
A documentation system is more than a model. It combines how the repository is represented, what the agent can retrieve, how work is planned and how pages are written. CodeWikiBench compares these complete systems, and so do we:
| DeepWiki [3] | CodeWiki [1] | Ours (HydraDB) | |
|---|---|---|---|
| Repository representation | Proprietary, closed source | Static AST dependency graph, used to split the repository into modules | HydraDB knowledge graph that the agent queries throughout the run |
| How the agent finds relationships | Not published | Follows the module tree built from the dependency graph; recursive sub-agents | Hybrid and graph-context retrieval at the start of every session and on demand |
| Agent structure | Not published | Recursive multi-agent system with delegation, bottom-up synthesis | One agent: survey, then outline, then 6 pages |
| Generator model | Proprietary | Kimi K2 Instruct (for json) | GPT-6 Astra |
All three systems are evaluated on the same pinned commit (4bc4e37f), the same 57-leaf rubric, the same three judge models, the same judge prompt, temperature 0 and the same hierarchical scoring code. Our judge setup reuses the upstream CodeWikiBench evaluator at a pinned commit [6]. It also adds stricter checks: a score is accepted only if the saved trace shows the judge actually read page content, and a failed judgment is never counted as a 0 or 1. All 171 judgments (3 judges × 57 criteria) were run from scratch on this wiki. We compare complete harnesses rather than individual models or databases. Only the HydraDB Astra harness was evaluated in our pipeline; the DeepWiki and CodeWiki scores are taken from the CodeWikiBench paper.
5. Results
5.1 Harness against harness on nlohmann/json

| Harness | Score / 100 | Criteria covered (all 3 judges) |
|---|---|---|
| GPT-6-Astra Harness (ours) | 90.73 ± 3.29 | 47 / 57 |
| DeepWiki | 66.06 ± 3.08 | 33 / 57 |
| CodeWiki (Kimi K2) | 61.28 ± 2.35 | 30 / 57 |
GPT-6-Astra Harness scores +24.7 points over DeepWiki and +29.5 points over CodeWiki, and covers 14 and 17 more criteria respectively. Judge disagreement (±3.29) is in the same range as the baselines, so the higher score does not come from one lenient judge.
This is also a hard repository for the other systems. The paper reports that C and C++ repositories are where both CodeWiki and DeepWiki struggle most [1, §5.2]. Across its main set, CodeWiki with Claude Sonnet 4 averages 53.24% on these languages and DeepWiki 56.39%. The cause the paper names is heavily cross-file, template-driven code, where the architecture lives in relationships between files rather than in any single file.
5.2 Scores by judge

| Judge | Score |
|---|---|
| Gemini 2.5 Flash | 94.73 |
| GPT-OSS 120B | 93.72 |
| Kimi K2 | 83.74 |
Even the strictest judge, Kimi K2, scores our wiki 17.7 points above DeepWiki's panel average.
5.3 Scores by area
| Area | Panel | Gemini | GPT-OSS | Kimi |
|---|---|---|---|---|
| Core object model and type system | 100.0 | 100 | 100 | 100 |
| Serialization engine (DOM, SAX, 5 binary formats) | 100.0 | 100 | 100 | 100 |
| Exception handling | 100.0 | 100 | 100 | 100 |
| JSON Patch and document modification | 95.2 | 100 | 85.7 | 100 |
| JSON Pointer and path navigation | 92.6 | 100 | 88.9 | 88.9 |
| Type conversion framework | 91.6 | 100 | 95.9 | 78.8 |
| STL-compatible container interface | 75.0 | 87.5 | 87.5 | 50.0 |
| Configuration and customization | 69.6 | 66.1 | 87.5 | 55.4 |
The three areas with perfect agreement all depend on connecting many files. Parsing runs from the input adapters through the lexer, parser and SAX handlers to the DOM. Serialization runs from dump through the output adapters and serializer, and on to the CBOR, MessagePack, BSON, UBJSON and BJData codecs. The type system runs from basic_json's template parameters through value_t to the union storage. The weak spots are routine API details. For example, the size, empty and clear capacity functions were never explained, and C++20 three-way comparison and some low-weight configuration macros were not covered.
6. Conclusion
Repository understanding depends on relationships between files. The explanation of how a JSON string becomes a basic_json object lives in the connections between the lexer, parser, SAX handler and DOM builder, not in any one of those files. We built a documentation harness that starts from those connections. The repository is ingested into HydraDB as a knowledge graph. Every agent session begins with a hybrid, graph-context retrieval, and the agent confirms what it retrieves by reading the exact source lines it cites.
On CodeWikiBench's nlohmann/json repository, a C++ codebase of the kind the paper found hardest, this harness scored 90.73, compared with 66.06 for DeepWiki and 61.28 for CodeWiki. All three judges agreed that it covers 47 of the 57 rubric criteria, against 33 and 30 for the baselines. The wiki cites 190 source lines, all valid at the pinned commit, and was written without the agent ever seeing the maintainers' documentation.
Next, we will run the full 22-repository suite, run a version of the same harness without HydraDB to measure the graph's contribution directly, and test repositories that change over time, where HydraDB's temporal memory can be exercised.
References
[1] A. Nguyen Hoang, M. Le-Anh, B. Le, N. D. Q. Bui. CodeWiki: Automated Repository-Level Documentation at Scale. arXiv:2510.24428 (v1), 2025. The nlohmann/json baselines are in Table 4. arxiv.org/abs/2510.24428v1
[2] CodeWikiBench dataset, Hugging Face, revision 6d215eb7d50a164e370a9a5703b813f9da345965. huggingface.co/datasets/anhnh2002/codewikibench
[3] Cognition. DeepWiki. deepwiki.com
[4] G. Starace et al. PaperBench: Evaluating AI's Ability to Replicate AI Research. arXiv:2504.01848, 2025. Source of the rubric-based evaluation approach CodeWikiBench adopts. arxiv.org/abs/2504.01848
[5] HydraDB. Harness integration: [src/hydra_agent/hydradb.py](src/hydra_agent/hydradb.py), [src/hydra_agent/codewiki_memory.py](src/hydra_agent/codewiki_memory.py).
[6] FSoft-AI4Code. CodeWikiBench evaluator, commit 5e728fb40492effb54d59041f908dbf9079fe238. github.com/FSoft-AI4Code/CodeWikiBench
[7] Google DeepMind. Gemini 2.5. arXiv:2507.06261, 2025. arxiv.org/abs/2507.06261
[8] OpenAI. gpt-oss-120b & gpt-oss-20b Model Card. arXiv:2508.10925, 2025. arxiv.org/abs/2508.10925
[9] Kimi Team. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534, 2025. arxiv.org/abs/2507.20534
[10] N. Lohmann. JSON for Modern C++. github.com/nlohmann/json