HydraDBBenchmarks

Building Coding Harness on HydraDB

A documentation agent built on graph-backed retrieval scores 90.73 on nlohmann/json, ahead of the published DeepWiki and CodeWiki results.

90.73
Score / 100
nlohmann/json
47 / 57
Criteria covered
all three judges
+24.7
Points over
DeepWiki
263 / 263
Repository files
indexed

Introduction

How a documentation agent built on HydraDB's graph-backed retrieval scored 90.7 on the nlohmann/json repository in CodeWikiBench, compared with 66.1 for DeepWiki and 61.3 for CodeWiki.

Abstract

A coding repository is a network of files, modules, types and functions connected by calls, includes, inheritance and data flow. Most coding tools still explore it like a folder. They list directories, grep for strings and read one file at a time, then try to reconstruct the relationships in the model's context window. We built a documentation agent on a different premise. First, the repository is ingested into HydraDB, which turns the code into a queryable knowledge graph. Every step of the agent then starts from graph-aware retrieval instead of a blank directory listing.

We tested this harness on CodeWikiBench [1, 2], a benchmark of repository-level documentation. Each repository has a hierarchical rubric derived from the maintainers' own documentation, and a panel of three LLM judges scores every generated wiki. We followed the benchmark's methodology and compared harness against harness: the DeepWiki harness [3], the CodeWiki harness [1], and our HydraDB-backed harness. All three are scored on the same repository commit, rubric, judge models and scoring formula. On the C++ library nlohmann/json, our harness scores 90.73 ± 3.29 / 100, and all three judges agree that 47 of 57 criteria are covered. The published results on the same rubric are 66.06 (33/57) for DeepWiki and 61.28 (30/57) for CodeWiki [1, Table 4]. We will publish the open-source harness repo when we will run full benchmark across 22 repos for reproduction of results.

Why a repository should be treated as a graph

Ask an engineer how nlohmann::json parses a string, and the answer crosses many files. basic_json::parse builds an input adapter and hands it to parser. The parser pulls tokens from lexer and emits events to a SAX handler, and json_sax_dom_parser turns those events back into a basic_json tree. No single file contains that explanation. It exists only in the relationships between files.

Coding tools approach those relationships in roughly three ways:

ApproachHow relationships are foundLimitation
File-system agentsls, grep, and reading files one at a time. The model infers connections in its own context.Every connection must be rediscovered in every session, and the context window fills with raw file contents before any structure appears.
Partial graphsA static AST or dependency graph is built once, then used to split the repository into modules. CodeWiki takes this approach [1].The graph shapes the plan but is not something the writing agent can query. It is computed once per run, with no retrieval over it and no memory across versions.
Graph-backed memory (HydraDB)The repository is ingested into a persistent knowledge graph with relations between entities. Retrieval combines semantic, keyword and graph signals.This is the approach tested here.

HydraDB is built as a memory layer rather than a search index. Ingested content is organized around an ontology of entities and relations, and time is part of the data model, so knowledge can be versioned and reasoned about over time. Retrieval is hybrid: vector similarity, keyword matching and graph context work together, and the relations connected to a query can be forced into the ranking. For code, this means an agent can ask "how does the parser hand values to the DOM builder?" and get evidence ranked with the code's relationships taken into account, instead of a list of files that happen to contain the word "parser".

Example of a graph of single source file on HydraDB:

Screenshot 2026-09-23 at 5.16.02 PM

The question we wanted to answer is whether this makes a measurable difference on a real benchmark.

2. Building the Harness

We built a complete harness around HydraDB: a controller, an indexing layer, a tool-restricted documentation agent and an evaluation layer. The agent's task is to read a repository snapshot and write a multi-page technical wiki with source citations and architecture diagrams, similar to what DeepWiki or CodeWiki produce.

Overview of Harness Pipeline:

Overview of the CodeWikiBench documentation harness pipeline

2.1 Agent architecture

The harness is a controller and one documentation agent. The controller prepares the snapshot, waits until HydraDB has indexed every eligible file, and then runs the agent through a fixed sequence of sessions. Session count, page count, and the start of evaluation are set by the controller. Inside a session, the agent's job is to produce that session's artifact.

Controller and agent session architecture

Each session starts with an empty message history. The survey note and the outline are the only state that later sessions inherit. The same model, the same tools, and the same step budget run the survey, the outline, and every page.

Inside a session.

Before the model speaks, the controller sends the session task to HydraDB and places the top 8 chunks in the context. The model then loops. On each step it may call tools, or it ends the session by calling finish alone with the artifact (a JSON survey, a JSON outline, or one Markdown page). A response may batch up to 12 tool calls. finish cannot be mixed with other calls. The session allows 20 model steps and 15 minutes. On the last step the controller tells the model to finish from the evidence it already has.

Documentation agent session and tool loop

Tools.

The agent has four ways to look at the repository, plus finish. It has no shell, no network, and no path that can leave the indexed allowlist.

ToolWhat it returnsBound
memory_searchHydraDB chunks for one natural-language question8 chunks; same query settings as §2.2
list_filesIndexed paths containing a substring150 paths per call, paginated
read_fileNumbered source lines and the commit-pinned URL160 lines from one allowlisted file
module_notesThe saved survey: summary, named modules, evidence ranges, open questionsAvailable after the survey is saved
finishThe session artifactMust be the only call in that response

read_file is a lookup in the in-memory allowlist. A path outside that list, a bad line range, or any other tool error comes back as an error object in the same turn, and the model can try again. Survey evidence is checked before it is saved: every cited range must be lines read_file actually returned in that session, and each range is 1–20 lines. Dependency labels (calls, uses, imports, and so on) are stored as inferences from those reads. Later sessions are told to treat them as a reading list and to confirm behavior with HydraDB and read_file before writing it down.

What each session produces.

SessionInput besides the mandatory searchOutput
SurveyDirectory inventory of the allowlisted corpusOne note: summary, up to 6 named modules, evidence ranges, open questions
OutlineThat note, via module_notesJSON plan of at most 6 pages: slug, title, description
PageThe note, plus that page's title and descriptionOne Markdown page, with commit-pinned source links and Mermaid where the code supports it

Pages are written one after another. A finished page is checkpointed by content hash, so a resumed run reuses it. After the sixth page, the controller exports the wiki into CodeWikiBench's navigation format and hands it to the judge panel in §3. The judges are a separate system: they read the finished wiki through docs_navigator and never call the documentation agent's tools.

2.2 Stage by stage

The controller downloads the benchmark record and passes only the repository URL and commit hash to the agent. It fetches that exact commit and applies a fixed input policy. For nlohmann/json, 263 files (3.6 MB) are eligible: 205 test files, 46 library headers under include/, and build and tooling files. 911 files are excluded, including all 564 documentation files. The agent never sees the maintainers' documentation or the rubric it will be graded against, so it has to learn the library from the code alone.

Build the HydraDB graph.

Each eligible file is uploaded to a dedicated HydraDB collection for the run. Its source ID is derived from the repository, commit, path and content hash. HydraDB extracts entities and relations and builds the graph and embedding index on its own servers. The harness does not supply any parsing rules for C++. Generation starts only after every source reports completed, and this run reached 263 of 263.

Every agent query is sent to HydraDB's /query endpoint with these settings:

SettingValueEffect
query_byhybridCombines semantic and keyword retrieval
modethinkingHydraDB's deeper retrieval mode
graph_contexttrueUses graph context in retrieval
query_forceful_relationstrueForces relations connected to the query into consideration
idsthe run's allowlisted source IDsKeeps retrieval inside this snapshot
max_results8Top 8 chunks per query

HydraDB accepts at most 200 source IDs per request. With 263 files, each logical query is therefore sent as two requests with disjoint allowlists, and the results are merged by score. Results are filtered back to the allowlist and the run's collection. Identical queries are answered from a cache.

One generator model (GPT-6 Astra, served through OpenRouter) works in three kinds of sessions, and each one begins with a required HydraDB query:

  1. Survey. Reads entry points and saves module notes: named modules, evidence line ranges, relationships inferred from the code, and open questions. Every piece of evidence must point to source lines the agent actually read during the session.
  2. Outline. Plans six pages using the notes and retrieved evidence.
  3. Pages. Six sessions of up to 20 model steps each. Each page ends with finish and includes commit-pinned citations and Mermaid diagrams.

The agent has no shell and no arbitrary filesystem access. read_file returns at most 160 lines from an allowlisted file, and memory_search goes to HydraDB.

Evaluate. The generated wiki is exported in CodeWikiBench's navigation format. It is then scored by the benchmark's judge code, pinned to the upstream commit [6].

2.3 What the Agent run looked like

MetricValue
Sources indexed in HydraDB263 / 263
Logical HydraDB queries19 (8 mandatory, 11 requested by the agent)
HTTP search requests38 (two shards per query)
Evidence chunks delivered152 from 41 distinct files
Tool calls206 read_file · 11 memory_search · 8 list_files · 7 module_notes
Generator tokens1.47 M across 49 model calls
Wall time (survey to final page)19 min 47 s
Output6 pages · 8,445 words · 6 Mermaid diagrams
Source citations190 / 190 valid path and line ranges at the pinned commit

Retrieval and file reading worked together. HydraDB returned relevant chunks from 41 files at the start of each session. The agent then spent its steps on targeted read_file calls to confirm exact lines before citing them.

3. Benchmark

CodeWikiBench was introduced with the CodeWiki paper by FPT Software AI Center and the University of Melbourne [1]. It targets a gap that metrics like BLEU and ROUGE cannot fill: judging whether repository-level documentation actually covers how a system works.

Dataset. The public release [2] contains 22 open-source repositories in seven languages: Python, Java, JavaScript, TypeScript, C, C++ and C#. Each record contains:

  • the repository URL and the exact commit evaluated;
  • docs_tree and structured_docs, the maintainers' official documentation parsed into a tree;
  • rubrics, a hierarchical rubric derived from that documentation.

Rubric-generator agents built on Claude Sonnet 4, Gemini 2.5 Pro and Kimi K2 each drafted a rubric from the official documentation, and the drafts were merged into one [1, §3.1]. Each node has a requirement and a weight from 1 to 3. The nlohmann/json rubric has 8 top-level areas and 57 leaf criteria:

AreaWeightCriteria
Core JSON object model and type system37
STL-compatible container interface37
Serialization and deserialization engine313
Type conversion and serialization framework39
JSON Pointer and path navigation26
JSON Patch and document modification25
Exception handling and error management24
Configuration and customization26

Three judge models from different model families score every leaf criterion: Gemini 2.5 Flash [7], GPT-OSS 120B [8] and Kimi K2 Instruct [9]. Each judge opens wiki pages with a docs_navigator tool and returns 1 if the criterion is explained, described or mentioned, and 0 otherwise, with reasoning and evidence.

Scoring. Every result in this post is written as a score and a spread, for example 90.73 ± 3.29. The scoring method comes from the CodeWikiBench paper [1, §3.3].

Step 1: score each criterion. The three judges each vote 1 (documented) or 0 (not documented) on a leaf criterion. The leaf's score is the average of the votes, so a 1/1/0 vote scores 0.67. The spread \sigma is the standard deviation of the three votes. It measures how much the judges disagreed:

VotesLeaf score\sigmaMeaning
1 / 1 / 11.000.00All judges agree it is documented
1 / 1 / 00.670.58Judges disagree
1 / 0 / 00.330.58Judges disagree
0 / 0 / 00.000.00All judges agree it is missing

Step 2: combine up the tree. Every node above the leaves takes the weighted average of its children's scores. A child with weight 3 counts three times as much as a child with weight 1. The spreads are combined with the same weights:

S(n)=∑iw(ci)S(ci)∑iw(ci),σn=∑iw(ci)2σci2∑iw(ci)S(n)=\frac{\sum_i w(c_i)S(c_i)}{\sum_i w(c_i)},\qquad \sigma_n=\frac{\sqrt{\sum_i w(c_i)^2\sigma_{c_i}^2}}{\sum_i w(c_i)}

In these formulas, n is a node, cic_i are its children, w(ci)w(c_i) their weights, S(ci)S(c_i) their scores and σci\sigma_{c_i}  their spreads. The spread formula squares the weights and takes a square root. This is the standard way to combine independent uncertainties, so one disputed criterion doesn't dominate the total.

Reading the result. In "90.73 ± 3.29":

  • 90.73 is the weighted share of the rubric that the judges consider documented.
  • ± 3.29 measures how much the judges disagreed on the way to that score. Lower means stronger agreement.

The spread is not a confidence interval. It doesn't estimate how much the score would change if the wiki were generated again. The paper also reports coverage, the number of criteria that are satisfied. Here, a criterion counts as satisfied only when all three judges give it a 1.

4. Comparing Harnesses

A documentation system is more than a model. It combines how the repository is represented, what the agent can retrieve, how work is planned and how pages are written. CodeWikiBench compares these complete systems, and so do we:

DeepWiki [3]CodeWiki [1]Ours (HydraDB)
Repository representationProprietary, closed sourceStatic AST dependency graph, used to split the repository into modulesHydraDB knowledge graph that the agent queries throughout the run
How the agent finds relationshipsNot publishedFollows the module tree built from the dependency graph; recursive sub-agentsHybrid and graph-context retrieval at the start of every session and on demand
Agent structureNot publishedRecursive multi-agent system with delegation, bottom-up synthesisOne agent: survey, then outline, then 6 pages
Generator modelProprietaryKimi K2 Instruct (for json)GPT-6 Astra

All three systems are evaluated on the same pinned commit (4bc4e37f), the same 57-leaf rubric, the same three judge models, the same judge prompt, temperature 0 and the same hierarchical scoring code. Our judge setup reuses the upstream CodeWikiBench evaluator at a pinned commit [6]. It also adds stricter checks: a score is accepted only if the saved trace shows the judge actually read page content, and a failed judgment is never counted as a 0 or 1. All 171 judgments (3 judges × 57 criteria) were run from scratch on this wiki. We compare complete harnesses rather than individual models or databases. Only the HydraDB Astra harness was evaluated in our pipeline; the DeepWiki and CodeWiki scores are taken from the CodeWikiBench paper.

5. Results

5.1 Harness against harness on nlohmann/json

Comparison of CodeWikiBench harness scores and criteria covered
HarnessScore / 100Criteria covered (all 3 judges)
GPT-6-Astra Harness (ours)90.73 ± 3.2947 / 57
DeepWiki66.06 ± 3.0833 / 57
CodeWiki (Kimi K2)61.28 ± 2.3530 / 57

GPT-6-Astra Harness scores +24.7 points over DeepWiki and +29.5 points over CodeWiki, and covers 14 and 17 more criteria respectively. Judge disagreement (±3.29) is in the same range as the baselines, so the higher score does not come from one lenient judge.

This is also a hard repository for the other systems. The paper reports that C and C++ repositories are where both CodeWiki and DeepWiki struggle most [1, §5.2]. Across its main set, CodeWiki with Claude Sonnet 4 averages 53.24% on these languages and DeepWiki 56.39%. The cause the paper names is heavily cross-file, template-driven code, where the architecture lives in relationships between files rather than in any single file.

5.2 Scores by judge

CodeWikiBench scores by judge model
JudgeScore
Gemini 2.5 Flash94.73
GPT-OSS 120B93.72
Kimi K283.74

Even the strictest judge, Kimi K2, scores our wiki 17.7 points above DeepWiki's panel average.

5.3 Scores by area

AreaPanelGeminiGPT-OSSKimi
Core object model and type system100.0100100100
Serialization engine (DOM, SAX, 5 binary formats)100.0100100100
Exception handling100.0100100100
JSON Patch and document modification95.210085.7100
JSON Pointer and path navigation92.610088.988.9
Type conversion framework91.610095.978.8
STL-compatible container interface75.087.587.550.0
Configuration and customization69.666.187.555.4

The three areas with perfect agreement all depend on connecting many files. Parsing runs from the input adapters through the lexer, parser and SAX handlers to the DOM. Serialization runs from dump through the output adapters and serializer, and on to the CBOR, MessagePack, BSON, UBJSON and BJData codecs. The type system runs from basic_json's template parameters through value_t to the union storage. The weak spots are routine API details. For example, the size, empty and clear capacity functions were never explained, and C++20 three-way comparison and some low-weight configuration macros were not covered.

6. Conclusion

Repository understanding depends on relationships between files. The explanation of how a JSON string becomes a basic_json object lives in the connections between the lexer, parser, SAX handler and DOM builder, not in any one of those files. We built a documentation harness that starts from those connections. The repository is ingested into HydraDB as a knowledge graph. Every agent session begins with a hybrid, graph-context retrieval, and the agent confirms what it retrieves by reading the exact source lines it cites.

On CodeWikiBench's nlohmann/json repository, a C++ codebase of the kind the paper found hardest, this harness scored 90.73, compared with 66.06 for DeepWiki and 61.28 for CodeWiki. All three judges agreed that it covers 47 of the 57 rubric criteria, against 33 and 30 for the baselines. The wiki cites 190 source lines, all valid at the pinned commit, and was written without the agent ever seeing the maintainers' documentation.

Next, we will run the full 22-repository suite, run a version of the same harness without HydraDB to measure the graph's contribution directly, and test repositories that change over time, where HydraDB's temporal memory can be exercised.

References

[1] A. Nguyen Hoang, M. Le-Anh, B. Le, N. D. Q. Bui. CodeWiki: Automated Repository-Level Documentation at Scale. arXiv:2510.24428 (v1), 2025. The nlohmann/json baselines are in Table 4. arxiv.org/abs/2510.24428v1

[2] CodeWikiBench dataset, Hugging Face, revision 6d215eb7d50a164e370a9a5703b813f9da345965. huggingface.co/datasets/anhnh2002/codewikibench

[3] Cognition. DeepWiki. deepwiki.com

[4] G. Starace et al. PaperBench: Evaluating AI's Ability to Replicate AI Research. arXiv:2504.01848, 2025. Source of the rubric-based evaluation approach CodeWikiBench adopts. arxiv.org/abs/2504.01848

[5] HydraDB. Harness integration: [src/hydra_agent/hydradb.py](src/hydra_agent/hydradb.py), [src/hydra_agent/codewiki_memory.py](src/hydra_agent/codewiki_memory.py).

[6] FSoft-AI4Code. CodeWikiBench evaluator, commit 5e728fb40492effb54d59041f908dbf9079fe238. github.com/FSoft-AI4Code/CodeWikiBench

[7] Google DeepMind. Gemini 2.5. arXiv:2507.06261, 2025. arxiv.org/abs/2507.06261

[8] OpenAI. gpt-oss-120b & gpt-oss-20b Model Card. arXiv:2508.10925, 2025. arxiv.org/abs/2508.10925

[9] Kimi Team. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534, 2025. arxiv.org/abs/2507.20534

[10] N. Lohmann. JSON for Modern C++. github.com/nlohmann/json