Metrics

Measured, not estimated

What it costs to connect a library, how fast connections are made, and how far retrieval reaches, measured on public NRC regulatory documents against plain vector search and LLM-built graph tools. Each figure says how it was measured. Runs still in progress are marked.

Cost

What it costs to connect a library

LLM spend to build a connected graph from the same 100 documents. LLM-built graph tools pay for every chunk they read. Spyndel's connections use no LLM at all.

SpyndelLLM-built graph toolsPlain vector search
LLM tokens to make connections015.6M–31.4M per 100 documentsnone (no connections)
LLM cost, 100 documents$0.68 all on optional topic names and summaries$91–$133$0
LLM cost per documentabout $0.007$0.91–$1.33$0
Projected, 1,717 documentsabout $11about $1,460–$2,130$0
LLM cost per question asked$0 retrieval uses no LLM$0.0006–$0.0037$0

Claude Sonnet 5 for every LLM step, at AWS list price in us-west-2 ($2.20 per million input tokens, $11 per million output). Projections scale the 100-document measurement by the size of the text.

Connection speed

How fast connections are made

Spyndel checks candidate connections with its own compact models on a GPU, so speed is set by your hardware, not by an AI provider's rate limits.

59 s

to check 13,393 candidate connections and keep 10,276, on one GPU (about 230 per second)

12 min

to build a connected graph of 100 documents from scratch, against 51–70 minutes for the LLM-built graph tools

< 2 min

for a 400-page document to join an existing graph on one GPU

76 min

to build a connected graph of all 1,717 documents from scratch: 209,884 statements and 392,371 verified connections

21–42 s

per document to join the 1,717-document graph; a 230-page guide made 611 connections into 124 existing documents in 42 s

72 s

to reorganize the 1,717-document graph after new documents arrive: only the 96 new pairs are checked, 514,140 earlier checks are reused

One NVIDIA RTX 5090. On CPUs the same work runs, more slowly; Settings shows what is in use.

Retrieval

Retrieval that follows connections

Real questions rarely live in one document. Plain search returns passages that look like the question. Spyndel returns the statements that match, and then follows their verified connections, so an answer can reach the related requirement in another document, even when it is written in different words.

Standard Review Plan 11.4 · Solid Waste Management

"In facilitating decommissioning, designs should minimize, to the extent practicable, embedding contaminated piping in concrete, consistent with maintaining radiation doses ALARA during operations and decommissioning."

Standard Review Plan 11.3 · Gaseous Waste Management

"Design features that facilitate decommissioning, including their role in the decommissioning process. These should include both design features (such as modular components and adequate space for equipment removal) and operating procedures to minimize the amount of residual radioactivity…"

A real connection from Spyndel's graph of public NRC documents. Plain keyword search would not pair these: they share almost no words. Spyndel's own models checked that they are about the same thing.

Does following connections reach what plain search misses?

The test uses the documents' own citations: where one NRC document relies on another, we ask a question in the first document's words whose answer needs the second. Then we check which systems reach the second document, across all 1,717 documents, with Spyndel's connections switched on and off.

63%

of cited guidance reached by Spyndel across all 1,717 documents, following connections. Plain vector search reached 62%: level.

+7.5 pts

more cited guidance reached because of the connections alone: Spyndel retrieving statements with connections switched on against off, confirmed on 120 fresh questions

93%

of cited guidance reached by Spyndel on 100 documents, against 63%–70% for the LLM-built graph tools

Level

answer quality with plain vector search, within measurement noise in two blind tests of 120 fresh questions each: 1.33 against 1.31 and 1.20 against 1.33 out of 2

8,557

citation links between the public NRC documents, found in their own text: the ground truth for the test

240

questions written from those citations before any system ran: 120 for the first run, then 120 fresh ones to confirm it

It gets better with use

We taught the graph with 240 answered questions, then asked 120 new questions about the same needs in different words, so nothing could be remembered word for word.

+4.2 pts

more cited guidance reached after use: 73% before, 78% after, with the search Spyndel ships

1.15 → 1.32

answer score (0 to 2, graded blind) when retrieving statements, before use and after 960 answered questions

0

questions got worse: the cited document was lost for none of the 120 in the shipped search

73%75%77%79%04080120160200 answered questions learned from peak 78.3% at 185 73.3% before use
Cited guidance reached by the shipped search, measured on the same 120 test questions after every 5 answered questions. Each step is one question.
~40

answered questions before it starts to improve; most of the gain by about 85, and it tops out around 180-200. After that more use neither helps nor hurts (measured up to 960).

1.151.201.251.301.350240480720960 answered questions learned from 1.151.221.241.251.32
Answer score (0 to 2, graded blind) when retrieving statements, on the same 120 test questions after each 240 answered questions. Every point after the first is significantly above it (paired 95% intervals).
+0.17

answer score after 960 answered questions, when retrieving statements: +0.07 after the first 240, and still rising at 960.

Learning nudges ranking toward what answers cite and away from what they flag as unhelpful; statements flagged repeatedly and never cited drop out of search. No model is retrained and no LLM is involved.

Every retrieval, without an LLM

SpyndelPlain vector searchLLM-built graph tools
Follows connections between documentsYes, verified, with a strength on eachNoYes, made by an LLM
Median retrieval time37 ms22 ms1.7–2.5 s
LLM calls per question to retrieve001 or more
Gets better from useYes: connections answers rely on rank higherNoNo
Finds the right document for a simple lookup97%97%95%

Simple lookups (one fact from one document, 100 documents, 60 questions) are a baseline every system passes; the difference is in questions that span documents.

Method

How these were measured

  • Same text for everyone. Every system indexed the same extracted text of the same public NRC documents, so PDF reading played no part.
  • Same model for every LLM step. Claude Sonnet 5 for extraction, naming, summaries and answers; Claude Opus 5 wrote the questions and graded the answers without knowing which system produced them.
  • Written down first. Questions, measures and the rule for calling a difference were fixed before the first run. A difference is only called when its 95% interval excludes zero.
  • Defaults for the other tools. The LLM-built graph tools ran with their default settings, with set-up fixes only where a default broke on this material.