Search engines

Microseconds in the stacks

A MARC21 catalogue parsed, indexed, graphed and drawn by one static Go binary — and an assistant that cannot cite a record the catalogue did not hand it.

Records
170,938
Graph
2.67 M edges
Cold build
3.6 s
Typeahead
71 µs
Throughput
31,685 req/s

01What it is

Take a MARC21 file, and serve it. Not “load it into Elasticsearch and serve that” — parse it, index it in-process, and answer queries out of the same memory the records were decoded into. No external search service, no CGO, no network dependency at query time.

make demo                       # sixty-record sample, no key, no account
make demo MARC=/path/to/your.mrc
ingested vub.mrc: 5572 records, 0 skipped in 32ms (12 workers)
indexed 30562 terms, 188831 postings
graph 26963 nodes, 77496 edges
index size 9.5 MiB (docs 3.0 MiB, folded 2.0 MiB, dict 727.3 KiB, postings 988.4 KiB,
  facets 1.8 MiB, graph 805.2 KiB, items 21.8 KiB, works 67.2 KiB) = 1782 bytes/doc
listening on :8080

The index is immutable once built. That single decision is where most of the performance comes from: no locks on the query path, no invalidation, and every structure free to be a flat arena instead of a graph of pointers the garbage collector has to walk forever. Rebuilding the demo corpus takes ~55 ms, which is why “just rebuild it” is a legitimate answer to most mutation questions.

Six things the corpus is served as: a search engine (inverted index, BM25 per field, facets, phrases, typo tolerance, sorting, over HTTP, h2c and Server-Sent Events); a graph, for readers with no query to type; a building, so a shelfmark becomes a walk; a vector space, for queries whose answers share none of their words; a bibliography, in the three formats a university actually uses; and an assistant that answers in prose over the catalogue’s own endpoints.

There is exactly one exception to “no network dependency”, and it is opt-in: the assistant and the embedding provider are somewhere else, unless you point them at a model on the same machine — which is the configuration the project recommends and documents first.

What is running on the demo box

Not the sample. The whole VUB repository harvest, read off the live /stats:

Documents 170,938
Terms / postings 240,416 / 5,928,882
Graph 565,774 nodes · 2,672,881 edges
Works clustered 4,430 works · 2,047 publications · 3,071 chapters joined to their book
Copies placed 9,300 in 5 buildings · 20 floors · 180 rooms · 560 racks
Resident index ~511 MiB, of which 250 MiB is vectors
Assistant gemini-3.1-flash-lite

Measured from outside, through Traefik, over the public internet — the server’s own took_us, so these are the work and not the round trip:

GET /suggest?q=mangro                       →      71 µs
GET /search?q=mangrove&sort=year            →   2,366 µs
GET /search?q=mangrove&facet=doctype,year   →   3,586 µs   389 hits; the year facet counts
                                                           every year over 170,938 records
GET /search?q=how+seeds+travel…&mode=vector →  276,185 µs  exhaustive exact scan of
                                                           170,938 × 384 float32

02The build: one pass, all cores, byte-identical

┌─ shard 1: decode ▸ store ▸ facet ▸ tokenize ▸ postings ─┐ vub.mrc ─mmap─┼─ shard 2: ... ┼─ merge ─▶ Index ─▶ server └─ shard N: ... ─┘ (parallel) (immutable)

The file is mapped, not read, cut at record terminators, and each shard gets a slice. A shard decodes a record with gomarc and immediately stores it, facets it and indexes its text, while the values are still in cache. There is no intermediate slice of parsed documents — parse and index are one phase, not two.

The interesting constraint is determinism. Shards must not change the answer:

  • Documents land in one pre-counted slice, filled by workers in disjoint ranges in file order, so document numbering does not depend on the worker count.
  • The dictionary merges by interleaving sorted term lists, split by first byte — since the sort is bytewise, all terms sharing a first byte lie in one contiguous run of every shard’s list, so byte ranges merge independently and concatenate in order.
  • Posting lists concatenate in shard order, which is document order, so only the delta joining one shard’s contribution to the previous one has to be rewritten.
The invariantThe finished index is byte-identical whatever the shard count — pinned by a test that builds at 1, 2, 12 and 37 shards and compares every structure.

A record with unreadable framing is skipped, parsing resynchronizes on the next terminator, and nothing is fatal. On a 3.78 GB union catalogue that came to 152 skipped records out of 3.59 million — 0.004%.

03Four arenas and a column store

Per-object allocation is the enemy here — not for speed of allocation, but because every pointer is something the garbage collector walks, forever, for an index that never changes.

  • Documents. Every display value sits in one byte arena, with uint32 (offset, length) refs per attribute. A handful of allocations for the whole corpus instead of ~8 per document, and nothing for the GC to scan.
  • Dictionary. Terms frozen into one sorted arena plus an offset table, so a prefix range is a contiguous slice found by binary search at both ends. An open-addressing hash table gives ~12 ns exact lookups; a parallel array of 4-byte character masks drives the fuzzy prefilter.
  • Folded text. A normalized copy of the searchable text, one run per document, so phrase verification is a byte search over stored bytes — no folding at query time — and the separators stop a phrase running out of one field into the next.
  • Postings. One arena of delta-encoded entries: varint(docDelta) + fieldMask + varint(tf). One list serves every field-boosted query; there is no per-field index, and cursors decode without allocating.

Facetable columns are dictionary-encoded: values interned once, documents holding uint32 codes. Single-valued columns collapse further to a direct per-document array — one load instead of three, 4 bytes per document less. Counting a facet is one increment per value with no string work at all; only the values actually returned are turned back into strings, chosen through a bounded heap rather than by sorting the column.

04The query path

  1. Query text goes through the same tokenizer as index time — lowercase, diacritics removed by NFD decomposition — so kinesitherapie finds Kinesithérapie.
  2. Each token resolves to a group of dictionary terms: exact, a prefix expansion capped at the 64 most frequent terms, or a typo rescue. Members OR, groups AND.
  3. A group is walked one of three ways, chosen from member count and document frequency: lazily with one cursor per member; merged once through a k-way heap; or — for expansions of 8 members or more — accumulated straight into per-document slots, one pass with no comparisons, which drops the log(members) the heap costs on every posting.
  4. The intersection leapfrogs from the rarest group. A multi-token AND that matches nothing falls back to a scored OR with a coordination factor, rather than returning an empty page.
  5. Scoring is BM25 per field (k1=1.2, b=0.75) with per-field weights — title 5, author 3, subject 2, journal 1.5, department 1 — plus a recency multiplier of at most 8%. Only the fields in a posting’s mask are visited.
  6. Top-K is a bounded min-heap, so memory is O(limit), not O(matches). Facets are counted in the same pass; filters are integer comparisons on the facet columns.

Typo tolerance runs three filters of rising cost over the terms sharing the token’s first byte: length from the offset table, then two popcounts over the character masks (a lower bound on edit distance, so it can only reject genuine non-matches), then bounded Damerau-Levenshtein over reusable DP rows. A handful of terms reach the last stage; candidate generation costs 1.9 µs, allocation-free.

Phrases keep no positions in the index. Each document carries a 64-byte bit set of adjacent token-pair hashes: a phrase can only occur where all of its pairs do, which rejects ~94% of probes before any scan. Survivors get a byte search anchored on the phrase’s rarest byte — byte frequencies counted from the corpus at build time — because anchoring on the first byte means comparing at every token boundary, folded text being one space per token.

Two dates, and an ordering that is not a second pass

260$c says when the work was published. 005 says when the cataloguing system last wrote the record — present on 170,839 of these records over 1,084 distinct days — and nothing read it. That is a time axis the catalogue never had, and one a search could never be ordered by, because there was no sort parameter at all.

Sorting is free: 389 matches for mangrove take 55 µs by relevance and 60 µs by year. Hit carries a Rank beside its Score and every ordering site compares Rank. That mattered more than the feature did — the order lived in four implementations, and four implementations only stay in agreement if they have one field to agree about. When they drift, a page boundary silently reorders and page 2 repeats a record. A test pages the whole result set three at a time and compares it against one page of two hundred.

05Rendering was the larger half

The uncomfortable measurementThe search was 8 µs of a 46 µs request. The rest was writing JSON.

So responses are not built with encoding/json. They are appended byte-wise into a pooled buffer straight out of the index arena: strings scanned eight bytes at a time for the characters JSON must escape, with the clean runs between them copied in bulk; object keys copied verbatim rather than escaped; scores rendered as fixed point rather than through general float formatting. A page of 20 hits went 36 µs → 19 µs (~980 MB/s), and the whole request 46 µs → 31 µs.

What is left is bandwidth, which is why include= exists: a client naming the nine fields it displays gets a proportionally smaller and cheaper response. Two experiments did not survive measurement and were reverted: precomputing per document whether its values need escaping at all (the scan is already hidden behind the copy), and dispatching through a function value to skip escaping (slower than the branch).

06The optimization log

Every line is a before and an after on the same machine — AMD Ryzen 5 3600, 12 threads, Go 1.25 — over the 5,572-record corpus.

Path Before After What did it
Phrase search 92 µs 18 µs keep a folded copy of the text; stop folding at query time
Phrase, matching 17 µs 12.0 µs bigram bit set + anchor on the rarest byte
Phrase, no match 41.6 µs 10.3 µs …and skip the OR fallback a phrase could never accept
Fuzzy 36 µs 13.7 µs reuse DP rows, narrow to bytes, character-mask prefilter
Prefix / typeahead 16 µs 10.2 µs stream single-token queries into scoring; merge only when it pays
Broad typeahead a* 637 µs 323 µs pack (doc, cursor) into heap words; per-document accumulation
3 facets 23.0 µs 11.6 µs reuse counter buffers
journal + subject facets 130.5 µs, 46.8 KB 15.8 µs, 1.0 KB bounded heap instead of sorting the column
Whole request 46 µs 31 µs the renderer, above
Ingest, end to end 38.2 ms 29.3 ms parallel dictionary interleave, measure-then-write posting merge
Parse 22.7 ms 13.1 ms frame and decode in place; flatten in one pass; pre-counted slice
Build, text phase 78 ms 28 ms tokenize once for three consumers; shard across cores
Graph build 4.8 ms 2.5 ms fold keys into an arena, bucket by first byte, MSB radix sort
Node typeahead 106 µs 7.5 µs bound the prefix range at both ends; keep labels out of the scan
Doc→doc graph path 24 µs 2.1 µs hub threshold as a fraction of the corpus, not a constant 512

Two of those generalize. Fusing ingest was worth ~2%, and that was not the point — both phases were already parallel. What the fusion exposed was that with shard work spread over twelve cores, the serial merges were 40% of the wall clock; the dictionary interleave alone was 11.5 ms. Parallelizing those is where the 9 ms came from. Amdahl gets his rent eventually.

Parsing was allocation-bound, not CPU-bound. 42 MB and 801k objects to read a 5.8 MB file: gomarc’s reader copied every record into two fresh buffers, and eight GetFields-style calls per record each built a tag map and rescanned every field. Framing records here, decoding in place, and flattening in a single pass took churn to 24.6 MB. A test checks that single pass against those convenience accessors for all 5,572 records, so they remain the specification.

And the ones that measured slower and were reverted, which is the honest half of any list like this: a 4-ary sift (extra comparisons per level cost more than halved depth), a heap over the shard heads in the dictionary merge (a dozen string comparisons beat its bookkeeping), splitting one column’s fold across cores, and building subgraph edges by visiting the union of the members’ documents instead of walking each member’s adjacency — 15% slower, because the union is nearly as large as the sum of the degrees.

07Measured, at both scales

The interesting question about a search engine is not how fast it is on one file, but what happens when the file gets thirty times bigger. The benchmarks take MARCSEARCH_CORPUS so the same numbers can be taken over either corpus.

170,938 records — 256.7 MiB index, 1,574 B/doc
Ingest, file to queryable index 1.03 s (191 MB/s)
Indexing already-parsed documents 767 ms (222,809 docs/s)
— of which the graph 62 ms for 2.48M edges (141 B/doc)
Rare term (482 matches) 21 µs
Ordinary query (4,678 matches) 175 µs
Phrase 247 µs
With 3 facets 243 µs
Broad prefix, 138,595 matches 10.6 ms
/suggest 2.7 µs
Barcode lookup 91 ns
Search, 12 cores in parallel 15 µs/query
End to end over HTTP Throughput p50 p99
The request the demo UI makes, 12 readers 31,685 req/s 0.33 ms 1.13 ms
The same at 48 connections 41,873 req/s 0.99 ms 4.32 ms
/suggest 60,219 req/s 0.17 ms —
A drawn graph map 32,268 req/s 0.33 ms —

The slowest thing a reader can ask for is not a query at all: a bare filter=doctype:article matches 143,000 records and takes 3.3 ms, because there is no term to drive an intersection and every matching document is visited. A query for the — 61,302 matches — takes 2.0 ms for the same reason. Both are still under one animation frame.

5,572 records — single-threaded, limit=20 Latency Allocations
Rare term (kinesitherapie, 19 hits) 1.5 µs 5
Common term (data, 176 hits) 7.9 µs 5
Fuzzy (oncolgy, distance 1) 9.8 µs 5
Phrase, 54 matches 11.1 µs 10
Two wide facets (journal 2500, subject 4654 values) 15.8 µs 9
Co-authors of an author (graph) 0.51 µs 1
Path, author to author, 4 hops 0.31 µs 3
Personalized PageRank, 930 nodes reached 14.7 µs 1

All twelve threads querying the same index: 1.2 µs/query ≈ 830k queries/s, and 0.13 µs per two-hop graph traversal. That is what immutability buys — there is nothing to contend on. At the other end of the scale, ANet’s union catalogue: 3,591,563 records and 4,594,786 real copies, 3.78 GB, 26 s to parse on 12 cores, 2.7 GiB of index at 822 bytes/doc. Search stays at 2 ms.

08The graph, which needs no query

Search assumes the reader already knows what to ask for. A bibliographic corpus is also a dense web of relations, and none of it was queryable: the facet columns map a document to its value codes, and nothing mapped a value back to its documents. The graph is those columns flattened into one CSR array plus its transpose — derived from the merged columns, so it inherits their shard-independence for free.

  • Bipartite: documents on one side, values on the other. Making author a facet column meant the graph reused the existing deterministic column merge unchanged — and /search gained facet=author at no cost.
  • One flat uint32 id space: documents take [0, docs), then each kind takes a contiguous run. A kind is a range check, not a tag — restricting a two-hop expansion to subjects costs two comparisons per candidate and no memory.
  • Plain uint32 adjacency, not varint. A traversal jumps to an arbitrary row and reads it whole, so a fixed stride keeps the walk a pointer chase. Varint would save ~40% of the adjacency and add a decode to every edge visit on the hot path of every primitive.
  • The transpose is built serially on purpose. Counting and scattering 68k edges is two linear passes of ~150 µs; partitioning it would have every worker rescan the whole forward array. Not every loop wants a go in front of it.
  • Every primitive uses the facet counter’s trick: a dense array over the node space plus a list of the slots it touched, so undoing a traversal costs what it reached, not the corpus size. In steady state a traversal allocates exactly one object — the slice it returns.

Hubs are the one thing a bibliographic graph gets wrong on its own. Every document is two hops from every other through its language or its year, so the shortest path is almost always that — and “both published in 2015” is not a connection. A value carried by more than 2% of the corpus is treated as carrying no signal and is not routed through. That is also what makes these queries fast: hubs are the expensive nodes to expand.

The map is a cross-filter, and a cross-filter has to be one query

Drawing a neighbourhood and scrubbing a time range are normally two views over two requests, and they drift — the picture is of the whole corpus while the slider counts only the current page, or the bars redraw themselves as they are dragged. /graph/map answers both halves off one walk: the subgraph is what the window contains, so a value carried only by excluded records leaves the picture on its own; the timeline is the seeds’ whole extent in time, window included and excluded alike, so the bars never move under the handle dragging them.

A window is a comparison per document, not a search per document: the graph keeps every document’s year in a []uint16. Expanding the busiest subject costs 6.7 µs with a window against 6.6 µs without — the test disappears into the row reads.

And a map is a map of something: /graph/map takes the search endpoint’s own q, filter and available, and what reaches the traversal is a *DocSet — the query’s matching stage returned as a bitset. In the interface that is one link in each direction: Draw as a map on a result list, Show as a list on a map, with the selection surviving both crossings, because a seed is a filter chip and the scrubber’s window is a year:2015..2020 filter. Two spellings of one address, not two applications.

09A record is not a thing

The modelling half, which is where a catalogue is actually won or lost. The same publication gets described more than once, and there are two different questions here that get conflated:

  • Same publication? Then one record can stand for the rest in a list — the reader cannot tell them apart either. Key: folded title + first author + year + journal + volume. 2,047 groups, 2,403 records behind a row.
  • Same writing, published more than once? Then they are different things the library holds, each keeps its row, and the record page names the others. Key: the same, minus journal and volume. 4,430 groups, 5,096 records.

The corpus forced the split: of the groups sharing title, author and year, 55% name a different journal. The clearest case is a 2025 editorial published simultaneously in eleven toxicology journals — collapsing on the work key alone makes the biggest number and hides ten real publications behind one row.

What it refuses to do matters as much. Fingerprinting the title alone finds 7,745 groups, which sounds better until you read them: 152 records are called Introduction, 49 Inleiding, 44 Preface.

What it still gets wrong, because a claim like this is worth nothing without one: the largest group is 37 records titled Profiles: 24 Essays for Vibraphone and Piano — one author, one year, no journal, no volume, nothing in the MARC that tells them apart. They are 37 separate deposits and they collapse into one row. The catalogue cannot separate them because the MARC does not, so the row says 37 records state the same thing and opens them, rather than claiming they are one work.

An ISBN is a host, not a workOne ISBN covers 39 records that are 39 chapters of one book. Read as a chapter→book link instead, it joins 3,071 chapters to their book — and states the finding plainly: 76% of the ISBNs in this catalogue name a book the catalogue does not have.

An institutional repository holds the chapter its staff wrote, not the book around it. That is also what fixed the holdings: a chapter is not a thing on a shelf — the book it is in is.

And a name is not a person. Roels, Frank has 53 records; Roels, F. has 17; they are two nodes because they are two strings, and there are 26 more people called Roels. There is no name authority in this codebase and inventing one is its own project — so the person page shows the other spellings beside the work, orders the ones whose initials agree first, and says plainly that it does not know which of them are the same person. An ordering, not a claim.

10Where the copy is

A shelfmark is a promise that the building can be read. Main Library · Open shelves · 616.99 SMI 2018 is a sentence the reader has to turn into a walk, so the catalogue writes the walk out instead:

Floor 4 › rack 001.4 › shelf 5 of 5 › slot 11 of 30

shelf 5 is meaningless without the of 5 that makes it the top one, so the rack’s grid is served beside the position — and those labels are kept once per floor and once per rack, not once per copy.

The building is generated, not surveyed. No repository export carries geometry, so it is derived from the columns that do exist:

library → building one per holding institution collection → floor the collections that hold something in it class group → room shelf classes that sign together: 700s, MAG, 600xxx shelf → rack one per shelf class, as a grid of shelves and slots copy → book placed on the first free slot from a hash of its item key

Two details are the whole difference between a drawing and a plausible one. A catalogue is not one building: the generator used to map library → floor, which is fine for five libraries and absurd for two hundred — a union catalogue came out as a 232-storey tower nobody was ever in. And accession-numbered shelfmarks are bucketed by their thousand first; without that, one ANet member produced 443,971 “classes” for 443,971 copies, one rack each.

The hash is a preference, not a placement. Shelf and slot used to be hash % 5 and hash % 30, which reads like a placement and is not one: over a rack of 150 slots that is the birthday problem, and it put 1,303 copies — 14% of them — inside another book. Now the hash chooses the slot a copy would like and the building gives it the first free one after that, which keeps the building a pure function of the corpus. A rack that genuinely fills up is reported as crowded rather than drawn as if it had fitted.

And it is drawn, not rendered. WebGL is the wrong price for “which floor is it on”. /plan is the 3D viewer’s own document minus the books — 63 KB against 383 KB — and the page draws it as SVG: an isometric projection, two multiplications per point.

x_screen = (x − z)·cos30 y_screen = (x + z)·sin30 − height

Same geometry as the renderer, so the two pictures cannot disagree; SVG, so it themes, prints, scales, and is hit-tested by the browser instead of by a raycaster. It is also a facet: every rack the query touches is lit, the rest go pale, and clicking a rack filters by its library, collection and shelf class — because a rack is a shelf class in a collection in a library, and filtering by the class alone lights the same class on every floor that stocks it.

The 3D wayfinder is still there for looking around — Three.js, vendored into the binary, no CDN — fetched only when somebody asks, and mounted once per session and aimed by postMessage rather than one iframe per copy. Three copies and a record change cost one fetch of the building instead of four. Walk to all 3 sorts the stops by floor and then by distance from that floor’s door, so the route climbs once instead of going up and down.

11Embeddings, and whether they were worth it

internal/embed speaks /v1/embeddings against any OpenAI-compatible base URL, caches what it buys, and feeds -map embedded. Three decisions in it are the interesting part.

The width is a memory decision. At 170,938 records, 3,072 dimensions — the natural width of a current model — is 2.10 GB of vectors against a 257 MiB index. 768 is 525 MB. The default, 256, is 175 MB. The models worth using are Matryoshka-trained, so a truncated vector is a supported output rather than a damaged one — and the width asked for is checked against what arrives, because a provider that ignores the parameter answers at its own width and would put two widths in one store.

The cache is the other half. Embedding this corpus is 40 MB of text and about 10M tokens: cheap to buy once, absurd to buy on every restart, slower than the entire rest of the build by two orders of magnitude. Vectors are keyed by the text they were made from, not by the document number, so the file survives a rebuild and only genuinely new or edited records cost anything. A run that dies at 90% keeps the 90%.

No provider, no quota: a model on the same machine. Ollama and llama.cpp serve /v1/embeddings too, so a local embedding model needs no code and no key. all-minilm at 384 dimensions over the full corpus: 47m33s on 12 CPU cores, index 257 → 509 MiB. That is what the live demo runs.

Was it worth it? The measurement that settles it is not a benchmark, it is a pair of cosines:

Cosine, all-minilm
dispersal ↔ population genetics — no shared words, same subject 0.742
dispersal ↔ antibiotic resistance — both academic, other field 0.584

Two papers on the same species in the same ocean, sharing authors and no vocabulary, recognised as nearer to each other than either is to the word that ought to find them both. And whole queries, against the keyword path:

Query mode=vector mode=keyword
how seeds travel on ocean currents Floating with seeds: hydrochorous mangrove propagule dispersal Distribution of Eocene bivalves
doctor patient conversation about dying Physician discussions with terminally ill patients doctor-patient relationship in chronic fatigue
how cities cope with too much rain Climate change impact on urban rainfall extremes The Too Little/Too Much Scale: A New Rating Format

The last row is the shape of the whole thing: keyword matched the words too much and returned a psychometrics rating scale. The first row reproduces against the live demo in 276 ms, because an exhaustive scan of 170,938 × 384 float32 is exactly what it sounds like — and is exact rather than approximate. HNSW starts paying at millions.

And where it does not help, which is the half that usually goes unpublished: more-like-this was already respectable on the provider-free -map lexical stand-in, because a whole record has enough words to survive hashing. The gain is concentrated in short queries and in crossing vocabularies. At full scale, though, the stand-in is not weaker — it is noise: 240,416 terms hashed into 256 buckets answers cancer with masterclass, three times.

12The assistant

/assist answers a question in prose over Server-Sent Events and shows the records the answer rests on. It is off by default; without both flags the endpoint is 501 and names what is missing.

ollama serve & ollama pull qwen2.5:7b
./marcsearchd -marc vub.mrc -assistant-url http://localhost:11434/v1 \
              -assistant-model qwen2.5:7b -assistant-timeout 15m

It speaks the OpenAI chat-completions protocol rather than any vendor’s SDK — 400 lines of net/http and encoding/json, adding nothing to go.mod, pointable at llama.cpp, vLLM, Ollama, LM Studio, OpenRouter, Azure or OpenAI. That protocol is a common shape rather than a standard, which is why -assistant-param exists: turning a reasoning model’s thinking off is spelled think=false by Ollama, chat_template_kwargs by vLLM and reasoning_effort=low by OpenAI. Fields set that way cannot overwrite the ones carrying the conversation, so no configuration can quietly turn the tools off.

The tools are the endpoints

Nine of them. Six that read — search, get_record, related, about_person, same_publication, cite — each calling the same Go functions the HTTP handlers call, in process. There is no second copy of the query logic and no loopback request, so the assistant cannot answer about a different catalogue than the one on screen.

Three that act — show_results, open_record, open_person — and these are deliberately offers. The server validates and the page performs: a search matching nothing is refused before the reader is offered it, a made-up record never becomes a button, and the button says how many records it would show, because the search was run here to find out. The page does not move on its own — a model that decided to navigate could lose somebody’s place mid-sentence.

A citation is not something the model writes

Every record any tool hands back is recorded by the server, and the sources frame is that list. The model never writes it, so it cannot put a record in it that the catalogue did not return: an invented “Smith 1999” can appear in the prose and will be absent from the sources, where the reader is looking.

The done frame carries grounded, false when nothing was looked up — and the page marks it, in those words: no records behind this answer; nothing was looked up, so there is nothing here to check it against. It says that even when the answer is a correct refusal, because the page cannot tell a good refusal from an invention, and the honest label is the same for both.

A model’s output is text from outside the program, so the page never puts it in innerHTML: the markdown a model writes is rendered by building nodes and setting textContent, which means a model that emits a script tag emits the characters of one. And the system prompt forbids the three lies this corpus makes easy — no abstracts, no disambiguated names, invented holdings. A test pins each of those sentences, because a prompt is the kind of file somebody tidies.

What one question costs

A tool loop is not one request: each round resends the system prompt, the tool schemas and every earlier tool result. The done frame reports it, so this is visible without putting a proxy in front of your own server:

event: done
data: {"took_ms":10425,"sources":7,"grounded":true,"rounds":3,
       "prompt_tokens":7080,"cached_tokens":0,"completion_tokens":952}

About 7,000 input and 450 output tokens per question, half of it a 1,131-token prefix re-sent once per round — exactly what a provider’s prompt cache serves at a tenth of the price. At August 2026 list prices, with gemini-3.1-flash-lite averaging 4,042 in and 273 out:

Model One question 1,000 50,000
gemini-3.1-flash-lite $0.0007 $0.71 $36
gpt-4o-mini $0.0008 $0.77 $39
gpt-5.4-mini $0.0043 $4.26 $213
gpt-5.4 $0.017 $17 $835
gpt-5.6 Sol $0.033 $33 $1,669

Fifty thousand questions is a busy year for a university library’s catalogue, and on a lite tier it is the price of lunch. The expensive tiers buy nothing the tools do not already supply — this is retrieval and citation over structured data, not reasoning.

Whether a cheap model can be trusted with it

Cost is the easy half. The question that matters is whether a model small enough to be cheap will still refuse to say the three things this corpus makes easy to say wrongly. Five questions, three of them guardrails, against the full corpus:

Question gemini-3.1-flash-lite gemini-3.5-flash-lite qwen2.5:7b (CPU) llama3.2:3b (CPU)
who publishes on X 2.2 s ✓ 4.2 s ✓ 59 s ✓ 31 s ✗ departments as authors
three papers + BibTeX 8.5 s ✓ 3.9 s ✓ 223 s ✓ —
an author who does not exist 2.1 s ✓ 2.3 s ✓ — —
can I borrow a copy today 1.3 s ✓ 0.6 s ✓ — —
is this all of their work 2.7 s ✓ 2.3 s ✓ — —
grounded when it should be yes yes no no

The guardrail answers are the reason to believe the rest:

The catalogue lists 17 records for “Roels, F.” published between 1981 and 2007. It is impossible to say if this is all of their work. The catalogue does not disambiguate author names; there are 26 other spellings of the surname “Roels” listed, such as “Roels, Frank,” which may or may not refer to the same person.

Both numbers in that answer are right, and it is not a hedge the model could have produced without looking. 7B instruct is the floor for this tool set. Below it, tool calling exists but the choices are wrong often enough to be worse than no assistant — llama3.2:3b asked related for departments and then presented department names as authors. Reasoning models are actively bad here: qwen3:4b spent more than ten minutes on thinking tokens for one question and timed out twice.

Two protocol details a provider needs that the protocol does not describe, both of which cost a debugging session. OpenAI numbers parallel tool calls with an index and sends the id once, while Gemini omits the index entirely — keying on the index puts every Gemini call at 0 and concatenates two searches into one tool named searchsearch. And Gemini’s OpenAI-compatible endpoint attaches a thought_signature to every tool call and rejects the next round with a 400 if it does not come back. So whatever a server hangs on a tool call is carried back unread — this package does not need to know what a thought signature is in order to return one.

13Honesty as a design constraint

This is a demo with invented holdings, and the code spends real effort on not lying about it.

  • Real and invented are kept apart. What records say about each other is computed from the MARC; everything about copies is invented. A derived number inherits its weakest input, so “the book this chapter is in is held by 3 libraries” is a real relation over invented holdings and is labelled invented. Every block in /doc and /stats carries its own computed_from, so a reader who deletes everything marked generated is left with exactly what the MARC said.
  • Not knowing is not the same as being on the shelf. With no availability feed, available=1 matches nothing and no copy claims a status, rather than the catalogue quietly asserting that everything is in.
  • A hold the library does not know about is not a hold. The request queue is real and ordered — server-side, under a cookie pseudonym, no account — but nothing behind it talks to a library system, so the API returns "simulated": true and the button says simulated request. Keeping “holds” in localStorage would look like a reservation and be nothing at all.
  • A demonstrator that lies about being ready is worse than one that will not start. A remote provider with no key used to start cleanly and log assistant … at https://… exactly as if configured — and then every question came back Missing or invalid Authorization header from a provider the reader has no reason to know exists. A non-loopback endpoint with no key is now refused at startup, naming what to set.
  • A citation that states something false is worse than an incomplete one. 12,315 records give the journal as “Unknown Journal”, “Unknown” or “-”, and 622 of 28,811 020 values are not ISBNs (“geen”, “NoISBN”, a WorldCat URL). /export drops them.

Availability is the same shape as everything else: a snapshot published whole behind an atomic pointer, exactly the way the server holds the index. A handler loads the pointer once, so the badge on a card, the count in the summary and the “available now” filter all describe the same instant. And “available now” is a real search dimension, not a filter on the page: the snapshot derives one bit per document and that DocSet enters the query as an ordinary compiled filter, which is what lets crossfilter=1 leave it alone and probe=available:1 answer what it would leave before anyone clicks it.

Two counting rules, and the view that had to go

A column must not narrow itself. Counted the ordinary way, pinning doctype:book leaves the doctype column showing one value — so the only way to look at articles is to clear the pin first. crossfilter=1 counts each column with its own filter relaxed and everyone else’s applied. A hover is a question, not a commitment: probe= counts the facets as if a filter were applied and reports how many records it would leave, without moving the results. Four columns over the whole corpus cost 330 µs with both rules against 315 µs with neither — interactivity for about 5%.

There used to be a whole view built on this: three live facet columns, a year window on two sliders, no search box, one /search per animation frame. It was deleted, and that is the part worth recording. Once the facet rail beside the ordinary results started counting the same way, the view’s distinguishing idea was on every page that has facets, and what was left was a second, worse way to reach it — its own hash, its own 277-line module, no query box, no place in the view switch. A view whose distinguishing idea is everywhere is not a distinct view. The counting rules stayed; the module did not.

14The front end is seven files and no build step

index.html is 413 lines: markup, and the list of files it loads. The script is cut along the section banners it already carried — as classic scripts, not modules, on purpose: consecutive classic scripts share one scope and run in order, so the split changed no semantics. The only thing that can break is hoisting across a boundary, so every cut was checked for exactly that. What it does not do is make the state explicit — thirty globals are still thirty globals, and nobody should mistake splitting for untangling.

Everything is //go:embed-ed and served through an allowlist, with the path cleaned so nothing outside the embedded tree is reachable however the request spells it. No bundler, no npm, no CDN. The force layout, the brush and the hit testing are about 300 lines of dependency-free JavaScript. And every view is a link — #search?q=…&sort=year, #record/<id>, #person/<name>, #map?…&from=…&to=…, #plan?q=…, #ask — where a record’s id is whatever identifies it: the repository GUID, the internal number, or a barcode, which is what makes a scanned code a link.

Keyboard and screen reader are not polish. The EU Web Accessibility Directive covers public-sector bodies, and a demonstrator shown to a university that cannot be operated by keyboard has a short conversation. There was no h1 anywhere (the masthead was a bold tag), no skip link, and not one aria-live in the codebase — a search returning nothing told a screen reader nothing at all. The map and the building were labelled role="img" while containing fifty-one described nodes and a set of focusable racks. The year brush could only be set by dragging and cleared by double-clicking, which are two gestures a keyboard does not have: arrows move the window, shift widens it, Home and End jump to the corpus’s ends, Escape clears it.

15Try it

The demo is at vatas.s.zauberto.nl, behind basic auth — vub / sollicitatie. Everything below works against it.

S=https://vatas.s.zauberto.nl
C='curl -s -u vub:sollicitatie'

$C "$S/stats" | jq                                          # the numbers at the top of this page
$C "$S/search?q=oncolgy&fuzzy=1" | jq '.terms'              # typo tolerance
$C "$S/search?q=mangrove&collapse=1" | jq '.total, .rows'   # records vs. publications
$C "$S/suggest?q=mangro" | jq                               # typeahead, ~70 µs
$C -N "$S/stream?q=mangrove&facet=doctype"                  # SSE, one frame per hit

# the embeddings, and the query keyword search cannot answer
$C "$S/search?q=how+seeds+travel+on+ocean+currents&mode=vector&limit=3&include=title" \
   | jq '.hits[].title'

$C "$S/graph/path?from=author:koedam,%20nico&to=subject:oncology" | jq
$C "$S/graph/map?q=oocyte&filter=doctype:article" | jq keys # picture + timeline + records
$C "$S/person?name=Roels,%20Frank" | jq '.records, .same_surname[0]'
$C "$S/shelves?q=mangrove" | jq                             # which floors hold the answers
$C "$S/export?format=bibtex&q=mangrove&collapse=1" -o mangroves.bib

$C -N "$S/assist?q=who+at+VUB+publishes+on+mangroves"       # the assistant, streaming

That last one answers, live, in about three seconds:

event: tool      {"tool":"search","arguments":"{\"q\":\"mangroves\"}"}
event: tool_done {"tool":"search","sources_so_far":10}
event: text      "…researchers who publish on mangroves include Farid Dahdouh-Guebas,
                  Nico Koedam, Katrien Quisthoudt, Viviana Otero, Tom Van Der Stocken…"
event: tool      {"tool":"show_results", …}

In the browser: search something, then use Map on the toolbar to draw the same result set as a graph, drag the timeline under it, and Show as a list to go back — the selection survives both crossings. Open a record and hit Where is it for the drawn floor plan, with the 3D tab beside it if you want to walk the building. Then ask the assistant something, and check its prose against the source list it could not have written.

What is not there

Deliberately out of scope: persistence of a built index (rebuilding 5,572 records is ~55 ms, and the whole 170,938-record corpus 3.6 s, from the MARC file, cold), incremental updates, and distribution across machines. An approximate vector index is not there either; the exhaustive scan is exact and only starts costing at millions of documents.

The broad typeahead is at the floor this design allows: exact totals and facet counts require visiting every match, so there is no WAND-style skipping to be had. What is left is decoding 14,816 postings and scoring 4,421 documents. That is the honest end of an optimization list — the point where the remaining time is the work itself.

Every number on this page is reproducible: go test ./internal/index -run XXX -bench . -benchmem for the benchmarks, and took_us in every /search response for the live ones.


<< Back to the work