Cataloguing

And back into the catalogue

marc-flow proposes; something else has to write. A worked example against a real Koha — one thin biblio out, a webhook receiver in, and the accepted proposals back on the same record.

Suggestions
8
Applied
7
Writable tags
11
Receiver
519 lines
Retries
5

01The missing half

marc-flow ends at a proposal. A record goes in thin, agents read whatever you pointed them at, and what comes back is 650$a with this heading, 300$a with this extent, each with a confidence and the text it was read from. Then you download a file.

Downloading a file is not what a library does. The record came out of an ILS, and it has to go back into the same one, on the same biblio, without disturbing anything the ILS considers its own.

That step is deliberately not in the service. A processor proposes; it never decides, and it certainly never writes into somebody else’s catalogue. So the repository carries a worked example instead — examples/koha, a 519-line receiver that closes the loop against a real Koha, and is the only implementation of that step anywhere in the project.

What is actually Koha-specificOne file. koha.go — get a biblio as MARCXML, put it back. The part worth reading, turning a suggestion into a field edit, is pkg/marcapply, shared with the API itself, so "accept this proposal" cannot mean one thing where it is offered and another where it is written.

02The loop

Koha biblio ──▶ POST /jobs ──▶ marc21 + sru ──▶ proposals │ Koha biblio ◀── PUT /biblios/{id} ◀── receiver ◀── webhook

The example seeds Koha with a deliberately thin acquisitions record for The Hobbit — a 245, an 020, a 250, and nothing else. No author, no extent, no summary, no subjects.

The job it submits is one record made of two files:

{"files": [
  {"source": "marc", "format": "xml", "data": "<the biblio as Koha holds it>"},
  {"source": "isbn", "isbn": "978-0-547-92822-7"}
]}

Both halves matter. The first is what your catalogue says; the second sends the SRU agent to the Library of Congress for the same ISBN. The proposals are the difference between the two, which is a far more useful thing to review than a description generated from nothing.

For the run in the repository that difference is eight suggestions: 100$a Tolkien, J. R. R. at 0.95, a fuller 245$a at 0.88, 300$a, 520$a, two 650$a headings, a corrected 250$a, and a 490$a series statement at 0.62 that is probably wrong.

03Which biblio

The jobs API has no caller-supplied metadata field, so there is nowhere in the request to write down “this is biblio 42”. It rides in the hook path instead:

POST /hooks/biblio/{biblio_id}

The submitting script chooses that URL, so the answer travels with the question. It also fits the delivery contract as it actually is: marc-flow does not sign webhooks. The URL is the secret, which is a thin one — so the receiver belongs on an internal network, and the compose file publishes its port only so the walkthrough can curl at it.

04What may be touched

An allowlist, and a short one:

100 245 250 260 264 300 490 520 650 655 700

Everything else is refused. 001, 003 and 005 belong to the ILS. Koha’s 999 carries the biblionumber in $c/$d, and a suggestion that damaged it would detach the record from its own holdings. Control fields — 008 and friends — are never written from here at all, even though the service is happy to propose byte-level edits to them: correcting 008/06 is a cataloguer’s call, not a webhook’s.

The other half of not doing damage is what gets sent back. The receiver fetches the biblio, edits that record in place, and puts the same record back — so every field outside the allowlist survives the round trip because nothing ever rebuilt it.

Within the allowlist, repeatability decides the verb:

020, 490, 650, 655, 700 a new value becomes a new field — a second subject heading does not evict the first
everything else the subfield is overwritten

A field created from scratch needs indicators, and a suggestion does not carry any, so they come from a small per-tag table — 650/655 get ind2 4, “source not specified”; most others blank. The example says out loud that real cataloguing would decide these properly. It is the honest gap in the piece.

05The mode travels with the job

The receiver does not invent an acceptance policy. It reads the one the job already carries — the same auto / manual / numeric confidence mode the interface sets — and applies it:

auto      apply everything
manual    apply nothing, report what it would have done
0.75      apply where confidence >= 0.75

An unrecognised mode is treated as manual, which is the direction a parsing failure should fall in when the outcome is a write.

On the canned job that is seven of eight applied: 100, 245, 250, 300, 520 and two 650s. The 490$a at 0.62 sits below the bar and is refused, and it is refused by name in the response, alongside everything that was written:

{"written": true, "applied": [...], "skipped": [{"field": "490$a", ...}]}

DRY_RUN=1 runs the whole decision and stops before the PUT. It is the setting to point at a real catalogue first.

06When it is delivered twice

A failed delivery is retried up to five times with backoff. So the receiver has to be idempotent, and the requirement is met where it belongs — in the applier: a value already on the record is a skip, not a second field. Delivering the same job ten times leaves the biblio in the state one delivery left it in.

The status codes are chosen for that retry policy rather than for tidiness:

200 applied, or nothing to apply, or a job that reported failed
404 no such biblio — permanent, do not retry
502 Koha 5xx, a timeout, a dropped connection — worth another attempt

A job that failed still fires its webhook, and answering 2xx to it is correct: there is nothing to apply, and the delivery itself worked.

07Run it

The example talks to a Koha you already run; it does not ship one. The community’s koha-testing-docker is the easy way to get one — turn on RESTBasicAuth, and make a patron with editcatalogue → edit_catalogue, because Koha’s database user is not a patron and cannot authenticate against the REST API.

cd examples/koha
cp .env.example .env
docker compose -f ../../docker-compose.yml -f docker-compose.yml up -d --build koha-webhook

./scripts/01-seed-biblio.sh                    # prints the biblionumber
./scripts/03-show-biblio.sh <biblionumber>     # before: 020, 245, 250
./scripts/02-submit-job.sh <biblionumber>
./scripts/03-show-biblio.sh <biblionumber>     # after: 100, 300, 520, 650 …

No LLM key, or no appetite for one: testdata/job-completed.json is a finished job exactly as the API sends it, carrying the real difference between the seeded record and the LoC record. Replay it straight at the receiver and the Koha half runs on its own.

One test replays that same payload against the same seed record and asserts the outcome described here — seven applied, one refused — so the walkthrough cannot quietly drift away from the code.


<< Back to the work