01What it is
If you work with library data, you work with MARC21 — the length-prefixed binary record format catalogues have run on since the 1960s, complete with a directory of field offsets, subfield delimiters, and a pre-Unicode character encoding called MARC-8 that needs a lookup table with thousands of entries to decode.
In Python that problem is solved. pymarc is mature, complete and pleasant to use. In Go it was not.
gomarc is a port of pymarc to Go. It covers the binary MARC21 transmission format, MARC-8 to Unicode conversion, MARCXML and MARC-in-JSON — and on real catalogue exports it runs 4x to 11x faster than the library it was ported from.
go get github.com/beyto1974/gomarc
This is the parser underneath the other two pieces here: it is what marc-engine decodes 170,938 records with, and what marc-flow is required to use rather than hand-rolling a MARC reader.
02It reads like pymarc
If you know pymarc, you already know this API. That is the point of a port: the reason to reach for it is that the Python you already wrote translates line for line.
Iterate records, pull the fields you want:
reader := marc.NewReader(f)
for {
record, err := reader.Next()
if errors.Is(err, io.EOF) {
break
}
if err != nil {
log.Println(err) // permissive: bad records are skipped, not fatal
continue
}
title, _ := record.Title()
fmt.Println(title)
}
The reader is permissive on purpose. A catalogue export is not a clean file, and a parser that fails the batch on record 40,001 is a parser nobody can run overnight.
Title, Author, ISBN, ISSN, Subjects, Publisher, PubYear and more are there as
methods. For anything else, go at the tag and subfield directly:
value, ok := record.Get("245").Subfield("a")
for _, f := range record.GetFields("650") {
fmt.Println(f)
}
Build records, modify them, write them back:
record.Get("245").SetSubfield("a", "The Zombie Programmer : ")
writer := marc.NewWriter(out)
writer.Write(record)
And convert to the formats the rest of a stack can actually read — both use UTF-8 throughout instead of MARC-8, which is what lets standard tooling touch them at all:
s, err := record.AsJSON() // MARC-in-JSON
records, err := marc.ParseXML(r) // MARCXML
A large MARCXML file streams one record at a time through marc.NewXMLReader rather than loading
into memory.
03The numbers
Two real catalogue exports — 138,076 records, 166 MB. AMD Ryzen 5 3600, Go 1.25.12, CPython 3.13.5, gomarc v0.1.0, pymarc 5.4.0, single-threaded, fastest of three repetitions.
50,000 records, 58 MB
| scenario | gomarc | pymarc | speedup | gomarc rec/s | pymarc rec/s |
|---|---|---|---|---|---|
| parse (MARC-8 to Unicode) | 1.96 s | 21.54 s | 11.0x | 25,545 | 2,321 |
| parse (force UTF-8) | 0.82 s | 5.10 s | 6.2x | 61,096 | 9,813 |
| parse + field access | 2.02 s | 23.01 s | 11.4x | 24,758 | 2,173 |
| parse + write MARC21 | 4.69 s | 26.01 s | 5.5x | 10,661 | 1,922 |
| parse + write MARCXML | 3.23 s | 36.78 s | 11.4x | 15,481 | 1,360 |
| parse + write MARC-in-JSON | 8.18 s | 32.22 s | 3.9x | 6,110 | 1,552 |
88,076 records, 108 MB
| scenario | gomarc | pymarc | speedup | gomarc rec/s | pymarc rec/s |
|---|---|---|---|---|---|
| parse (MARC-8 to Unicode) | 3.33 s | 35.74 s | 10.7x | 26,459 | 2,464 |
| parse (force UTF-8) | 1.53 s | 7.12 s | 4.7x | 57,662 | 12,370 |
| parse + field access | 3.72 s | 34.55 s | 9.3x | 23,663 | 2,550 |
| parse + write MARC21 | 5.48 s | 38.45 s | 7.0x | 16,059 | 2,290 |
| parse + write MARCXML | 5.62 s | 53.66 s | 9.5x | 15,674 | 1,641 |
| parse + write MARC-in-JSON | 7.09 s | 45.67 s | 6.4x | 12,430 | 1,929 |
In practical terms: a full MARC-8 parse of 88,076 records drops from 36 seconds to 3.3. A catalogue-to-MARCXML conversion drops from 54 seconds to 5.6.
Peak memory stays between 11 and 23 MB across every scenario, because gomarc streams — file size does not drive memory, which is the difference between converting a catalogue on a laptop and needing a machine for it.
04Why you can trust them
Speed claims about a port are cheap. A benchmark of two libraries is really a benchmark of two programs somebody wrote, and if one parses lazily while the other eagerly materialises everything, the ratio means nothing at all.
So the suite proves equivalence before it reports a single timing. Three checks, in order:
- Identical record acceptance. Both libraries run permissively and report the same counts — 50,000 and 88,076 records, zero errors — so they agree on exactly which records are well-formed. Two parsers disagreeing about what is broken would make every later number a comparison of different work.
- Identical decoded text. Summing the codepoint length of every extracted title, author, ISBN and subject gives the same total from both libraries: 671,999. MARC-8 decoding agrees character for character, which is the check the thousands-of-entries lookup table exists to fail.
- Byte-identical output. Read a file with each library, write every record back out as binary MARC21, and compare:
$ cmp py.marc go.marc && echo BYTE_IDENTICAL
BYTE_IDENTICAL
That last one is the practical point. Swapping gomarc into a pipeline that pymarc currently feeds does not change the bytes coming out the far end — so the thing downstream of it, which is usually somebody’s ILS, never finds out.
The harnesses, the runner, the raw per-run JSON and the generated tables are all published alongside the library, so the whole thing re-runs:
REPS=3 ./run.sh
python report.py results.jsonl
The report carries a spread column — slowest repetition over fastest — because a benchmark that
publishes an estimator without its noise is not worth much.
05Where the time goes
The biggest win is the MARC-8 decode path, which is where real catalogue data spends most of its time — 10.7x and 11.0x on the two files, against 4.7x and 6.2x for the same parse with UTF-8 forced.
That is not gomarc being clever. It is what happens when thousands of per-character table lookups stop being Python-level operations: force UTF-8 and most of them disappear, and the gap narrows to roughly five times, which is about what any Go-versus-CPython port of the same algorithm should look like.
Which is worth stating plainly, because it decides whether this library is worth adopting. If your records are already UTF-8 you are buying a 5x parser. If they are MARC-8 — and a catalogue that has been running since before Unicode existed generally is — you are buying an 11x one.
06What shipped since
The benchmarks above are v0.1.0, and the library is at v1.0.0 now. Four things changed that a reader of the original write-up should know about:
to_unicode=falseis implemented. pymarc’s raw mode was the first version’s headline gap;WithToUnicodeandWithForceUTF8are record options now, so undecoded bytes are reachable rather than an error.- The readers were rewritten to stream tokens. Zero-allocation tag normalisation, integer parsing straight off the byte slice, and a reflection-free MARCXML decoder — the paths the benchmark measured are not the paths that ship today, and the numbers above therefore understate the current library rather than flattering it.
- An oversize record is refused rather than written. MARC21 encodes a record’s length in five digits, so a record over 99,999 bytes cannot be represented — the writer used to emit a leader that lied about it. It now returns an error, which is the only honest thing a writer can do with a record the format cannot hold.
- A
schemasub-package. Machine-readable descriptions of every named leader position, the 008 fixed field broken down by material type, and the field, indicator and subfield vocabulary for every defined MARC21 tag. It exists because a language model asked to reason about a record needs to be told what008/06means, and that is a dictionary rather than a prompt — marc-flow reads it to decide which subfield codes a tag actually defines.
MARC-in-JSON encoding is still the least optimised corner, at 3.9x to 6.4x. Comfortably ahead, and the obvious next thing to tune.
07Try it
go get github.com/beyto1974/gomarc
Repository, docs and the full benchmark suite:
github.com/beyto1974/gomarc. Test fixtures under
testdata/ are copied verbatim from pymarc’s own suite, so behaviour can be cross-checked against
the original.
Issues and pull requests welcome — particularly if you have MARC data that breaks it. That is the one thing a port of somebody else’s mature library genuinely needs and cannot generate for itself.
A shorter version of this piece was published on dev.to in August 2026.