01What it is
A MARC21 file arrives from an OAI-PMH endpoint. It is one blob of length-prefixed binary, tens of thousands of records long, and the questions you have about it are ordinary ones: what is in it, is any of it broken, and does record 4,021 look the way the repository claims.
The tools for that are a converter and a text editor, which is two steps too many. lunette is the missing middle — a terminal browser with the record list on the left and the record on the right, reading binary MARC21 and MARCXML, converting between them, and running where the file already is.
lunette records.mrc
lunette records.marcxml # the format is detected, not assumed
lunette part-*.mrc # several files read as one set
It is one static binary with nothing needed at runtime, and it was split out of a harvesting repo — which is the fact that shapes everything below. The records it was written for are not clean library exports. They are what a repository actually emits.
02Five ways to read a record
a c r J X cycle the record pane, and they are five genuinely different questions:
| Annotated | decoded leader, MARC21 field labels, one subfield per line, blank indicators as # — and 880 vernacular fields rendered beside the field they translate rather than stranded at the end |
| Compact | one field per line, subfields inline: 245 10 $a Title $b subtitle. The form to grep |
| Raw | pymarc breaker format, =245 10$aTitle |
| JSON | MARC-in-JSON |
| XML | MARCXML |
Tags, indicators, subfield codes, labels and the leader each carry their own colour, in ANSI-256 indices so they follow the terminal theme instead of fighting it. In the JSON and XML views a search match is marked by inverting the cell, so the syntax colour underneath survives being highlighted.
Long values wrap with continuation lines indented under the value, not under the tag, and an 856 URL too long for the pane is broken rather than left to overflow.
make golden rewrites them; the rule is that you read the diff.03Searching costs something
/ filters, and s cycles where it looks — because the two panes hold different text and the
two indexes cost different amounts.
| Scope | Reads | In the browser |
|---|---|---|
titles (default) |
control number, title, author, year | no prefix |
record |
every subfield and control field | rec: |
both |
either | all: |
Titles is the default because its key is built during load, while the record index has to walk every field of every record and is therefore built only when a scope asks for it.
The distinction is not academic. On a 5,572-record harvest, privacy matches 21 records by title
and 54 by record body — most of the extra hits are 650 subject headings, which is exactly the
kind of thing you were looking for and exactly what a title search cannot see.
tag:856 narrows by field presence and combines with either scope. f filters by the field under
the cursor, o opens its URL, y copies it.
04The leader lies
MARC21 records say what encoding they use in one byte: leader/09, a for UTF-8, blank for MARC-8.
Plenty of OAI-PMH repositories emit UTF-8 and leave that byte blank. Read literally, données
becomes donn©♭es, and it stays that way through every downstream system that trusted the leader.
lunette samples the file, and when the bytes are valid multi-byte UTF-8 with no MARC-8 escape sequences it decodes them as UTF-8 whatever the leader says — then reports the override in the title bar rather than quietly being right.
lunette encoding gives the whole picture, and exits non-zero on a conflict so a harvest script can
gate on it:
lunette encoding records.mrc
It reports the leader/09 distribution, how many records hold non-ASCII, how many carry real MARC-8 escape sequences, how many hold bytes that are neither ASCII nor valid UTF-8, and which records are mislabelled. Notably it reads raw record bytes rather than decoded records, because escape sequences and invalid UTF-8 are precisely what a decoder removes on the way past.
And everything lunette writes is UTF-8, so every exported record is stamped leader/09 a. Passing a
blank leader through is how the problem spreads in the first place.
05Somebody else's bytes
A MARC file came from a repository you do not run, and it is being rendered into a terminal, which is an interpreter. So record content is treated as untrusted:
- Control characters are rewritten as caret notation —
^[— before anything is displayed or copied. A record carryingESC [ 2 Jwould otherwise clear the screen; one carrying an OSC 52 sequence would write to the user’s clipboard. There is a fixture record that does exactly this, and the tests use it. - Links reach the desktop only if they are http or https, carry no control characters, and are of sane length.
- Damaged records are reported with their ordinal and byte offset rather than dropped silently, and anything that reduces the output count says so on stderr.
- A record short enough to crash the parser is contained. gomarc panics on a declared record length below 5 — a negative-length slice allocation — so the reader recovers and ends the walk, since a panic that consumed no input would otherwise repeat forever.
export -orefuses to write over its own input, and over any existing file without-force.
06Watching a harvest arrive
lunette -follow harvest.mrc
Records appear as they are written, through inotify on Linux and kqueue on BSD and macOS, with a burst of writes coalesced into one read.
Two details are the whole feature. It watches the directory, not the file, because a writer that replaces a file by renaming over it leaves a watch on the old inode pointing at nothing. And it reads only whole records — a half-written one at the end of the file is left for the next read rather than reported as damage.
A file that shrinks was replaced rather than appended to, so following stops and says so. Where a watch cannot be established at all, the browser falls back to a one-second timer and says that too; a slow re-check runs alongside the watch regardless, since a watch on a network filesystem can miss writes made by another host.
Following is binary-only. A MARCXML document is not a document until its closing tag.
07Without the browser
The same reader, the same five views, no TUI:
lunette validate records.mrc # exit 1 if a record fails to decode
lunette encoding records.mrc # what encoding the file really uses
lunette show -mode compact records.mrc # print records
lunette show -scope record -filter privacy f.mrc # search inside records, not just titles
lunette export -format xml records.mrc > out.xml # mrc, xml or json
metha-cat -format marc21 "$URL" | lunette show - # "-" is standard input
Colour is on when output is a terminal and off when it is piped, which -color, -no-color and
NO_COLOR override.
A harvest arrives in pieces, so every subcommand except encoding takes any number of files and
reads them as one set — record numbering runs across the whole set, and each reported issue names the
file it came from, so “record 402” means something once there is more than one input. Concatenating
first would work for binary MARC and not for MARCXML, whose documents cannot be joined end to end.
The browser is the one thing that refuses a pipe, because it re-reads the file as the cursor moves
and -follow seeks back into it.
08Try it
Without installing anything:
go run github.com/beyto1974/lunette@latest records.mrc
Or the installer, which fetches the binary for your platform, checks it against the published
checksums.txt and puts it in ~/.local/bin:
curl -fsSL https://raw.githubusercontent.com/beyto1974/lunette/main/install.sh | sh
Reading a script before piping it into a shell is reasonable, and the README says how.
MARC parsing is gomarc, pinned at v0.2.0 because it is pre-1.0 and its API can move; the interface is the Charm stack. Every package was built test-first — 177 tests, an 86.8% line coverage against a 70% CI floor, and roughly as many lines of test as of code.
Source, notes and roadmap: github.com/beyto1974/lunette. MIT.