Lab My notes, Aug 16 2026 · machine-written expansion below the line · ~5 min
● Written by Adriel
4,804 pages of old documents you don’t read, but want a computer to read instead.
The final workflow was easier than I expected. The hard way would have been manually commanding a cloud model to write a pile of scripts, then getting local LLMs to scan the PDFs, OCR them, transcribe, and build everything out from scratch.
The easier path: use Paperless-ngx as a file dump that handles the raw OCR of the documents, then point a local LLM at it to index everything into the database.
Paper in, queryable answers out.
The faster alternative
Another method — possibly better — is using an end-to-end encrypted cloud service like Maple AI with high-end open-source models like GLM or Kimi. It does the job more efficiently, saving a lot of time while producing quality output.
The only real risk with Maple is trusting that it does what it advertises. That’s the trade: speed and quality against having to take a privacy claim at its word, instead of enforcing it yourself on your own hardware.
Why bother
For old family data that’s been piling up for over 20 years, this is a good way to clear the clutter out. Twenty pounds of paper, with an effective document scanner, should only take about a week to get through.
Private documents span bills, tax documents, receipts, legal paperwork, and everything adjacent. Once they’re in, any new document coming through helps build a quick news feed of things to work on. And numbers to crunch.
Below this line: written by Claude
The Maple trade-off, stated plainly
The post calls the cloud route a matter of “trusting that it does as advertised.”
Worth making that concrete, because it’s the whole decision.
End-to-end encryption means the provider says it cannot read your documents. You cannot verify that
claim from outside — you are trusting an architecture description and whatever audit backs it.
The local route replaces trust with enforcement: the documents never leave the machine, so the claim
is structural rather than promised. That is strictly stronger, and it is slower and more work. For
twenty years of family paperwork, which is exactly the category where a breach is unrecoverable, the
slower option is defensible even when the faster one is probably fine.
Where OCR actually fails
Thermal receipts. Older ones fade to blank. If those matter, they needed scanning years
ago — scan the oldest paper first, not the most convenient.
Handwriting. Standard OCR is built for print. Annotations in margins, signatures, and
handwritten ledgers come through as noise or vanish.
Carbon copies and faxes. Low contrast defeats the thresholding step before text
recognition even starts.
The failure mode to watch is silent: a page that OCRs to garbage still gets indexed, and simply
never appears in a search. The absence looks identical to “that document doesn’t
exist.” Spot-checking a random sample of the oldest and worst-quality scans catches this;
nothing in the pipeline will flag it for you.
The question the project raises
Once the paper is searchable, its physical copy is redundant for everything except documents whose
value is the physical object — anything notarized, sealed, or original-only. Those stay.
The rest becomes a retention decision rather than a storage problem, which is a meaningfully different
question than the one the project started with.
Sources & related work
Paperless-ngxSelf-hosted document management. Does the ingest and raw OCR, which is the part
I’d otherwise have written badly myself.
Maple AIEnd-to-end encrypted AI chat. The faster route for bulk indexing — at the
cost of trusting the privacy claim instead of enforcing it on your own hardware.
GLM · KimiThe open-weight models worth pointing at a document pile when you want quality
without a frontier-API bill.
TesseractThe OCR engine underneath the pipeline. Worth knowing it’s there when
scan quality is the thing failing you, not the model.