~/blog/documents-database.html

Private Documents to Database

Lab  My notes, Aug 16 2026 · machine-written expansion below the line · ~5 min

4,804 pages of old documents you don’t read, but want a computer to read instead.

The final workflow was easier than I expected. The hard way would have been manually commanding a cloud model to write a pile of scripts, then getting local LLMs to scan the PDFs, OCR them, transcribe, and build everything out from scratch.

The easier path: use Paperless-ngx as a file dump that handles the raw OCR of the documents, then point a local LLM at it to index everything into the database.

The document pipeline Paper is scanned, OCR'd by Paperless-ngx, indexed by a local LLM or an encrypted cloud alternative, and ends as a searchable database. 20 lbs of paper · 20 years bills, tax, receipts, legal Document scanner the only genuinely manual step about one week of feeding Paperless-ngx a file dump that does the raw OCR Local LLM — indexing or Maple AI with GLM / Kimi cloud is faster — but you trust the claim Searchable database 4,804 pages · a news feed of things to do
Paper in, queryable answers out.

The faster alternative

Another method — possibly better — is using an end-to-end encrypted cloud service like Maple AI with high-end open-source models like GLM or Kimi. It does the job more efficiently, saving a lot of time while producing quality output.

The only real risk with Maple is trusting that it does what it advertises. That’s the trade: speed and quality against having to take a privacy claim at its word, instead of enforcing it yourself on your own hardware.

Why bother

For old family data that’s been piling up for over 20 years, this is a good way to clear the clutter out. Twenty pounds of paper, with an effective document scanner, should only take about a week to get through.

Private documents span bills, tax documents, receipts, legal paperwork, and everything adjacent. Once they’re in, any new document coming through helps build a quick news feed of things to work on. And numbers to crunch.

Below this line: written by Claude

The Maple trade-off, stated plainly

The post calls the cloud route a matter of “trusting that it does as advertised.” Worth making that concrete, because it’s the whole decision.

End-to-end encryption means the provider says it cannot read your documents. You cannot verify that claim from outside — you are trusting an architecture description and whatever audit backs it. The local route replaces trust with enforcement: the documents never leave the machine, so the claim is structural rather than promised. That is strictly stronger, and it is slower and more work. For twenty years of family paperwork, which is exactly the category where a breach is unrecoverable, the slower option is defensible even when the faster one is probably fine.

Where OCR actually fails

The failure mode to watch is silent: a page that OCRs to garbage still gets indexed, and simply never appears in a search. The absence looks identical to “that document doesn’t exist.” Spot-checking a random sample of the oldest and worst-quality scans catches this; nothing in the pipeline will flag it for you.

The question the project raises

Once the paper is searchable, its physical copy is redundant for everything except documents whose value is the physical object — anything notarized, sealed, or original-only. Those stay. The rest becomes a retention decision rather than a storage problem, which is a meaningfully different question than the one the project started with.

Sources & related work

← The Weekly Newspaper & the CorpusAll posts →