~/blog/newspaper-corpus.html

The Weekly Newspaper & the Corpus

Lab  My notes, Aug 16 2026 · machine-written expansion below the line · ~6 min

11,204 podcast episodes. About 10% signal. Searchable.

This is a simple project that a lot of people and power users can benefit from: letting your computer or a cloud model process countless hours of podcasts and other content from your favourite creators.

A good example is the Bitcoin space. I was very active listening to podcasts and creators for about a year. After that I got more focused on other things, and started to notice the “noise” in what I was subscribed to. That’s when I got fully exhausted and decided to let the machines watch them for me.

The corpus pipeline Podcast sources flow into local Whisper transcription, then a raw transcript database, then a local or cloud model acting as a librarian, producing the weekly newspaper. 40 creators · 11,204 episodes YouTube + podcast sources, back to 2016 Whisper — local transcription my own GPUs, better than YouTube captions 12,000+ hours in 3 to 21 days Raw transcript database timestamped, searchable, mine Local or cloud model the personalized librarian ranks for signal — most of it is filler The Weekly Newspaper a year of listening → ~15 minutes a week
The whole pipeline, end to end.

How it runs

I run Whisper on my own computers to load the content and transcribe it locally, which does a better job than the captions YouTube provides. With just a good computer or two, I can feed 12,000+ hours through it in anywhere from 21 days down to 3, depending on how much I throw at it.

The end product is a database of raw transcriptions you can point a cloud or local model at — which becomes your personalized librarian. With new content pumping out daily, that’s enough to fill a week’s worth of material and give it its own newspaper.

The corpus problem

My database spans 11,000+ podcast episodes from 40 creators, stretching as far back as 2016. A large fraction of that is filler — quick to age, sponsor reads, and speculation. It also includes the clickbait titles and thumbnails that make things a nuisance to bother watching in the first place.

Only a tiny fraction is genuinely useful, or insightful — “signal,” as they call it. Imagine spending years as an active listener and pouring that energy into forgettable garbage.

It generalizes

This expands to a lot of other interests. Mindset and self-help, AI building — more examples of things I used to watch a lot of. Adult life is a chaotic one, and it doesn’t help with my scattered attention span.

I highly recommend this as a way to cut as much noise as possible and build a condensed way of digesting information — down to as little as 15 minutes a week.

Related projects:
Character AI lab - Local personas

Below this line: written by Claude

What it takes to copy this

The hardware bar is lower than the numbers suggest. One consumer GPU with enough VRAM to hold a Whisper model does the work; the 3-to-21-day spread above is almost entirely a function of how many machines are running at once and whether the GPU is also doing something else. Storage is trivial — 11,000 episodes of plain-text transcripts is a few gigabytes, far less than the audio it came from.

The real cost is curation. Picking the roster, deciding what counts as signal, and re-checking that judgment as creators drift is the part that can’t be automated, and it’s the part that determines whether the output is worth reading.

Three honest limitations

Why this shape keeps reappearing

The pattern — transcribe everything locally, index it, query it with a model — is being rebuilt independently by a lot of people right now, which is why pullthatupjamie.ai exists and why arriving at it separately isn’t a coincidence. Transcription got cheap enough to run at home before search over your own media got good. That gap is what everyone is filling, and the differentiator is no longer the pipeline; it’s whose roster and whose definition of signal you trust.

Sources & related work

← All postsPrivate Documents to Database →