Adriel's Lab > AI > 3090 Field Manual
One RTX 3090 · Ryzen 9 3900X · 64GB RAM · Ollama
Every rule below came from watching the same box misbehave and then timing exactly why. Nothing
here is a best practice borrowed from a blog post — it’s what this specific 24GB card
actually does when you push it.
- Fit in VRAM or pay for it in minutes, not seconds. The card has ~23,300 MiB usable.
A model that fits runs fully resident; one that doesn’t spills to CPU and the same prompt
that answers in 1.8 seconds resident takes 4–6 minutes once it spills. There is no
graceful middle.
- Context length is a second model you forgot to budget for. Pushing a 27B model to a
131K-token context costs an additional 8.3GB of VRAM on top of the weights. Context isn’t
free headroom — it competes with the model for the same 24GB.
- One big tenant at a time, no exceptions. Two models sharing the card degrades both
rather than dividing cleanly — the fleet-wide rule (one heavy job on the GPU at once, everything
else queues) exists because of exactly this measurement.
- “Thinking” mode is not a free upgrade. The same prompt, same model, with
reasoning/thinking tokens enabled: 305 seconds. With it off: 1.8 seconds. Extended reasoning is a
real, billable-in-time feature — turn it on because you need it, not by default.
- Pick the architecture for the job, not out of habit. Mixture-of-experts models earn
their keep on agentic, multi-step work; dense models win on narrow specialist tasks like
translation, recall, or vision. The roster on this box is deliberately mixed for exactly that
reason.
- Don’t trust the naive disk listing.
ollama list reported roughly
750GB of models on disk. The actual store was 284GB — quantizations of the same base model
share blobs, and the naive per-model listing double- and triple-counts them.
- A stale config fails silently, not loudly. The most expensive bugs on this box
weren’t crashes — they were a service quietly pointed at the wrong port or a model id
that stopped existing, running “successfully” against nothing for weeks before anyone
noticed the output was empty.
Every number above is measured on this box, not sourced from
a spec sheet or a benchmark someone else ran.