Adriel's Lab > AI > 3090 Field Manual

[ glyph ]

3090 Field Manual

Seven laws, all of them measured, none of them theoretical.

One RTX 3090 · Ryzen 9 3900X · 64GB RAM · Ollama

Every rule below came from watching the same box misbehave and then timing exactly why. Nothing here is a best practice borrowed from a blog post — it’s what this specific 24GB card actually does when you push it.

  1. Fit in VRAM or pay for it in minutes, not seconds. The card has ~23,300 MiB usable. A model that fits runs fully resident; one that doesn’t spills to CPU and the same prompt that answers in 1.8 seconds resident takes 4–6 minutes once it spills. There is no graceful middle.
  2. Context length is a second model you forgot to budget for. Pushing a 27B model to a 131K-token context costs an additional 8.3GB of VRAM on top of the weights. Context isn’t free headroom — it competes with the model for the same 24GB.
  3. One big tenant at a time, no exceptions. Two models sharing the card degrades both rather than dividing cleanly — the fleet-wide rule (one heavy job on the GPU at once, everything else queues) exists because of exactly this measurement.
  4. “Thinking” mode is not a free upgrade. The same prompt, same model, with reasoning/thinking tokens enabled: 305 seconds. With it off: 1.8 seconds. Extended reasoning is a real, billable-in-time feature — turn it on because you need it, not by default.
  5. Pick the architecture for the job, not out of habit. Mixture-of-experts models earn their keep on agentic, multi-step work; dense models win on narrow specialist tasks like translation, recall, or vision. The roster on this box is deliberately mixed for exactly that reason.
  6. Don’t trust the naive disk listing. ollama list reported roughly 750GB of models on disk. The actual store was 284GB — quantizations of the same base model share blobs, and the naive per-model listing double- and triple-counts them.
  7. A stale config fails silently, not loudly. The most expensive bugs on this box weren’t crashes — they were a service quietly pointed at the wrong port or a model id that stopped existing, running “successfully” against nothing for weeks before anyone noticed the output was empty.

Every number above is measured on this box, not sourced from a spec sheet or a benchmark someone else ran.