← All posts

Choosing a GPU for local AI: what the numbers actually mean

Graphics card spec sheets lead with the numbers gamers care about: cores, clock speeds, "AI TOPS." For running a language model on your own machine, two quieter lines matter far more — how much memory the card has, and how fast it can read that memory. This post explains why, shows how to size a model against a card with simple arithmetic, and says when buying a GPU is the wrong move.

The one number that decides what you can run: memory

A language model is, physically, a very large pile of numbers called weights. To generate text at full speed, all of them need to sit in the GPU's own memory (VRAM). The rule for how much space they take is plain multiplication: parameter count times bytes per parameter. NVIDIA's own inference guide gives the example of a 7-billion-parameter model at 16-bit precision (2 bytes per weight) needing about 14 GB[4].

Hugging Face published the same arithmetic for the Llama 3.1 family, and the numbers show how quickly it escalates[3]:

Model16-bit (FP16)8-bit (FP8)4-bit (INT4)
Llama 3.1 8B16 GB8 GB4 GB
Llama 3.1 70B140 GB70 GB35 GB
Llama 3.1 405B810 GB405 GB203 GB

Hugging Face adds an important footnote: those figures are only what it takes "to load the model checkpoint," and don't include working space the software needs on top[3]. Treat them as a floor, not a budget.

What happens if the model doesn't fit? It doesn't necessarily fail. llama.cpp, the engine underneath many popular local-AI apps, supports "CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity"[6]. The layers that fit go on the GPU; the rest run from ordinary system RAM on the CPU. It works, but the slow part sets the pace, and the next section explains why that part is much slower.

The number that decides how fast: memory bandwidth

Generating text happens in two phases, and they stress the hardware differently. First the model reads your whole prompt at once (the "prefill"). NVIDIA describes this as a highly parallel matrix-matrix operation — the kind of work GPUs are built for. Then it produces the answer one token at a time (the "decode"), and each step is closer to a matrix-vector operation that "underutilizes the GPU compute ability"[4].

In decode, the bottleneck is moving data, not doing maths. In NVIDIA's words, the speed at which weights and cached values are transferred from memory "dominates the latency, not how fast the computation actually happens … this is a memory-bound operation"[4]. To produce each new word, the card has to stream essentially the whole model out of memory again.

That gives you a handy back-of-envelope ceiling. Divide the card's memory bandwidth by the size of the model in memory, and you get the most tokens per second you could hope for if nothing else got in the way. It's an upper bound from our own arithmetic, not a benchmark — real software lands below it — but it ranks cards correctly and explains results that otherwise look strange.

Here's one such result. llama.cpp's documentation publishes speed figures for Llama 3.1 8B at different compression levels (on hardware the page doesn't name)[2]:

FormatSizePrompt reading (tokens/s)Text generation (tokens/s)
F16 (16-bit)14.96 GiB92329
Q8_0 (~8.5 bits)7.95 GiB86551
Q4_K_M (~4.9 bits)4.58 GiB82272

Prompt reading, which is compute-bound, barely moves. Text generation, which is memory-bound, speeds up by roughly 2.5 times as the file shrinks by roughly 3 times. That is bandwidth at work: less data to stream per token, more tokens per second.

Reading a spec sheet with that in mind

NVIDIA's comparison page lists both numbers for its current desktop cards. As of September 2026 it shows[1]:

CardMemoryBus widthBandwidth
RTX 509032 GB GDDR7512-bit1,792 GB/s
RTX 508016 GB GDDR7256-bit960 GB/s
RTX 5070 Ti16 GB GDDR7256-bit896 GB/s
RTX 507012 GB GDDR7192-bit672 GB/s
RTX 5060 Ti16 GB or 8 GB GDDR7128-bit448 GB/s
RTX 50608 GB GDDR7128-bit448 GB/s

Three things jump out once you know what to look for.

  1. Capacity and speed are separate axes. The 16 GB RTX 5060 Ti holds exactly as much model as the RTX 5080, but at under half the bandwidth (448 vs 960 GB/s)[1]. It will run the same models, more slowly. For local AI that's often a good trade: a model that fits slowly beats a model that doesn't fit at all.
  2. Bus width is a decent shorthand. On this generation, bandwidth tracks the memory bus width closely, from 128-bit at the bottom to 512-bit at the top[1].
  3. "AI TOPS" is the least useful headline. The RTX 5090 is marketed at 3,352 AI TOPS[7]. That figure measures compute, which mainly helps prompt reading. For the part you actually wait on — words appearing — bandwidth decides.

Power belongs on the list too. NVIDIA rates the RTX 5090 at 575 W of graphics power and recommends a 1,000 W system power supply[7]. The top card may mean a new PSU and a noisier room, not just a bigger purchase.

The hidden cost: context length and the KV cache

Weights aren't the only thing in memory. As a model reads and writes, it keeps a record of every token in the conversation so it doesn't recompute them each step: the key-value (KV) cache. Hugging Face describes it as storing "keys and values of all the tokens in the model's context"[3]. It grows with every token of context.

For Llama 3.1 at 16-bit, Hugging Face lists the cache sizes[3]:

Model1k tokens16k tokens128k tokens
8B0.125 GB1.95 GB15.62 GB
70B0.313 GB4.88 GB39.06 GB

Their own summary: for the small model, "the cache uses as much memory as the weights when approaching the context length maximum"[3]. So a 4-bit 8B model that is only about 5 GB on disk can still overflow a 16 GB card if you ask it to hold a whole codebase or a long document. In practice local apps let you cap the context length, and that setting is often the difference between "fits" and "falls back to the CPU." NVIDIA's guide gives the underlying formula, which scales with the number of layers, the model's hidden size and the sequence length[4].

Quantization: how people fit big models on small cards

The 4-bit column in the first table is why local AI is practical at all. Quantization stores each weight with fewer bits. llama.cpp's own figures for the Llama 3.1 family, using its popular Q4_K_M format, show the scale of it: the 8B model goes from 32.1 GB to 4.9 GB, and the 70B from 280.9 GB to 43.1 GB[2].

It isn't free. The llama.cpp docs note that accuracy loss is measured with perplexity or KL divergence, and that it "can be minimized" with a calibration file (an "imatrix")[2] — minimized, not eliminated. Our opinion, for what it's worth: a 4- to 6-bit version of a bigger model is usually a better buy than an 8-bit version of a smaller one, but the lowest settings (2–3 bits) are worth testing on your own tasks before trusting them.

Putting it together: three worked examples

Here's the sizing method applied end to end. The "ceiling" is bandwidth divided by model size, as explained above — an optimistic bound, not a measured speed.

The pattern: discrete NVIDIA cards give you high bandwidth but cap out at 32 GB per card on this list[1]; large unified-memory machines trade peak speed for the ability to load much bigger models.

When not to buy a GPU for this

Honest limits, because a graphics card bought for this is a serious purchase.

If you do buy: pick the capacity for the biggest model you want to run, including its context; then pick the highest bandwidth you can afford at that capacity. Ignore the TOPS number. That order of priorities will serve you better than any single benchmark chart.

Sources

  1. NVIDIA — Compare GeForce graphics cards (memory size, interface width, bandwidth), accessed September 2026
  2. llama.cpp — quantize tool README (quantization types, sizes and speeds for Llama 3.1), accessed September 2026
  3. Hugging Face — "Llama 3.1 – 405B, 70B & 8B with multilinguality and long context" (inference memory requirements), accessed September 2026
  4. NVIDIA Technical Blog — "Mastering LLM Techniques: Inference Optimization", accessed September 2026
  5. Apple — Mac Studio technical specifications, accessed September 2026
  6. llama.cpp — project README (supported backends, CPU+GPU hybrid inference), accessed September 2026
  7. NVIDIA — GeForce RTX 5090 product page (specs and power requirements), accessed September 2026

Related: Running AI on your own hardware: local LLMs explained · Self-hosting 101: what's worth running on your own hardware in 2026 · How large language models actually work