Choosing a GPU for local AI: what the numbers actually mean
Graphics card spec sheets lead with the numbers gamers care about: cores, clock speeds, "AI TOPS." For running a language model on your own machine, two quieter lines matter far more — how much memory the card has, and how fast it can read that memory. This post explains why, shows how to size a model against a card with simple arithmetic, and says when buying a GPU is the wrong move.
The one number that decides what you can run: memory
A language model is, physically, a very large pile of numbers called weights. To generate text at full speed, all of them need to sit in the GPU's own memory (VRAM). The rule for how much space they take is plain multiplication: parameter count times bytes per parameter. NVIDIA's own inference guide gives the example of a 7-billion-parameter model at 16-bit precision (2 bytes per weight) needing about 14 GB[4].
Hugging Face published the same arithmetic for the Llama 3.1 family, and the numbers show how quickly it escalates[3]:
| Model | 16-bit (FP16) | 8-bit (FP8) | 4-bit (INT4) |
|---|---|---|---|
| Llama 3.1 8B | 16 GB | 8 GB | 4 GB |
| Llama 3.1 70B | 140 GB | 70 GB | 35 GB |
| Llama 3.1 405B | 810 GB | 405 GB | 203 GB |
Hugging Face adds an important footnote: those figures are only what it takes "to load the model checkpoint," and don't include working space the software needs on top[3]. Treat them as a floor, not a budget.
What happens if the model doesn't fit? It doesn't necessarily fail. llama.cpp, the engine underneath many popular local-AI apps, supports "CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity"[6]. The layers that fit go on the GPU; the rest run from ordinary system RAM on the CPU. It works, but the slow part sets the pace, and the next section explains why that part is much slower.
The number that decides how fast: memory bandwidth
Generating text happens in two phases, and they stress the hardware differently. First the model reads your whole prompt at once (the "prefill"). NVIDIA describes this as a highly parallel matrix-matrix operation — the kind of work GPUs are built for. Then it produces the answer one token at a time (the "decode"), and each step is closer to a matrix-vector operation that "underutilizes the GPU compute ability"[4].
In decode, the bottleneck is moving data, not doing maths. In NVIDIA's words, the speed at which weights and cached values are transferred from memory "dominates the latency, not how fast the computation actually happens … this is a memory-bound operation"[4]. To produce each new word, the card has to stream essentially the whole model out of memory again.
That gives you a handy back-of-envelope ceiling. Divide the card's memory bandwidth by the size of the model in memory, and you get the most tokens per second you could hope for if nothing else got in the way. It's an upper bound from our own arithmetic, not a benchmark — real software lands below it — but it ranks cards correctly and explains results that otherwise look strange.
Here's one such result. llama.cpp's documentation publishes speed figures for Llama 3.1 8B at different compression levels (on hardware the page doesn't name)[2]:
| Format | Size | Prompt reading (tokens/s) | Text generation (tokens/s) |
|---|---|---|---|
| F16 (16-bit) | 14.96 GiB | 923 | 29 |
| Q8_0 (~8.5 bits) | 7.95 GiB | 865 | 51 |
| Q4_K_M (~4.9 bits) | 4.58 GiB | 822 | 72 |
Prompt reading, which is compute-bound, barely moves. Text generation, which is memory-bound, speeds up by roughly 2.5 times as the file shrinks by roughly 3 times. That is bandwidth at work: less data to stream per token, more tokens per second.
Reading a spec sheet with that in mind
NVIDIA's comparison page lists both numbers for its current desktop cards. As of September 2026 it shows[1]:
| Card | Memory | Bus width | Bandwidth |
|---|---|---|---|
| RTX 5090 | 32 GB GDDR7 | 512-bit | 1,792 GB/s |
| RTX 5080 | 16 GB GDDR7 | 256-bit | 960 GB/s |
| RTX 5070 Ti | 16 GB GDDR7 | 256-bit | 896 GB/s |
| RTX 5070 | 12 GB GDDR7 | 192-bit | 672 GB/s |
| RTX 5060 Ti | 16 GB or 8 GB GDDR7 | 128-bit | 448 GB/s |
| RTX 5060 | 8 GB GDDR7 | 128-bit | 448 GB/s |
Three things jump out once you know what to look for.
- Capacity and speed are separate axes. The 16 GB RTX 5060 Ti holds exactly as much model as the RTX 5080, but at under half the bandwidth (448 vs 960 GB/s)[1]. It will run the same models, more slowly. For local AI that's often a good trade: a model that fits slowly beats a model that doesn't fit at all.
- Bus width is a decent shorthand. On this generation, bandwidth tracks the memory bus width closely, from 128-bit at the bottom to 512-bit at the top[1].
- "AI TOPS" is the least useful headline. The RTX 5090 is marketed at 3,352 AI TOPS[7]. That figure measures compute, which mainly helps prompt reading. For the part you actually wait on — words appearing — bandwidth decides.
Power belongs on the list too. NVIDIA rates the RTX 5090 at 575 W of graphics power and recommends a 1,000 W system power supply[7]. The top card may mean a new PSU and a noisier room, not just a bigger purchase.
The hidden cost: context length and the KV cache
Weights aren't the only thing in memory. As a model reads and writes, it keeps a record of every token in the conversation so it doesn't recompute them each step: the key-value (KV) cache. Hugging Face describes it as storing "keys and values of all the tokens in the model's context"[3]. It grows with every token of context.
For Llama 3.1 at 16-bit, Hugging Face lists the cache sizes[3]:
| Model | 1k tokens | 16k tokens | 128k tokens |
|---|---|---|---|
| 8B | 0.125 GB | 1.95 GB | 15.62 GB |
| 70B | 0.313 GB | 4.88 GB | 39.06 GB |
Their own summary: for the small model, "the cache uses as much memory as the weights when approaching the context length maximum"[3]. So a 4-bit 8B model that is only about 5 GB on disk can still overflow a 16 GB card if you ask it to hold a whole codebase or a long document. In practice local apps let you cap the context length, and that setting is often the difference between "fits" and "falls back to the CPU." NVIDIA's guide gives the underlying formula, which scales with the number of layers, the model's hidden size and the sequence length[4].
Quantization: how people fit big models on small cards
The 4-bit column in the first table is why local AI is practical at all. Quantization stores each weight with fewer bits. llama.cpp's own figures for the Llama 3.1 family, using its popular Q4_K_M format, show the scale of it: the 8B model goes from 32.1 GB to 4.9 GB, and the 70B from 280.9 GB to 43.1 GB[2].
It isn't free. The llama.cpp docs note that accuracy loss is measured with perplexity or KL divergence, and that it "can be minimized" with a calibration file (an "imatrix")[2] — minimized, not eliminated. Our opinion, for what it's worth: a 4- to 6-bit version of a bigger model is usually a better buy than an 8-bit version of a smaller one, but the lowest settings (2–3 bits) are worth testing on your own tasks before trusting them.
Putting it together: three worked examples
Here's the sizing method applied end to end. The "ceiling" is bandwidth divided by model size, as explained above — an optimistic bound, not a measured speed.
- An 8B model at Q4_K_M (4.58 GiB[2]) on a 16 GB RTX 5060 Ti. It fits with room left for roughly 16k tokens of context (about 2 GB of cache at 16-bit[3]). Ceiling: 448 GB/s ÷ ~4.9 GB ≈ 90 tokens/s. Comfortable for chat and light coding help.
- A 70B model at Q4_K_M (43.1 GB[2]) on a 32 GB RTX 5090. It doesn't fit, despite this being the most powerful consumer card on the list. You'd be relying on CPU offload[6], and a big chunk of every token would crawl through system RAM. This is where a second card, or a different kind of machine, enters the conversation.
- The same 70B model on a Mac with unified memory. Apple's Mac Studio shares one memory pool between CPU and GPU: up to 128 GB with the M5 Max (460 or 614 GB/s) and up to 512 GB with the top M5 Ultra configuration (1.2 TB/s)[5]. A 43 GB model fits with space for long context. Ceiling on the faster M5 Max: 614 ÷ 43 ≈ 14 tokens/s — usable, not fast. llama.cpp supports Apple's Metal backend[6], so the software side is covered.
The pattern: discrete NVIDIA cards give you high bandwidth but cap out at 32 GB per card on this list[1]; large unified-memory machines trade peak speed for the ability to load much bigger models.
When not to buy a GPU for this
Honest limits, because a graphics card bought for this is a serious purchase.
- If you want the best possible answers. The very largest open models (the 405B row above needs 203 GB even at 4-bit[3]) are out of reach of any single consumer card. If quality is the goal, a hosted model will beat what fits on your desk.
- If you haven't tried what you already own. Hybrid CPU+GPU inference[6] and a 4-bit file of around 5 GB[2] mean a small model can run on a lot of ordinary machines. Try an 8B model first; you'll learn what speed you actually need before paying for it.
- If you need your card's vendor to be fully supported. llama.cpp lists backends for CUDA, HIP (AMD), Vulkan, SYCL (Intel) and Metal[6], but other tools — especially training and fine-tuning frameworks — are often CUDA-first. Check the specific software you plan to use before buying anything that isn't NVIDIA.
- If you'll use it an hour a month. A 575 W card[7] sitting mostly idle is an expensive hobby. In our view, rented cloud GPUs or API access make more sense for occasional use.
If you do buy: pick the capacity for the biggest model you want to run, including its context; then pick the highest bandwidth you can afford at that capacity. Ignore the TOPS number. That order of priorities will serve you better than any single benchmark chart.
Sources
- NVIDIA — Compare GeForce graphics cards (memory size, interface width, bandwidth), accessed September 2026
- llama.cpp — quantize tool README (quantization types, sizes and speeds for Llama 3.1), accessed September 2026
- Hugging Face — "Llama 3.1 – 405B, 70B & 8B with multilinguality and long context" (inference memory requirements), accessed September 2026
- NVIDIA Technical Blog — "Mastering LLM Techniques: Inference Optimization", accessed September 2026
- Apple — Mac Studio technical specifications, accessed September 2026
- llama.cpp — project README (supported backends, CPU+GPU hybrid inference), accessed September 2026
- NVIDIA — GeForce RTX 5090 product page (specs and power requirements), accessed September 2026
Related: Running AI on your own hardware: local LLMs explained · Self-hosting 101: what's worth running on your own hardware in 2026 · How large language models actually work