Best GPUs for AI and Local LLMs
Pick the model you want to run and see the cheapest set of cards that can run it, with combined VRAM, estimated speed and total power for each one. Memory pools across cards, so the answer is often two or four of something rather than one of anything.
| Setup | Buy | |||||||
|---|---|---|---|---|---|---|---|---|
| Intel Arc A750 Intel Arc · oneAPI | 8GB | 0.7GB | 77-89 tok/s | $209.99 | 225W | 21.6 | ||
| GeForce RTX 3050 8GB Nvidia · CUDA | 8GB | 0.7GB | 34-39 tok/s | $259.97 | 130W | 13.6 | ||
| Intel Arc B570 Intel Arc · oneAPI | 10GB | 2.6GB | 57-66 tok/s | $259.99 | 150W | 24.6 | ||
| Intel Arc B580 Intel Arc · oneAPI | 12GB | 4.5GB | 69-80 tok/s | $329.99 | 190W | 30.2 | ||
| Radeon RX 7600 AMD · Vulkan | 8GB | 0.7GB | 43-50 tok/s | $339.99 | 165W | 19.7 | ||
| Radeon RX 6600 XT AMD · Vulkan | 8GB | 0.7GB | 38-45 tok/s | $349.97 | 160W | 18.5 | ||
| Intel Arc A770 8GB Intel Arc · oneAPI | 8GB | 0.7GB | 77-89 tok/s | $359.00 | 225W | 22.1 | ||
| GeForce RTX 5050 Nvidia · CUDA | 8GB | 0.7GB | 48-56 tok/s | $369.99 | 130W | 23.0 | ||
| Radeon RX 9060 XT 8 GB AMD · ROCm | 8GB | 0.7GB | 48-56 tok/s | $438.82 | 150W | 36.1 | ||
| GeForce RTX 3070 Nvidia · CUDA | 8GB | 0.7GB | 67-78 tok/s | $453.87 | 220W | 31.8 | ||
| 2 × Intel Arc A380 Intel Arc · oneAPI | 12GB | 3.7GB | 28-33 tok/s | $458.00$229.00 each | 150W | 6.9 | ||
| Radeon RX 6600 AMD · Vulkan | 8GB | 0.7GB | 34-39 tok/s | $459.00 | 132W | 16.9 | ||
| 2 × GeForce RTX 3050 6GB Nvidia · CUDA | 12GB | 3.7GB | 25-29 tok/s | $459.98$229.99 each | 140W | 12.3 | ||
| GeForce RTX 5060 Nvidia · CUDA | 8GB | 0.7GB | 67-78 tok/s | $459.99 | 145W | 31.8 | ||
| GeForce RTX 3060 Ti Nvidia · CUDA | 8GB | 0.7GB | 67-78 tok/s | $464.79 | 200W | 29.1 | ||
| GeForce RTX 3060 12 GB Nvidia · CUDA | 12GB | 4.5GB | 54-63 tok/s | $479.99 | 170W | 24.4 | ||
| GeForce RTX 2080 Ti Nvidia · CUDA | 11GB | 3.5GB | 79-93 tok/s | $486.96 | 250W | 30.9 | ||
| GeForce RTX 5060 Ti 8 GB Nvidia · CUDA | 8GB | 0.7GB | 67-78 tok/s | $499.99 | 180W | 39.2 | ||
| Radeon RX 9060 XT 16 GB AMD · ROCm | 16GB | 8.2GB | 48-56 tok/s | $519.99 | 160W | 40.9 | ||
| Radeon RX 5700 XT AMD · Vulkan | 8GB | 0.7GB | 67-78 tok/s | $529.00 | 225W | 17.6 | ||
| Radeon RX 9070 GRE AMD · ROCm | 12GB | 4.5GB | 65-76 tok/s | $539.99 | 220W | 50.6 | ||
| Radeon RX 7800 XT AMD · ROCm | 16GB | 8.2GB | 80-94 tok/s | $539.99 | 263W | 48.6 | ||
| GeForce RTX 4060 Ti 8GB Nvidia · CUDA | 8GB | 0.7GB | 43-50 tok/s | $599.99 | 160W | 34.2 | ||
| GeForce RTX 4060 Nvidia · CUDA | 8GB | 0.7GB | 41-48 tok/s | $609.99 | 115W | 27.5 | ||
| Radeon RX 6700 XT AMD · Vulkan | 12GB | 4.5GB | 58-67 tok/s | $680.00 | 230W | 32.6 | ||
| Radeon RX 9070 AMD · ROCm | 16GB | 8.2GB | 82-96 tok/s | $699.99 | 220W | 61.1 | ||
| GeForce RTX 5060 Ti 16GB Nvidia · CUDA | 16GB | 8.2GB | 67-78 tok/s | $788.99 | 180W | 43.1 | ||
| GeForce RTX 3080 12GB Nvidia · CUDA | 12GB | 4.5GB | 117-137 tok/s | $789.00 | 350W | 48.0 | ||
| Radeon RX 9070 XT AMD · ROCm | 16GB | 8.2GB | 82-96 tok/s | $790.23 | 304W | 68.7 | ||
| GeForce RTX 4070 Nvidia · CUDA | 12GB | 4.5GB | 76-88 tok/s | $799.00 | 200W | 47.9 | ||
| Radeon RX 7600 XT AMD · Vulkan | 16GB | 8.2GB | 43-50 tok/s | $799.00 | 190W | 29.4 | ||
| GeForce RTX 5070 Nvidia · CUDA | 12GB | 4.5GB | 86-101 tok/s | $849.00 | 250W | 56.8 | ||
| Radeon RX 7900 XT AMD · ROCm | 20GB | 12.0GB | 102-120 tok/s | $849.99 | 300W | 63.8 | ||
| 3 × Radeon RX 6500 XT AMD · Vulkan | 12GB | 2.9GB | 22-25 tok/s | $869.97$289.99 each | 321W | 6.3 | ||
| Radeon RX 6900 XT AMD · Vulkan | 16GB | 8.2GB | 77-89 tok/s | $889.00 | 300W | 48.3 | ||
| Radeon RX 6800 XT AMD · Vulkan | 16GB | 8.2GB | 77-89 tok/s | $890.00 | 300W | 48.1 | ||
| GeForce RTX 3080 Ti Nvidia · CUDA | 12GB | 4.5GB | 117-137 tok/s | $909.99 | 350W | 48.2 | ||
| 3 × Radeon RX 6400 AMD · Vulkan | 12GB | 2.9GB | 19-22 tok/s | $910.41$303.47 each | 159W | — | ||
| GeForce RTX 4070 Super Nvidia · CUDA | 12GB | 4.5GB | 76-88 tok/s | $969.00 | 220W | 55.8 | ||
| GeForce RTX 4070 Ti Nvidia · CUDA | 12GB | 4.5GB | 76-88 tok/s | $1,149.00 | 285W | 59.8 | ||
| GeForce RTX 5070 Ti Nvidia · CUDA | 16GB | 8.2GB | 115-135 tok/s | $1,169.99 | 300W | 71.0 | ||
| Titan RTX Nvidia · CUDA | 24GB | 15.8GB | 86-101 tok/s | $1,395.00 | 280W | 33.5 | ||
| GeForce RTX 5080 Nvidia · CUDA | 16GB | 8.2GB | 123-144 tok/s | $1,599.99 | 360W | 78.0 | ||
| GeForce RTX 4080 Super Nvidia · CUDA | 16GB | 8.2GB | 94-111 tok/s | $1,649.00 | 320W | 73.5 | ||
| GeForce RTX 4080 Nvidia · CUDA | 16GB | 8.2GB | 92-108 tok/s | $1,745.00 | 320W | 68.7 | ||
| GeForce RTX 3090 Nvidia · CUDA | 24GB | 15.8GB | 120-141 tok/s | $1,969.99 | 350W | 48.4 | ||
| GeForce RTX 3090 Ti Nvidia · CUDA | 24GB | 15.8GB | 129-152 tok/s | $2,395.00 | 450W | 51.1 | ||
| GeForce RTX 4090 Nvidia · CUDA | 24GB | 15.8GB | 129-152 tok/s | $4,499.99 | 450W | 88.2 | ||
| GeForce RTX 5090 Nvidia · CUDA | 32GB | 23.3GB | 200-240 tok/s | $7,100.00 | 575W | 100.0 |
VRAM decides what you can run, bandwidth decides how fast
Running a language model on your own machine is mostly a memory problem. The model’s weights have to fit in the card’s VRAM, along with its attention cache, which grows with the length of the conversation. If the model does not fit, it spills into system memory and slows down by a factor of ten or more, so capacity is the first question and everything else is secondary. The gaming answer is different and much lower: our guide to how much VRAM you need covers the 8GB, 12GB and 16GB tiers that matter for frame rates.
Speed is set by memory bandwidth rather than raw compute, because producing each word reads the whole active model from memory once. That is why a card with modest gaming performance and fast memory can generate text quicker than a faster-on-paper card with slower memory, and why the speed figures here are derived from bandwidth instead of from frame rates.
What quantisation changes
Quantisation compresses the weights so a larger model fits a smaller card. Q4 is the usual choice: it cuts memory to roughly a third of the original with a quality cost most people do not notice in ordinary use. Q8 sits closer to the original at about half, and FP16 is the uncompressed form, which is mostly of interest for training and research rather than for running a model day to day.
A “4-bit” build is not 4 bits per weight. Q4_K_M mixes 4-bit and 6-bit tensors, carries scaling factors for each block, and leaves some tensors at full precision, which works out near 4.9 bits in practice. The sizes here are measured from published builds rather than calculated from the bit width, so the figures match what you would actually download.
Does the file format matter? Less than the precision does. GGUF, GPTQ, AWQ and EXL2 differ in which software runs them and in small details of how the weights are packed, but at the same precision they land within a few percent of each other on memory, so the choice of precision above is what moves the answer. Two practical differences are worth knowing: GGUF can spill part of a model into system memory when it does not quite fit, at a heavy cost in speed, while GPTQ and AWQ expect the whole thing to sit in VRAM; and EXL2 lets you pick a precision anywhere on the scale rather than at fixed steps, which is a finer version of the same slider.
Why context length matters
The attention cache holds the conversation so far, and it grows with every word. Its size depends on the model’s own architecture, so the figures here are worked out per model rather than from a rule of thumb. On a 24GB card at Q4 the largest model you can hold falls from around 36B at a short context to around 8B at a very long one, which is why a capacity figure quoted without its context length is close to meaningless.
Several cards, and what that really buys you
Memory pools across graphics cards for this kind of work. The model is split by layer, each card holds its share, and no special link between them is required, which is why two 24GB cards will run a model that no single consumer card can hold. That is what makes the cheapest answer to a large model usually a few mid-range cards rather than one expensive one.
What more cards do not buy is speed. For one conversation at a time the cards take turns, and each word still reads the whole model once, so throughput tracks a single card’s memory bandwidth however many you add. Four cards give you four times the capacity and roughly the same tokens per second, which is the most commonly misunderstood thing about these setups.
Splitting a model does cost something, and the figures here account for the part that can be pinned down. Each card keeps its own working memory, so four 16GB cards hold a little less than the 64GB on the box, and the capacity column is calculated that way rather than by simple addition. On speed, the estimate assumes a multi-card setup runs no faster than a single card would, which matches how the default splitting works. Handing the conversation from one card to the next does add a small delay on top, too small to measure reliably from published figures, so treat the speed estimate as the optimistic end for a multi-card build. There is a second mode that divides each layer across cards instead of giving each card its own layers, which can be meaningfully quicker on two matched cards, but it asks more of the connection between them and is not the default.
How many cards you can fit depends on the board. An ordinary desktop board has one full-length slot and sometimes a second, so two is the realistic limit before the slots themselves run out. Workstation boards for Threadripper and Xeon processors commonly carry four to seven, which is why serious local-AI machines tend to be built on them. Because splitting a model by layer sends very little data between cards, a narrow connection costs less here than it would in a game, and builders do go further using riser cables. Physical space runs out before slots do on most builds, so check the GPU size and case clearance guide before committing to more than two cards.
Power is the constraint people meet last and should check first. A single 15A household circuit in the US carries about 1,440W continuously, and four high-end cards exceed that on their own before the rest of the machine is counted. The combined wattage column turns amber past that point, because at that stage it has stopped being a question about the power supply and become a question about the wiring, and beyond roughly eight cards it stops being a household question at all. If you are sizing a power supply for one of these builds, the guide to TDP and PSU sizing has the arithmetic, and the GPU efficiency rankings show which cards deliver the most performance per watt.
Nvidia, AMD and Intel are not interchangeable here
The memory and bandwidth figures in this table apply to any card, but the software around them does not. Nvidia’s CUDA is what almost every tool is written against first, so an Nvidia card is the one that works with the least thought, and it is effectively required for training, fine-tuning or serving many users at once.
For simply running a model on your own machine, which is what most people want, AMD and Intel are genuinely usable. The popular desktop applications support them, through AMD’s ROCm and Intel’s oneAPI, and through a portable Vulkan path that runs on almost anything. AMD officially supports ROCm on its recent cards and its professional line rather than across the whole back catalogue, so an older Radeon tends to rely on that portable path, which works but gives up some speed.
So read a cheap Intel or AMD card near the top of this table for what it is: the least expensive way to reach that much memory, with a little more setup and a narrower choice of software than the Nvidia equivalent. If you intend to train rather than run, or you want the broadest software support, that is worth paying the Nvidia premium for. For gaming value rather than AI capacity, the GPU performance per dollar rankings rank the same cards on frames per dollar instead.
How to read the estimates
Speed is shown as a range because the relationship between bandwidth and tokens per second is close but not exact: very fast cards running small models lose some of their advantage to overheads that have nothing to do with memory. The ranges here were checked against published measurements across eight cards spanning entry-level to workstation, and they sit within about a fifth either way.
Mixture-of-experts models, where only part of the model runs for each word, show no speed figure at all. A bandwidth estimate is not dependable for them, and a number we cannot stand behind is worse than a blank. Capacity for those models is still exact, because all of their weights occupy memory even though only a slice is read per word.
All Graphics Cards (74)
Every GPU we track, grouped by brand and generation. Open any model for its VRAM, memory bandwidth and what it can run locally.
25 models are out of stock at every retailer we track right now. Specs and benchmarks still apply, and they rejoin the table above automatically when stock returns.
Nvidia
GeForce RTX 50 (8)
GeForce RTX 40 (10)
GeForce RTX 30 (12)
- GeForce RTX 3050 6GB
- GeForce RTX 3050 8GB
- GeForce RTX 3060 12 GB
- GeForce RTX 3060 8 GB (currently out of stock)
- GeForce RTX 3060 Ti
- GeForce RTX 3070
- GeForce RTX 3070 Ti (currently out of stock)
- GeForce RTX 3080 (currently out of stock)
- GeForce RTX 3080 12GB
- GeForce RTX 3080 Ti
- GeForce RTX 3090
- GeForce RTX 3090 Ti
GeForce RTX 20 (5)
GeForce GTX 16 (1)
GeForce 10 (4)
AMD
Radeon RX 9000 (5)
Radeon RX 7000 (7)
Radeon RX 6000 (8)
Radeon RX 5000 (2)
Radeon R9 (1)
Affiliate disclosure: Some links on this page are affiliate links — MaxMyBuild may earn a commission at no extra cost to you if you buy through them. As an Amazon Associate, MaxMyBuild earns from qualifying purchases. For Amazon, current price and availability are shown on the Amazon product page; product prices and availability are accurate as of the date and time of purchase and are subject to change. Prices shown for other retailers are approximate and may differ at checkout. “Search” links are non-affiliate and provided for convenience only.
