Best GPUs for AI and Local LLMs

Pick the model you want to run and see the cheapest set of cards that can run it, with combined VRAM, estimated speed and total power for each one. Memory pools across cards, so the answer is often two or four of something rather than one of anything.

Intel Arc A750
Intel Arc · oneAPI
$209.99
Combined VRAM
8GB
Est. speed
77-89 tok/s
Combined watts
225W
Buy
GeForce RTX 3050 8GB
Nvidia · CUDA
$259.97
Combined VRAM
8GB
Est. speed
34-39 tok/s
Combined watts
130W
Buy
Intel Arc B570
Intel Arc · oneAPI
$259.99
Combined VRAM
10GB
Est. speed
57-66 tok/s
Combined watts
150W
Buy
Intel Arc B580
Intel Arc · oneAPI
$329.99
Combined VRAM
12GB
Est. speed
69-80 tok/s
Combined watts
190W
Buy
Radeon RX 7600
AMD · Vulkan
$339.99
Combined VRAM
8GB
Est. speed
43-50 tok/s
Combined watts
165W
Buy
Radeon RX 6600 XT
AMD · Vulkan
$349.97
Combined VRAM
8GB
Est. speed
38-45 tok/s
Combined watts
160W
Buy
Intel Arc A770 8GB
Intel Arc · oneAPI
$359.00
Combined VRAM
8GB
Est. speed
77-89 tok/s
Combined watts
225W
Buy
GeForce RTX 5050
Nvidia · CUDA
$369.99
Combined VRAM
8GB
Est. speed
48-56 tok/s
Combined watts
130W
Buy
$438.82
Combined VRAM
8GB
Est. speed
48-56 tok/s
Combined watts
150W
Buy
GeForce RTX 3070
Nvidia · CUDA
$453.87
Combined VRAM
8GB
Est. speed
67-78 tok/s
Combined watts
220W
Buy
2 × Intel Arc A380
Intel Arc · oneAPI
$458.00
$229.00 each
Combined VRAM
12GB
Est. speed
28-33 tok/s
Combined watts
150W
Buy
Radeon RX 6600
AMD · Vulkan
$459.00
Combined VRAM
8GB
Est. speed
34-39 tok/s
Combined watts
132W
Buy
$459.98
$229.99 each
Combined VRAM
12GB
Est. speed
25-29 tok/s
Combined watts
140W
Buy
GeForce RTX 5060
Nvidia · CUDA
$459.99
Combined VRAM
8GB
Est. speed
67-78 tok/s
Combined watts
145W
Buy
GeForce RTX 3060 Ti
Nvidia · CUDA
$464.79
Combined VRAM
8GB
Est. speed
67-78 tok/s
Combined watts
200W
Buy
$479.99
Combined VRAM
12GB
Est. speed
54-63 tok/s
Combined watts
170W
Buy
GeForce RTX 2080 Ti
Nvidia · CUDA
$486.96
Combined VRAM
11GB
Est. speed
79-93 tok/s
Combined watts
250W
Buy
$499.99
Combined VRAM
8GB
Est. speed
67-78 tok/s
Combined watts
180W
Buy
$519.99
Combined VRAM
16GB
Est. speed
48-56 tok/s
Combined watts
160W
Buy
Radeon RX 5700 XT
AMD · Vulkan
$529.00
Combined VRAM
8GB
Est. speed
67-78 tok/s
Combined watts
225W
Buy
$539.99
Combined VRAM
12GB
Est. speed
65-76 tok/s
Combined watts
220W
Buy
$539.99
Combined VRAM
16GB
Est. speed
80-94 tok/s
Combined watts
263W
Buy
$599.99
Combined VRAM
8GB
Est. speed
43-50 tok/s
Combined watts
160W
Buy
GeForce RTX 4060
Nvidia · CUDA
$609.99
Combined VRAM
8GB
Est. speed
41-48 tok/s
Combined watts
115W
Buy
Radeon RX 6700 XT
AMD · Vulkan
$680.00
Combined VRAM
12GB
Est. speed
58-67 tok/s
Combined watts
230W
Buy
Radeon RX 9070
AMD · ROCm
$699.99
Combined VRAM
16GB
Est. speed
82-96 tok/s
Combined watts
220W
Buy
$788.99
Combined VRAM
16GB
Est. speed
67-78 tok/s
Combined watts
180W
Buy
GeForce RTX 3080 12GB
Nvidia · CUDA
$789.00
Combined VRAM
12GB
Est. speed
117-137 tok/s
Combined watts
350W
Buy
$790.23
Combined VRAM
16GB
Est. speed
82-96 tok/s
Combined watts
304W
Buy
GeForce RTX 4070
Nvidia · CUDA
$799.00
Combined VRAM
12GB
Est. speed
76-88 tok/s
Combined watts
200W
Buy
Radeon RX 7600 XT
AMD · Vulkan
$799.00
Combined VRAM
16GB
Est. speed
43-50 tok/s
Combined watts
190W
Buy
GeForce RTX 5070
Nvidia · CUDA
$849.00
Combined VRAM
12GB
Est. speed
86-101 tok/s
Combined watts
250W
Buy
$849.99
Combined VRAM
20GB
Est. speed
102-120 tok/s
Combined watts
300W
Buy
$869.97
$289.99 each
Combined VRAM
12GB
Est. speed
22-25 tok/s
Combined watts
321W
Buy
Radeon RX 6900 XT
AMD · Vulkan
$889.00
Combined VRAM
16GB
Est. speed
77-89 tok/s
Combined watts
300W
Buy
Radeon RX 6800 XT
AMD · Vulkan
$890.00
Combined VRAM
16GB
Est. speed
77-89 tok/s
Combined watts
300W
Buy
GeForce RTX 3080 Ti
Nvidia · CUDA
$909.99
Combined VRAM
12GB
Est. speed
117-137 tok/s
Combined watts
350W
Buy
3 × Radeon RX 6400
AMD · Vulkan
$910.41
$303.47 each
Combined VRAM
12GB
Est. speed
19-22 tok/s
Combined watts
159W
Buy
$969.00
Combined VRAM
12GB
Est. speed
76-88 tok/s
Combined watts
220W
Buy
GeForce RTX 4070 Ti
Nvidia · CUDA
$1,149.00
Combined VRAM
12GB
Est. speed
76-88 tok/s
Combined watts
285W
Buy
GeForce RTX 5070 Ti
Nvidia · CUDA
$1,169.99
Combined VRAM
16GB
Est. speed
115-135 tok/s
Combined watts
300W
Buy
Titan RTX
Nvidia · CUDA
$1,395.00
Combined VRAM
24GB
Est. speed
86-101 tok/s
Combined watts
280W
Buy
GeForce RTX 5080
Nvidia · CUDA
$1,599.99
Combined VRAM
16GB
Est. speed
123-144 tok/s
Combined watts
360W
Buy
$1,649.00
Combined VRAM
16GB
Est. speed
94-111 tok/s
Combined watts
320W
Buy
GeForce RTX 4080
Nvidia · CUDA
$1,745.00
Combined VRAM
16GB
Est. speed
92-108 tok/s
Combined watts
320W
Buy
GeForce RTX 3090
Nvidia · CUDA
$1,969.99
Combined VRAM
24GB
Est. speed
120-141 tok/s
Combined watts
350W
Buy
GeForce RTX 3090 Ti
Nvidia · CUDA
$2,395.00
Combined VRAM
24GB
Est. speed
129-152 tok/s
Combined watts
450W
Buy
GeForce RTX 4090
Nvidia · CUDA
$4,499.99
Combined VRAM
24GB
Est. speed
129-152 tok/s
Combined watts
450W
Buy
GeForce RTX 5090
Nvidia · CUDA
$7,100.00
Combined VRAM
32GB
Est. speed
200-240 tok/s
Combined watts
575W
Buy

VRAM decides what you can run, bandwidth decides how fast

Running a language model on your own machine is mostly a memory problem. The model’s weights have to fit in the card’s VRAM, along with its attention cache, which grows with the length of the conversation. If the model does not fit, it spills into system memory and slows down by a factor of ten or more, so capacity is the first question and everything else is secondary. The gaming answer is different and much lower: our guide to how much VRAM you need covers the 8GB, 12GB and 16GB tiers that matter for frame rates.

Speed is set by memory bandwidth rather than raw compute, because producing each word reads the whole active model from memory once. That is why a card with modest gaming performance and fast memory can generate text quicker than a faster-on-paper card with slower memory, and why the speed figures here are derived from bandwidth instead of from frame rates.

What quantisation changes

Quantisation compresses the weights so a larger model fits a smaller card. Q4 is the usual choice: it cuts memory to roughly a third of the original with a quality cost most people do not notice in ordinary use. Q8 sits closer to the original at about half, and FP16 is the uncompressed form, which is mostly of interest for training and research rather than for running a model day to day.

A “4-bit” build is not 4 bits per weight. Q4_K_M mixes 4-bit and 6-bit tensors, carries scaling factors for each block, and leaves some tensors at full precision, which works out near 4.9 bits in practice. The sizes here are measured from published builds rather than calculated from the bit width, so the figures match what you would actually download.

Does the file format matter? Less than the precision does. GGUF, GPTQ, AWQ and EXL2 differ in which software runs them and in small details of how the weights are packed, but at the same precision they land within a few percent of each other on memory, so the choice of precision above is what moves the answer. Two practical differences are worth knowing: GGUF can spill part of a model into system memory when it does not quite fit, at a heavy cost in speed, while GPTQ and AWQ expect the whole thing to sit in VRAM; and EXL2 lets you pick a precision anywhere on the scale rather than at fixed steps, which is a finer version of the same slider.

Why context length matters

The attention cache holds the conversation so far, and it grows with every word. Its size depends on the model’s own architecture, so the figures here are worked out per model rather than from a rule of thumb. On a 24GB card at Q4 the largest model you can hold falls from around 36B at a short context to around 8B at a very long one, which is why a capacity figure quoted without its context length is close to meaningless.

Several cards, and what that really buys you

Memory pools across graphics cards for this kind of work. The model is split by layer, each card holds its share, and no special link between them is required, which is why two 24GB cards will run a model that no single consumer card can hold. That is what makes the cheapest answer to a large model usually a few mid-range cards rather than one expensive one.

What more cards do not buy is speed. For one conversation at a time the cards take turns, and each word still reads the whole model once, so throughput tracks a single card’s memory bandwidth however many you add. Four cards give you four times the capacity and roughly the same tokens per second, which is the most commonly misunderstood thing about these setups.

Splitting a model does cost something, and the figures here account for the part that can be pinned down. Each card keeps its own working memory, so four 16GB cards hold a little less than the 64GB on the box, and the capacity column is calculated that way rather than by simple addition. On speed, the estimate assumes a multi-card setup runs no faster than a single card would, which matches how the default splitting works. Handing the conversation from one card to the next does add a small delay on top, too small to measure reliably from published figures, so treat the speed estimate as the optimistic end for a multi-card build. There is a second mode that divides each layer across cards instead of giving each card its own layers, which can be meaningfully quicker on two matched cards, but it asks more of the connection between them and is not the default.

How many cards you can fit depends on the board. An ordinary desktop board has one full-length slot and sometimes a second, so two is the realistic limit before the slots themselves run out. Workstation boards for Threadripper and Xeon processors commonly carry four to seven, which is why serious local-AI machines tend to be built on them. Because splitting a model by layer sends very little data between cards, a narrow connection costs less here than it would in a game, and builders do go further using riser cables. Physical space runs out before slots do on most builds, so check the GPU size and case clearance guide before committing to more than two cards.

Power is the constraint people meet last and should check first. A single 15A household circuit in the US carries about 1,440W continuously, and four high-end cards exceed that on their own before the rest of the machine is counted. The combined wattage column turns amber past that point, because at that stage it has stopped being a question about the power supply and become a question about the wiring, and beyond roughly eight cards it stops being a household question at all. If you are sizing a power supply for one of these builds, the guide to TDP and PSU sizing has the arithmetic, and the GPU efficiency rankings show which cards deliver the most performance per watt.

Nvidia, AMD and Intel are not interchangeable here

The memory and bandwidth figures in this table apply to any card, but the software around them does not. Nvidia’s CUDA is what almost every tool is written against first, so an Nvidia card is the one that works with the least thought, and it is effectively required for training, fine-tuning or serving many users at once.

For simply running a model on your own machine, which is what most people want, AMD and Intel are genuinely usable. The popular desktop applications support them, through AMD’s ROCm and Intel’s oneAPI, and through a portable Vulkan path that runs on almost anything. AMD officially supports ROCm on its recent cards and its professional line rather than across the whole back catalogue, so an older Radeon tends to rely on that portable path, which works but gives up some speed.

So read a cheap Intel or AMD card near the top of this table for what it is: the least expensive way to reach that much memory, with a little more setup and a narrower choice of software than the Nvidia equivalent. If you intend to train rather than run, or you want the broadest software support, that is worth paying the Nvidia premium for. For gaming value rather than AI capacity, the GPU performance per dollar rankings rank the same cards on frames per dollar instead.

How to read the estimates

Speed is shown as a range because the relationship between bandwidth and tokens per second is close but not exact: very fast cards running small models lose some of their advantage to overheads that have nothing to do with memory. The ranges here were checked against published measurements across eight cards spanning entry-level to workstation, and they sit within about a fifth either way.

Mixture-of-experts models, where only part of the model runs for each word, show no speed figure at all. A bandwidth estimate is not dependable for them, and a number we cannot stand behind is worse than a blank. Capacity for those models is still exact, because all of their weights occupy memory even though only a slice is read per word.

All Graphics Cards (74)

Every GPU we track, grouped by brand and generation. Open any model for its VRAM, memory bandwidth and what it can run locally.

25 models are out of stock at every retailer we track right now. Specs and benchmarks still apply, and they rejoin the table above automatically when stock returns.

Nvidia

Affiliate disclosure: Some links on this page are affiliate links — MaxMyBuild may earn a commission at no extra cost to you if you buy through them. As an Amazon Associate, MaxMyBuild earns from qualifying purchases. For Amazon, current price and availability are shown on the Amazon product page; product prices and availability are accurate as of the date and time of purchase and are subject to change. Prices shown for other retailers are approximate and may differ at checkout. “Search” links are non-affiliate and provided for convenience only.

Advertisement