What you get with the weights
How much GPU memory each size needs
0.5 byte per parameter
at 4-bit, plus +25% for the KV cache and the serving runtime. It is a floor.
A 128K prompt, a big batch or a different engine all push it up, and a mixture-of-experts
model needs the whole parameter count in memory even though only a fraction of it
does work on any given token.
| Model | Parameters | Weights at 4-bit | With headroom | Fits |
|---|---|---|---|---|
| R1-Distill-Qwen 1.5B | 1.5B | ≈ 0.8 GB | ≈ 1 GB | Anything with a GPU |
| R1-Distill-Qwen 7B | 7B | ≈ 3.5 GB | ≈ 4 GB | An entry card |
| R1-Distill-Llama 8B | 8B | ≈ 4 GB | ≈ 5 GB | An entry card |
| R1-Distill-Qwen 14B | 14B | ≈ 7 GB | ≈ 9 GB | One 12–16 GB card |
| R1-Distill-Qwen 32B | 32B | ≈ 16 GB | ≈ 20 GB | One 24 GB card |
| R1-Distill-Llama 70B | 70B | ≈ 35 GB | ≈ 44 GB | Two 24 GB cards, or one 48 GB |
| DeepSeek-R1 (full) | 671B | ≈ 336 GB | ≈ 420 GB | A multi-GPU node |
The “fits” column is a rule of thumb about memory only. It says nothing about how fast the answer comes back — that is the card’s bandwidth and generation, which our catalogue records unevenly, so we do not pretend to rank it here.
Cheapest GPU at each tier
All GPU plans →The cheapest live GPU plan in the DohoHub catalogue at each memory tier, out of 142 plans whose provider publishes a GPU memory figure. Prices are the ones our feed tracks, normalised to be comparable and shown in the currency picked in the header. Where one machine is the cheapest way to clear several tiers, it is listed once and the note says which tiers it covers.
| Provider | Plan | Specs | Price | Visit |
|---|---|---|---|---|
|
Bee GPU VPS
cheapest at 8–24 GB+ · 24 GB on board
|
6 Cores24 GB vRAM30 GB RAM400 GB NVME100 TB traffic10 Gbps |
$79.00/mo
|
Visit | |
|
Supermicro X11 10SFF (GPU)
cheapest at 48 GB+ · 64 GB on board
|
64 GB vRAM25 TB traffic |
$355.99/mo
renews at $402.52
|
Visit | |
|
Enterprise GPU VPS - RTX Pro 6000
cheapest at 96 GB+ · 96 GB on board
|
32 Cores96 GB vRAM84 GB RAM400 GB SSDUnlimited traffic1 Gbps |
$649.00/mo
|
Visit | |
|
Supermicro H12 4SFF+U.3 (8 GPU)
cheapest at 192 GB+ · 384 GB on board
|
384 GB vRAM25 TB traffic50 Gbps |
$3,138.73/mo
|
Visit |
We record the memory a plan advertises, not which card supplies it, and 48 GB spread over two older cards behaves differently from 48 GB on one modern card. The full 671B model wants roughly 420 GB even at 4-bit — more than any single listing here holds, so that one is a multi-node question these plans do not answer. None of these providers sells “DeepSeek hosting”; they sell GPU servers, and the install is yours.
Filter the catalogue by GPU memory, cores, storage and location — live prices, tracked every six hours.
Find a GPU server →