What OpenAI actually released
The memory figure, published by OpenAI
into a single 80GB GPU,
naming the NVIDIA H100 and AMD MI300X, and that gpt-oss-20b runs within 16GB of memory.
Those are their numbers, not ours, and they are only possible because the weights are already
MXFP4. Serve the same models at a higher precision and you are back to doing the sum yourself.
| Model | Parameters | Active per token | Memory stated by OpenAI | Download size |
|---|---|---|---|---|
| gpt-oss-20b | 21B | 3.6B | within 16 GB | 14 GB |
| gpt-oss-120b | 117B | 5.1B | a single 80 GB GPU | 65 GB |
Both models carry a 128K context. The download sizes are the ones Ollama lists in its own library, and a download size is not a memory requirement — the context and the runtime sit on top of it. On performance we report only what OpenAI claims about its own models: that gpt-oss-120b reaches near-parity with o4-mini on core reasoning benchmarks and that gpt-oss-20b performs similarly to o3-mini on common ones. DohoHub does not run benchmarks, so those figures are theirs and are labelled as such.
Cheapest GPU that clears the bar
All GPU plans →The cheapest live GPU plan in the DohoHub catalogue at each memory tier, out of 142 plans whose provider publishes a GPU memory figure. The two tiers that matter here are the ones OpenAI names: 16 GB for the 20b and 80 GB for the 120b. Prices come from our own feed, normalised so they are comparable and shown in the currency picked in the header.
| Provider | Plan | Specs | Price | Visit |
|---|---|---|---|---|
|
Bee GPU VPS
cheapest at 16–24 GB+ · 24 GB on board
|
6 Cores24 GB vRAM30 GB RAM400 GB NVME100 TB traffic10 Gbps |
$79.00/mo
|
Visit | |
|
Supermicro X11 10SFF (GPU)
cheapest at 48 GB+ · 64 GB on board
|
64 GB vRAM25 TB traffic |
$355.99/mo
renews at $402.52
|
Visit | |
|
Enterprise GPU VPS - RTX Pro 6000
cheapest at 80–96 GB+ · 96 GB on board
|
32 Cores96 GB vRAM84 GB RAM400 GB SSDUnlimited traffic1 Gbps |
$649.00/mo
|
Visit |
“Within 16 GB” means the weights and a working context, not the weights and a crowd — if several people will use it at once, take the next tier up and read the vLLM page on why. We record the memory a plan advertises, not which card supplies it, and 80 GB spread over two cards is not the single 80 GB GPU OpenAI describes. None of these providers sells “GPT-OSS hosting”; they sell GPU servers, and the weights are a download away.
Filter the catalogue by GPU memory, cores, storage and location — live prices, tracked every six hours.
Find a GPU server →