GPT-OSS
How much VRAM does gpt-oss-120b need?
gpt-oss-120b has 117B parameters, so its weights alone take about 235 GB of VRAM at FP16 — or roughly 59 GB quantized to INT4. Real serving adds KV cache on top, which scales with your context length and concurrent requests. Size the exact figure for your workload below.
OpenAI's open-weight MoE release, natively quantized to MXFP4 for the MoE weights. All 128 experts resident in memory; 4 active per token.
Total params
117B
Active params
5.1B
Layers
36
Hidden size
2880
Attention heads
64
KV heads (GQA)
8
Vocab size
201,088
Native context window
131,072
Native precision
BF16
Experts (total)
128
Experts active / token
4
Size gpt-oss-120b for your workload
Opens the free calculator with gpt-oss-120b loaded at BF16 weights and FP16 KV cache. Set your context length and traffic to get exact GPU memory and a ranked list of cloud instances that fit — no signup.
Size this model →