GPT-OSS
gpt-oss-120b
OpenAI's open-weight MoE release, natively quantized to MXFP4 for the MoE weights. All 128 experts resident in memory; 4 active per token.
Total params
117B
Active params
5.1B
Layers
36
Hidden size
2880
Attention heads
64
KV heads (GQA)
8
Vocab size
201,088
Native context window
131,072
Native precision
BF16
Experts (total)
128
Experts active / token
4
Size this model
Opens the sizing calculator pre-filled with this model at BF16 weights / FP16 KV cache and a typical workload — sign in to run it and see GPU/cloud recommendations.
Size this model →