Nemotron 3 Super 120B-A12B: local hardware requirements and GPU compatibility
← All models

Local AI model

Nemotron 3 Super 120B-A12B

by NVIDIA

The weights fit entirely in accelerator memory on 0 of 22 reference machines. The KV cache or runtime state is not fully modelled, so this is a memory floor, not a "runs" verdict.

124B Parameters 256k Maximum context 1 Quantization Partial Compatibility coverage
Check my PC

Overview

Publisher
NVIDIA
Family
Nemotron
Series
Nemotron 3
Task
Text generation
Architecture
mixture of experts · nemotron_h
Parameters
124B
Maximum context
262,144 tokens
Licence
nvidia-nemotron-open-model-license
Published
2026-03-10
Formats
gguf, safetensors
KV cache per token (f16)
unknown
Compatibility coverage
Partial: memory floor only
Upstream repository
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
Official page
https://www.nvidia.com/en-us/ai-data-science/foundation-models/

Quantizations and file sizes

Quantization Runtime Weights Memory floor at 4k Identity
Q4_K_Mgguf ollamaollama run nemotron-3-super:120b 80.9 GB 80.9 GB digest verified

Sizes come from the runtime registry or repository listing. The memory floor adds the KV cache for 4,096 tokens when that figure cannot over-count; it is an estimate, not a requirement.

Runtimes

Compatibility on reference machines

8k context, 32 GB system RAM, best catalogued quantization per machine.

Machine Memory Answer Known facts
GeForce RTX 3060 Laptop GPU 6GB 6.00 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 6.00 GB VRAM: do not fitKV cache at 8k: unknown
Radeon RX 7600 8GB 8.00 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 8.00 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 4060 8GB 8.00 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 8.00 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 3080 10GB 10.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 10.0 GB VRAM: do not fitKV cache at 8k: unknown
Intel Arc B580 12GB 12.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: unknownruntime support on this machine: unknown
GeForce RTX 3060 12GB 12.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 4070 12GB 12.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 5070 12GB 12.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: unknown
Radeon RX 7800 XT 16GB 16.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: unknown
Radeon RX 9070 XT 16GB 16.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: unknown
Intel Arc A770 16GB 16.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: unknownruntime support on this machine: unknown
GeForce RTX 4060 Ti 16GB 16.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 4080 16GB 16.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 5060 Ti 16GB 16.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 5080 16GB 16.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: unknown
Radeon RX 7900 XTX 24GB 24.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 24.0 GB VRAM: do not fitKV cache at 8k: unknown
Mac mini M4 24GB 24.0 GB unified Does not fit estimated · Q4_K_M Weights 80.9 GB vs 24.0 GB unified: do not fitKV cache at 8k: unknown
Mac mini M4 Pro 24GB 24.0 GB unified Does not fit estimated · Q4_K_M Weights 80.9 GB vs 24.0 GB unified: do not fitKV cache at 8k: unknown
GeForce RTX 3090 24GB 24.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 24.0 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 4090 24GB 24.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 24.0 GB VRAM: do not fitKV cache at 8k: unknown
GeForce RTX 5090 32GB 32.0 GB VRAM Does not fit estimated · Q4_K_M Weights 80.9 GB vs 32.0 GB VRAM: do not fitKV cache at 8k: unknown
MacBook Pro 14-inch M4 Max 36GB 36.0 GB unified Does not fit estimated · Q4_K_M Weights 80.9 GB vs 36.0 GB unified: do not fitKV cache at 8k: unknown

Definitive answers ("runs", "does not fit") come from the calibrated engine or a physical run. "Potential" rows only compare the weights with memory: the rest of the requirement is not modelled, so they are not verdicts. Unknown is never a failure.

Hardware for this model

The smallest artifact (80.9 GB) exceeds every catalogued consumer GPU. No cloud GPU provider is configured, so no offer is shown.

  • Smallest catalogued artifact: 80.9 GB of weights (Q4_K_M).
  • Larger than every catalogued consumer GPU (32 GB). Locally it needs CPU offload with 128 GB of RAM or more; otherwise a rented datacenter GPU.
  • Disk space: 500 GB or more to keep this model and its variants.
  • Not modelled yet: unknown.

Cloud GPU — LocalAIReady has no cloud GPU partner, so no provider or price is shown here. Holding these weights in one accelerator needs at least 81 GB of memory, which today means a rented datacenter GPU.

These links open a hardware category search, not a specific product recommendation. Affiliate link — I may earn a commission from qualifying purchases. These links never change a compatibility answer. They follow the memory the data defends, not commission.

Sources and evidence

0 measured, 26 sourced, 2 estimated and 4 unknown fields.

Catalogue record verified 2026-09-17.

What else can your PC run?

Pick your GPU and system RAM to see every catalogued model that fits, with tested, estimated and potential answers kept apart.

Check my PC