Granite 4.0 H Small: local hardware requirements and GPU compatibility
← All models

Local AI model

Granite 4.0 H Small

by IBM

The weights fit entirely in accelerator memory on 4 of 22 reference machines. The KV cache or runtime state is not fully modelled, so this is a memory floor, not a "runs" verdict.

32.2B Parameters 128k Maximum context 1 Quantization Partial Compatibility coverage
Check my PC

Overview

Publisher
IBM
Family
Granite
Series
Granite 4.0
Task
Text generation
Architecture
mixture of experts · granitemoehybrid
Parameters
32.2B
Maximum context
131,072 tokens
Licence
apache-2.0
Published
2025-09-16
Formats
gguf, safetensors
KV cache per token (f16)
16,384 B · attention layers only: recurrent state not included
Compatibility coverage
Partial: memory floor only
Upstream repository
ibm-granite/granite-4.0-h-small
Official page
https://www.ibm.com/granite

Quantizations and file sizes

Quantization Runtime Weights Memory floor at 4k Identity
Q4_K_Mgguf ollamaollama run granite4:small-h 18.1 GB 18.2 GB digest verified

Sizes come from the runtime registry or repository listing. The memory floor adds the KV cache for 4,096 tokens when that figure cannot over-count; it is an estimate, not a requirement.

Runtimes

Compatibility on reference machines

8k context, 32 GB system RAM, best catalogued quantization per machine.

Machine Memory Answer Known facts
Radeon RX 7900 XTX 24GB 24.0 GB VRAM Potential: weights fit not verified · Q4_K_M Weights 18.1 GB vs 24.0 GB VRAM: fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 3090 24GB 24.0 GB VRAM Potential: weights fit not verified · Q4_K_M Weights 18.1 GB vs 24.0 GB VRAM: fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 4090 24GB 24.0 GB VRAM Potential: weights fit not verified · Q4_K_M Weights 18.1 GB vs 24.0 GB VRAM: fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 5090 32GB 32.0 GB VRAM Potential: weights fit not verified · Q4_K_M Weights 18.1 GB vs 32.0 GB VRAM: fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 3060 Laptop GPU 6GB 6.00 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 6.00 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
Radeon RX 7600 8GB 8.00 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 8.00 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 4060 8GB 8.00 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 8.00 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 3080 10GB 10.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 10.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 3060 12GB 12.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 4070 12GB 12.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 5070 12GB 12.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
Radeon RX 7800 XT 16GB 16.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
Radeon RX 9070 XT 16GB 16.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 4060 Ti 16GB 16.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 4080 16GB 16.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 5060 Ti 16GB 16.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
GeForce RTX 5080 16GB 16.0 GB VRAM Potential: weights need offload not verified · Q4_K_M Weights 18.1 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not included
Intel Arc B580 12GB 12.0 GB VRAM Not enough data not verified · Q4_K_M Weights 18.1 GB vs 12.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not includedruntime support on this machine: unknown
Intel Arc A770 16GB 16.0 GB VRAM Not enough data not verified · Q4_K_M Weights 18.1 GB vs 16.0 GB VRAM: do not fitKV cache at 8k: attention layers only: recurrent state not includedruntime support on this machine: unknown
Mac mini M4 24GB 24.0 GB unified Not enough data not verified · Q4_K_M Weights 18.1 GB vs 24.0 GB unified: fitKV cache at 8k: attention layers only: recurrent state not included
Mac mini M4 Pro 24GB 24.0 GB unified Not enough data not verified · Q4_K_M Weights 18.1 GB vs 24.0 GB unified: fitKV cache at 8k: attention layers only: recurrent state not included
MacBook Pro 14-inch M4 Max 36GB 36.0 GB unified Not enough data not verified · Q4_K_M Weights 18.1 GB vs 36.0 GB unified: fitKV cache at 8k: attention layers only: recurrent state not included

Definitive answers ("runs", "does not fit") come from the calibrated engine or a physical run. "Potential" rows only compare the weights with memory: the rest of the requirement is not modelled, so they are not verdicts. Unknown is never a failure.

Hardware for this model

At least 24 GB of VRAM is needed just to hold the smallest weights (18.1 GB). The complete requirement is higher and not modelled yet.

  • Smallest catalogued artifact: 18.1 GB of weights (Q4_K_M).
  • A GPU with at least 24 GB of VRAM holds those weights.
  • Not modelled yet: attention layers only: recurrent state not included.
These links open a hardware category search, not a specific product recommendation. Affiliate link — I may earn a commission from qualifying purchases. These links never change a compatibility answer. They follow the memory the data defends, not commission.

Sources and evidence

0 measured, 23 sourced, 4 estimated and 3 unknown fields.

Catalogue record verified 2026-09-17.

What else can your PC run?

Pick your GPU and system RAM to see every catalogued model that fits, with tested, estimated and potential answers kept apart.

Check my PC