Method
What is included
The comparison includes the exact model blob, an explicit 8k KV cache estimate, a conservative planning margin and possible system-RAM offload. It does not predict speed or claim a physical test where none exists.
Exact Ollama variant
This page evaluates the immutable model blob behind qwen3:4b-instruct-2507-q4_K_M. Results below use 8k context and 32 GB system RAM.
Exact model build verified
SHA-256 digest: sha256:85e4a5b7b8ef0e48af0e8658f5aaab9c2324c76c1641493f4d1e25fce54b18b9
ollama run qwen3:4b-instruct-2507-q4_K_M| Exact GPU | VRAM | Context | Status | Main limitation |
|---|---|---|---|---|
| GeForce RTX 3060 Laptop GPU 6GB | 6.00 GB | 8k | Tested Runs entirely on GPU | No blocker observed |
| GeForce RTX 4060 8GB | 8.00 GB | 8k | Estimated Likely full GPU | physical test data |
| GeForce RTX 3060 12GB | 12.0 GB | 8k | Estimated Likely full GPU | physical test data |
| GeForce RTX 5070 12GB | 12.0 GB | 8k | Estimated Likely full GPU | physical test data |
| GeForce RTX 4060 Ti 16GB | 16.0 GB | 8k | Estimated Likely full GPU | physical test data |
| GeForce RTX 5060 Ti 16GB | 16.0 GB | 8k | Estimated Likely full GPU | physical test data |
| GeForce RTX 3090 24GB | 24.0 GB | 8k | Estimated Likely full GPU | physical test data |
| GeForce RTX 4090 24GB | 24.0 GB | 8k | Estimated Likely full GPU | physical test data |
| GeForce RTX 5090 32GB | 32.0 GB | 8k | Estimated Likely full GPU | physical test data |
Tested means the exact GPU, model digest, runtime and context were physically tested. Estimated never means tested.
Method
The comparison includes the exact model blob, an explicit 8k KV cache estimate, a conservative planning margin and possible system-RAM offload. It does not predict speed or claim a physical test where none exists.
Provenance