GeForce RTX 4060 8GB: Ollama model compatibility | LocalAIReady
← All NVIDIA GPUs

NVIDIA GPU compatibility

GeForce RTX 4060 8GB

At 8k context with 32 GB system RAM, 7 of 15 listed Ollama variants fit the full-GPU planning envelope or have an exact full-GPU test.

8.00 GB physical VRAM CUDA Ollama backend 8k tokens comparison context
Check a custom RAM or context setup →

Available Ollama models

Exact model Weights Context Status Main limitation
Phi-4 mini phi4-mini:3.8b-q4_K_M · Q4_K_M 2.32 GB 8k Estimated Likely full GPU physical test data
Qwen3.5 4B qwen3.5:4b-q4_K_M · Q4_K_M 3.16 GB 8k Estimated Likely full GPU physical test data
Qwen3 qwen3:4b-instruct-2507-q4_K_M · Q4_K_M 2.33 GB 8k Estimated Likely full GPU physical test data
Gemma 3 4B gemma3:4b · Q4_K_M 3.11 GB 8k Estimated Likely full GPU physical test data
Qwen2.5 Coder qwen2.5-coder:7b-instruct-q4_K_M · Q4_K_M 4.36 GB 8k Estimated Likely full GPU physical test data
DeepSeek-R1 8B deepseek-r1:8b · Q4_K_M 4.87 GB 8k Estimated Likely full GPU physical test data
Qwen3.5 9B qwen3.5:9b-q4_K_M · Q4_K_M 6.14 GB 8k Estimated Likely full GPU physical test data
Mistral Nemo mistral-nemo:12b-instruct-2407-q4_K_M · Q4_K_M 6.96 GB 8k Estimated Likely CPU/GPU offload VRAM
DeepSeek-R1 14B deepseek-r1:14b · Q4_K_M 8.37 GB 8k Estimated Likely CPU/GPU offload VRAM
Gemma 3 12B gemma3:12b · Q4_K_M 7.59 GB 8k Estimated Likely CPU/GPU offload VRAM
GPT-OSS gpt-oss:20b · MXFP4 12.8 GB 8k Estimated Likely CPU/GPU offload VRAM
Qwen3.5 27B qwen3.5:27b-q4_K_M · Q4_K_M 16.2 GB 8k Estimated Likely CPU/GPU offload VRAM
Gemma 3 27B gemma3:27b · Q4_K_M 16.2 GB 8k Estimated Likely CPU/GPU offload VRAM
DeepSeek-R1 32B deepseek-r1:32b · Q4_K_M 18.5 GB 8k Estimated Likely CPU/GPU offload VRAM
Qwen3.5 35B-A3B qwen3.5:35b-a3b-q4_K_M · Q4_K_M 22.2 GB 8k Estimated Likely CPU/GPU offload VRAM

Estimated results use a conservative memory planning margin and are not physical tests. Change RAM or context in the checker for a personalized result.

How to read this page

Method and limitations

Weights and the selected 8k KV cache are compared with physical VRAM. System RAM is considered only for possible CPU/GPU offload. Runtime overhead, driver differences and current memory pressure can change a real run, so an estimate never becomes a test.