VRAM Calculator
Find out what AI models your hardware can actually run.
This tool is coming soon. For now, here’s the manual version:
The Math
VRAM needed = (model parameters × quantization bits) ÷ 8 + overhead
| Quantization | Bits per parameter | Use case |
|---|---|---|
| Q4_K_M | 4 | Best quality/size ratio — recommended |
| Q5_K_M | 5 | Higher quality, more VRAM |
| Q8_0 | 8 | Near-original quality, 2x size |
| F16 | 16 | Full precision, rarely needed locally |
Quick Reference Table
| Model Size | Q4_K_M VRAM | Q8 VRAM | Min RAM (no GPU) |
|---|---|---|---|
| 3B | 2.5 GB | 4 GB | 8 GB |
| 7B | 5 GB | 8 GB | 16 GB |
| 13B | 8 GB | 14 GB | 32 GB |
| 30B | 18 GB | 32 GB | 64 GB |
| 70B | 42 GB | 75 GB | 128 GB |
GPU Tier Guide
| GPU | VRAM | Best Models (Q4) |
|---|---|---|
| RTX 3060 / 4060 | 8 GB | 3B, 7B |
| RTX 3080 / 4070 | 10-12 GB | 7B, 9B |
| RTX 3090 / 4090 | 24 GB | 13B, 30B (tight) |
| 2× RTX 3090 | 48 GB | 30B, 70B (tight) |
Coming Soon: Interactive Calculator
We’re building a free tool that will:
- Input your hardware (GPU model, VRAM, RAM)
- Output compatible models with expected performance
- Show optimal settings (quantization, context length, batch size)
- Generate deployment scripts for your exact setup
Paid tier ($29/mo) unlocks:
- Full Docker Compose scripts for your hardware
- Monitoring dashboard templates (Grafana + Prometheus)
- Pre-built n8n AI workflow templates
- Priority tutorial requests