LLM model memory guides
Compare memory scenarios before choosing a GPU or changing precision and context.
These guides use pinned metadata and our calculator engine. Totals are planning estimates, not measured peak VRAM. Only models with complete text-generation scenarios are included. Search or paste any HF model in the calculator to inspect other models.
- deepseek-ai/DeepSeek-R1-Distill-Qwen-14B14.77B parameters · qwen2 · context up to 131,072
- meta-llama/Llama-3.3-70B-Instruct70.554B parameters · llama · context up to 131,072
- mistralai/Mistral-Large-Instruct-2411122.61B parameters · mistral · context up to 131,072
- mistralai/Mistral-Small-Instruct-240922.247B parameters · mistral · context up to 32,768
- microsoft/Phi-3-medium-128k-instruct13.96B parameters · phi3 · context up to 131,072
- microsoft/Phi-3.5-mini-instruct3.821B parameters · phi3 · context up to 131,072
- microsoft/Phi-3.5-MoE-instruct41.873B parameters · phimoe · context up to 131,072
- microsoft/Phi-4-mini-instruct3.836B parameters · phi3 · context up to 131,072
- Qwen/Qwen3-Coder-Next79.674B parameters · qwen3_next · context up to 262,144
Reading these comparisons
Each guide separates weight storage, cache/state and an explicit runtime budget. Higher precision increases weight storage; longer contexts usually increase cache memory. Use the calculator links to change a scenario and select your GPU. Memory capacity alone does not establish runtime support or generation speed.