Model memory guides / microsoft/Phi-3.5-MoE-instruct
microsoft/Phi-3.5-MoE-instruct VRAM estimates
For microsoft/Phi-3.5-MoE-instruct, compare how precision and context change memory planning for text generation.
What changes the memory requirement?
In the Q4_K_M scenario (4.85 effective bits/weight), reported parameters account for approximately 23.64 GiB of weights. At 2,048 tokens, cache and recurrent state add 0.25 GiB. At 32,768 tokens, they add 4.00 GiB. The metadata declares a maximum context of 131,072 tokens; this is not a guarantee of useful output quality at that length.
Reported parameters
41.873 billion
Cache architecture
phimoe
Metadata snapshot
2026-08-31
Weight, cache and planning totals
One sequence, F16 cache, full residency on one device. Context includes both prompt and generated tokens. All memory values are GiB (2³⁰ bytes).
| Precision / effective BPW | Context tokens | Weights | Cache & state | Planning total (+1 GiB budget) | Adjust |
|---|---|---|---|---|---|
| Q4_K_M scenario4.85 bits/weight | 2,048 | 23.64 | 0.25 | 24.89 | Calculator |
| Q4_K_M scenario4.85 bits/weight | 8,192 | 23.64 | 1.00 | 25.64 | Calculator |
| Q4_K_M scenario4.85 bits/weight | 32,768 | 23.64 | 4.00 | 28.64 | Calculator |
| Q8_0 scenario8.5 bits/weight | 2,048 | 41.43 | 0.25 | 42.68 | Calculator |
| Q8_0 scenario8.5 bits/weight | 8,192 | 41.43 | 1.00 | 43.43 | Calculator |
| Q8_0 scenario8.5 bits/weight | 32,768 | 41.43 | 4.00 | 46.43 | Calculator |
| 16-bit scenario16 bits/weight | 2,048 | 77.99 | 0.25 | 79.24 | Calculator |
| 16-bit scenario16 bits/weight | 8,192 | 77.99 | 1.00 | 79.99 | Calculator |
| 16-bit scenario16 bits/weight | 32,768 | 77.99 | 4.00 | 82.99 | Calculator |
Q4_K_M and Q8_0 here are assumed effective-BPW scenarios, not inspected GGUF files. Tensor mixtures and conversions can change actual weight storage. The calculator lets you inspect an actual artifact and choose your GPU before presenting a personalized result.
How these numbers are calculated
Weights = reported parameter count × effective bits per weight ÷ 8. Cache and state use the architecture-specific equations in the same engine as our calculator. Planning total = weights + cache/state + the stated runtime budget. For MoE, full residency includes all experts, not just active parameters per token.
Profile: llama.cpp / CUDA, Flash Attention enabled, micro-batch 512, prompt batch 2,048, no recurrent rollback snapshots. Engine 3.5.0; pinned runtime source. Multi-GPU, partial offload, training, speculative decoding and image/audio processing are outside these scenarios.
Model-specific assumptions and limitations
- head_dim uses the documented architecture default (128).
- The pinned llama.cpp loader uses full cache for this architecture despite the declared sliding window. The estimate follows that native allocation, not the Hugging Face attention implementation.
- Uses native llama.cpp architecture semantics; custom Hugging Face code is not executed.
- MoE weights include all experts in this full-residency scenario, not only the experts active per token.
- Weights use parameter count × assumed effective BPW; mixed tensor types and conversion can change artifact size.
- Runtime allowance is an editable planning budget, not a measured compute-buffer prediction or an error bound.
- No model execution or hardware compatibility is certified. Available device memory may be lower than listed capacity.
Sources and reproducibility
Configuration at the pinned model revision · Model repository at that revision · Equations and external evidence
The parameter count and configuration come from our bundled metadata snapshot. This page is generated from that snapshot; it does not claim a live metadata check or a GPU benchmark.
Revision: 43688451b462a3351d8580625ebe1931adb3986d. Open a table scenario to inspect that revision in the calculator.
Compare other models
- mistralai/Mistral-Small-Instruct-2409 — 22.247B parameters, mistral
- deepseek-ai/DeepSeek-R1-Distill-Qwen-14B — 14.77B parameters, qwen2
- microsoft/Phi-3-medium-128k-instruct — 13.96B parameters, phi3