microsoft · text

phi-4

microsoft/phi-4

phi-4 at Q4_K_M is exactly 8,890,306,112 bytes (8.28 GiB / 8.89 GB) — an effective 4.852 bits per weight, not the nominal 4. Its KV cache at 32K is 6.25 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
14.7B
Architecture
phi3
40 layers
Context
16,384
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XS4.18 GiB4,485,317,0562.448bartowski
IQ2_S4.41 GiB4,731,548,0962.582bartowski
IQ2_M4.76 GiB5,110,428,0962.789bartowski
Q2_K5.17 GiB5,547,348,4163.027bartowski
Q2_K5.22 GiB5,608,795,7123.061unsloth
Q2_K_L5.34 GiB5,729,218,1123.127unsloth
Q2_K_L5.63 GiB6,049,108,4163.301bartowski
IQ3_XS5.82 GiB6,246,699,4563.409bartowski
Q3_K_S6.06 GiB6,504,747,4563.550bartowski
IQ3_M6.44 GiB6,913,835,4563.773bartowski
Q3_K_M6.70 GiB7,190,834,7523.924363unsloth
Q3_K_M6.86 GiB7,363,269,0564.018243bartowski
Q3_K_L7.39 GiB7,930,155,2004.328lmstudio-community
Q3_K_L7.39 GiB7,930,155,4564.328bartowski
IQ4_XS7.40 GiB7,941,378,4964.334243bartowski
IQ4_NL7.81 GiB8,383,418,8164.575bartowski
Q4_07.83 GiB8,412,090,8164.591243bartowski
Q4_K_S7.86 GiB8,440,762,8164.606bartowski
Q4_K_M8.28 GiB8,890,306,1124.852363unsloth
Q4_K_M8.43 GiB9,053,114,5604.941lmstudio-community
Q4_K_M8.43 GiB9,053,114,8164.941243bartowski
Q4_18.63 GiB9,267,499,4565.058bartowski
Q4_K_L8.79 GiB9,434,452,4165.149bartowski
Q5_K_S9.45 GiB10,151,580,0965.540bartowski
Q5_K_M9.70 GiB10,412,707,3925.682363unsloth
Q5_K_M9.88 GiB10,604,188,0965.787243bartowski
Q5_K_L10.17 GiB10,921,300,4165.960bartowski
Q6_K11.20 GiB12,030,251,2006.565243lmstudio-community
Q6_K11.20 GiB12,030,251,4566.565bartowski
Q6_K11.20 GiB12,030,258,7526.565363unsloth
Q6_K_L11.44 GiB12,279,124,4166.701bartowski
Q8_014.51 GiB15,580,500,1608.503243lmstudio-community
Q8_014.51 GiB15,580,500,4168.503243bartowski
Q8_014.51 GiB15,580,507,7128.503unsloth
F1627.31 GiB29,323,399,61616.002bartowski
F1627.31 GiB29,323,406,91216.002363unsloth
F322 shards54.61 GiB58,641,584,54432.002bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.78 GiB0.78 GiB40 / 0 / 0
8,1921.56 GiB1.56 GiB40 / 0 / 0
16,3843.13 GiB3.13 GiB40 / 0 / 0
32,7686.25 GiB6.25 GiB40 / 0 / 0
65,53612.50 GiB12.50 GiB40 / 0 / 0
131,07225.00 GiB25.00 GiB40 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 7.68 GiB. The real file is 8.28 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
40
Attention heads
40
KV heads
10
Head dim
128
Hidden size
5120
Vocab
100,352
Sliding window
none
SWA period
1
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does phi-4 need?
Q4_K_M is exactly 8,890,306,112 bytes (8.28 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is phi-4's KV cache?
6.25 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of phi-4 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.