EpistemeAI · text

Athene-Phi-3.5-mini-instruct-orpo

EpistemeAI/Athene-Phi-3.5-mini-instruct-orpo

Athene-Phi-3.5-mini-instruct-orpo at Q4_K_M is exactly 2,318,920,352 bytes (2.16 GiB / 2.32 GB) — an effective 4.855 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
3.8B
Architecture
llama
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M1.26 GiB1,349,430,9442.825bartowski
Q2_K1.35 GiB1,446,880,9283.029bartowski
Q2_K_L1.44 GiB1,543,072,9283.231bartowski
IQ3_XS1.49 GiB1,596,869,7923.343bartowski
IQ3_XS1.49 GiB1,596,869,9203.343mradermacher
Q3_K_S1.57 GiB1,681,804,4483.521bartowski
IQ3_S1.57 GiB1,681,804,5763.521mradermacher
I1-IQ1_S2 shards1.64 GiB1,763,447,8723.692mradermacher
IQ3_M1.65 GiB1,775,389,8563.717bartowski
IQ3_M1.65 GiB1,775,389,9843.717mradermacher
Q3_K_M1.75 GiB1,877,626,0163.931bartowski
I1-IQ1_M2 shards1.77 GiB1,900,287,0403.978mradermacher
Q3_K_L1.90 GiB2,045,136,0324.282bartowski
IQ4_XS1.92 GiB2,059,858,5924.313bartowski
I1-IQ2_XXS2 shards1.98 GiB2,128,352,3204.456mradermacher
Q4_02.03 GiB2,182,474,4004.569bartowski
Q4_K_S2.04 GiB2,193,484,4484.592bartowski
Q4_K_M2.16 GiB2,318,920,3524.855bartowski
I1-IQ2_XS2 shards2.17 GiB2,329,678,9124.878mradermacher
Q4_K_L2.23 GiB2,392,026,2725.008bartowski
I1-IQ2_S2 shards2.34 GiB2,516,410,4325.269mradermacher
Q5_K_S2.46 GiB2,641,480,3525.530bartowski
I1-IQ2_M2 shards2.51 GiB2,698,862,6565.651mradermacher
Q5_K_M2.53 GiB2,715,011,7445.684bartowski
Q5_K_L2.59 GiB2,775,805,0885.812bartowski
Q2_K2 shards2.70 GiB2,893,762,1446.059mradermacher
I1-Q2_K2 shards2.70 GiB2,893,762,6246.059mradermacher
I1-IQ3_XXS2 shards2.75 GiB2,950,520,8966.177mradermacher
Q6_K2.92 GiB3,135,858,8486.565bartowski
Q6_K_L2.96 GiB3,183,570,0806.665bartowski
I1-IQ3_XS2 shards2.97 GiB3,193,740,3526.687mradermacher
Q3_K_S2 shards3.13 GiB3,363,609,1847.042mradermacher
I1-IQ3_S2 shards3.13 GiB3,363,609,6647.042mradermacher
I1-Q3_K_S2 shards3.13 GiB3,363,609,6647.042mradermacher
I1-IQ3_M2 shards3.31 GiB3,550,780,4807.434mradermacher
Q3_K_M2 shards3.50 GiB3,755,252,3207.862mradermacher
I1-Q3_K_M2 shards3.50 GiB3,755,252,8007.862mradermacher
Q8_03.78 GiB4,061,228,1928.503bartowski
Q3_K_L2 shards3.81 GiB4,090,272,3528.564mradermacher
I1-Q3_K_L2 shards3.81 GiB4,090,272,8328.564mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 2.00 GiB. The real file is 2.16 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Athene-Phi-3.5-mini-instruct-orpo need?
Q4_K_M is exactly 2,318,920,352 bytes (2.16 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Athene-Phi-3.5-mini-instruct-orpo should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.