Model comparison

granite-embedding-107m-multilingual vs gemma-3-270m-it

At Q4_K_M, granite-embedding-107m-multilingual is the smaller download — 117,011,136 bytes against 253,115,168. At long context the gap widens: gemma-3-270m-it's KV cache at 32K is 2.6× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

granite-embedding-107m-multilingualgemma-3-270m-it
Parameters107M268M
Architecturebertgemma3
Layers618
Native context51432,768
Mixture of expertsnono
Quantizations published1939
Smallest quantization0.11 GiB0.17 GiB
Q4_K_M0.11 GiB0.24 GiB
Licenceapache-2.0gemma

KV cache by context

the term that decides long-context viability
Contextgranite-embedding-107m-multilingualgemma-3-270m-itRatio
4,0960.04 GiB0.03 GiB1.33×
8,1920.07 GiB0.04 GiB1.85×
16,3840.14 GiB0.06 GiB2.29×
32,7680.28 GiB0.11 GiB2.59×
65,5360.56 GiB0.20 GiB2.78×
131,0721.13 GiB0.39 GiB2.89×