Model comparison

bge-reranker-v2-m3 vs KaLM-embedding-multilingual-mini-instruct-v2.5

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: KaLM-embedding-multilingual-mini-instruct-v2.5's KV cache at 32K is 8.0× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

bge-reranker-v2-m3KaLM-embedding-multilingual-mini-instruct-v2.5
Parameters568M494M
Architecturebertqwen2
Layers2424
Native context8,194131,072
Mixture of expertsnono
Quantizations published51
Smallest quantization0.41 GiB0.49 GiB
Q4_K_M0.41 GiB
Licenceapache-2.0apache-2.0

KV cache by context

the term that decides long-context viability
Contextbge-reranker-v2-m3KaLM-embedding-multilingual-mini-instruct-v2.5Ratio
4,0960.38 GiB0.05 GiB8.00×
8,1920.75 GiB0.09 GiB8.00×
16,3841.50 GiB0.19 GiB8.00×
32,7683.00 GiB0.38 GiB8.00×
65,5366.00 GiB0.75 GiB8.00×
131,07212.00 GiB1.50 GiB8.00×