deepreinforce-ai · text

Ornith-1.0-35B-GGUF

deepreinforce-ai/Ornith-1.0-35B-GGUF

Ornith-1.0-35B-GGUF at Q4_K_M is exactly 21,166,757,760 bytes (19.71 GiB / 21.17 GB) — an effective 4.886 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
34.7B
Architecture
qwen35moe
Context
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S6.97 GiB7,484,141,4401.727liodon-ai
IQ2_XS9.79 GiB10,507,030,4002.425liodon-ai
IQ2_S9.92 GiB10,652,479,3602.459liodon-ai
IQ2_M10.86 GiB11,659,235,2002.691liodon-ai
Q2_K12.05 GiB12,939,593,6002.987liodon-ai
IQ3_XS13.49 GiB14,484,144,0003.343liodon-ai
IQ3_M14.38 GiB15,440,519,0403.564liodon-ai
Q3_K_M15.61 GiB16,764,764,0323.869733liodon-ai
IQ4_XS17.44 GiB18,728,777,6004.323733liodon-ai
Q4_K_M19.71 GiB21,166,757,7604.886733liodon-ai
Q5_K_M23.03 GiB24,729,130,8805.708733liodon-ai
Q6_K26.56 GiB28,514,152,3206.581733liodon-ai
Q8_034.37 GiB36,903,139,2008.518733liodon-ai

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 18.16 GiB. The real file is 19.71 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Ornith-1.0-35B-GGUF need?
Q4_K_M is exactly 21,166,757,760 bytes (19.71 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Ornith-1.0-35B-GGUF should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.
Ornith-1.0-35B-GGUF — VRAM requirements, exact quant sizes — ossmodeldb