DavidAU · text

L3.2-Rogue-Creative-Instruct-Uncensored-Abliterated-7B

DavidAU/L3.2-Rogue-Creative-Instruct-Uncensored-Abliterated-7B

L3.2-Rogue-Creative-Instruct-Uncensored-Abliterated-7B at Q4_K_M is exactly 4,588,966,592 bytes (4.27 GiB / 4.59 GB) — an effective 4.873 bits per weight, not the nominal 4. Its KV cache at 32K is 8.38 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
7.5B
Architecture
llama
67 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K2.73 GiB2,931,892,9283.114DavidAU
Q3_K_S3.17 GiB3,400,022,7203.611DavidAU
Q3_K_M3.49 GiB3,749,296,8323.982DavidAU
IQ4_XS3.87 GiB4,156,478,1444.414DavidAU
Q4_K_S4.07 GiB4,374,835,9044.646DavidAU
Q4_K_M4.27 GiB4,588,966,5924.873DavidAU
Q5_K_S4.88 GiB5,240,402,6245.565DavidAU
Q5_K_M5.00 GiB5,364,486,8485.697DavidAU
Q6_K5.76 GiB6,188,477,1206.572DavidAU
Q8_07.46 GiB8,012,741,3128.510DavidAU

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.05 GiB1.05 GiB67 / 0 / 0
8,1922.09 GiB2.09 GiB67 / 0 / 0
16,3844.19 GiB4.19 GiB67 / 0 / 0
32,7688.38 GiB8.38 GiB67 / 0 / 0
65,53616.75 GiB16.75 GiB67 / 0 / 0
131,07233.50 GiB33.50 GiB67 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 3.95 GiB. The real file is 4.27 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
67
Attention heads
24
KV heads
8
Head dim
128
Hidden size
3072
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does L3.2-Rogue-Creative-Instruct-Uncensored-Abliterated-7B need?
Q4_K_M is exactly 4,588,966,592 bytes (4.27 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is L3.2-Rogue-Creative-Instruct-Uncensored-Abliterated-7B's KV cache?
8.38 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of L3.2-Rogue-Creative-Instruct-Uncensored-Abliterated-7B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.