Intel · consumer

Arc B570 10GB

Arc B570 10GB has 10 GB of VRAM at 380 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1658 of 2118 indexed models fit at 16K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
380 GB/s
160-bit bus
Tensor FP16
dense
TDP
150 W
$219 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1433vision language 125image 2video 12audio asr 39embedding 26audio tts 21

What fits at 16K context

largest quantization that fits, per model · 1658 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Devstral-Small-2-24B-Instruct-2512UD-IQ2_M24.0B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
Mistral-Small-3.2-24B-Instruct-2506UD-IQ2_M24.0B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
Devstral-Small-2507UD-IQ2_M23.6B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
Devstral-Small-2505UD-IQ2_M23.6B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
Magistral-Small-2509UD-IQ2_M24.0B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
Magistral-Small-2507UD-IQ2_M23.6B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
Mistral-Small-3.1-24B-Instruct-2503UD-IQ2_M24.0B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
Magistral-Small-2506UD-IQ2_M23.6B7.68 GiB0.70 GiB9.30 GiB0.00 GiB24±30%
OmniAtlas-Qwen3-30B-A3BI1-IQ2_XS31.7B8.45 GiB0.00 GiB9.30 GiB0.00 GiB24±30%
Qwen3-Omni-30B-A3B-CaptionerI1-IQ2_XS31.7B8.45 GiB0.00 GiB9.30 GiB0.00 GiB24±30%
v6-Finch-7B-HFQ6_K7.6B6.19 GiB2.25 GiB9.28 GiB0.02 GiB24±30%
rwkv-6-world-7bQ6_K7.6B6.19 GiB2.25 GiB9.28 GiB0.02 GiB24±30%
Qwen3-VL-30B-A3B-ThinkingMoEIQ2_XS31.1B8.07 GiB0.42 GiB9.28 GiB0.02 GiB75±37%
MiroThinker-v1.0-30BMoEIQ2_XS30.5B8.07 GiB0.42 GiB9.28 GiB0.02 GiB75±37%
Qwen3-30B-A3B-Instruct-2507MoEIQ2_XS30.5B8.07 GiB0.42 GiB9.28 GiB0.02 GiB75±37%
Qwen3-30B-A3B-Thinking-2507MoEIQ2_XS30.5B8.07 GiB0.42 GiB9.28 GiB0.02 GiB75±37%
Tiger-Gemma-12B-v3Q4_K_L12.8B8.02 GiB0.41 GiB9.28 GiB0.02 GiB24±30%
spoomplesmaxx-v2.1-30BI1-IQ2_XXS28.9B7.25 GiB1.13 GiB9.28 GiB0.02 GiB24±30%
Huihui-granite-4.1-30b-abliteratedI1-IQ2_XXS28.9B7.25 GiB1.13 GiB9.28 GiB0.02 GiB24±30%
granite-4.1-30b-hereticI1-IQ2_XXS28.9B7.25 GiB1.13 GiB9.28 GiB0.02 GiB24±30%
Tongyi-DeepResearch-30B-A3BMoEIQ2_XS30.5B8.07 GiB0.42 GiB9.28 GiB0.02 GiB75±37%
GLM-4.7-Flash-REAP-23B-A3BMoEQ2_K_L23.0B8.24 GiB0.23 GiB9.28 GiB0.02 GiB75±37%
Muse-Glimmer-30BIQ2_XXS29.8B8.31 GiB0.08 GiB9.28 GiB0.02 GiB24±30%
AceReason-Nemotron-14BIQ4_XS14.8B7.58 GiB0.84 GiB9.27 GiB0.03 GiB24±30%
Skywork-R1V3-38BIQ2_XXS38.4B8.41 GiB0.00 GiB9.27 GiB0.03 GiB24±30%
HunyuanImage-2.1Q3_K_S17.5B8.42 GiB0.00 GiB9.27 GiB0.03 GiB24±30%
EuroLLM-22B-Instruct-2512IQ2_M22.6B7.45 GiB0.95 GiB9.27 GiB0.03 GiB24±30%
medgemma-27b-itI1-IQ2_XS28.8B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
gemma-3-27b-it-abliterated-refined-visionI1-IQ2_XS27.4B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
Nidum-Gemma-3-27B-it-UncensoredI1-IQ2_XS27.4B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
gemma-3-27b-it-abliteratedIQ2_XS27.4B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
AtomicGPT-gemma3-27bI1-IQ2_XS27.4B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
NVIDIA-Nemotron-Nano-12B-v2Q4_112.3B7.30 GiB1.09 GiB9.26 GiB0.04 GiB24±30%
Unbound-v1.12.0-27BI1-IQ2_XS27.4B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
Mira-v1.12-Ties-27BI1-IQ2_XS27.4B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
gemma-3-27b-itIQ2_XS27.4B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
Medgamma27BI1-IQ2_XS27.0B7.86 GiB0.52 GiB9.26 GiB0.04 GiB24±30%
gemma-4-A4B-98e-v7-coder-itMoEQ2_K20.5B8.22 GiB0.26 GiB9.26 GiB0.04 GiB24±30%
gemma-4-A4B-98e-v7-coderx-itMoEQ2_K20.5B8.22 GiB0.26 GiB9.26 GiB0.04 GiB24±30%
DeepSeek-Coder-V2-Lite-BaseMoEI1-Q4_015.7B8.32 GiB0.13 GiB9.26 GiB0.04 GiB75±37%
OLMo-2-1124-13B-InstructQ2_K13.7B4.90 GiB3.52 GiB9.26 GiB0.04 GiB24±30%
codegeex4-all-9bQ4_19.4B5.59 GiB2.81 GiB9.25 GiB0.05 GiB24±30%
EVA-abliterated-TIES-Qwen2.5-14BI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Neuron-V1-14B-InstructI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Ektome-Qwen2.5-Coder-14B-Instruct-PristinelyUncensoredI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Qwen2.5-14B-Instruct-1M-abliteratedI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
DeepCoder-14B-PreviewIQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Deepseeker-Kunou-Qwen2.5-14bI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
SuperNova-MediusIQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
14B-Qwen2.5-Kunou-v1I1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Sugoi-14B-Ultra-HFI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Qwen2.5-Coder-14B-Instruct-abliteratedIQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
OpenCodeReasoning-Nemotron-14BIQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Qwen2.5-14B-InstructIQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
DeepSeek-R1-Distill-Qwen-14B-abliterated-v2I1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
C1-TachuI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
DeepSeek-R1-Distill-Qwen-14B-abliteratedI1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
0x-liteIQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Qwen2.5-Coder-14B-InstructIQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
Tessera-4I1-IQ4_XS14.8B7.56 GiB0.84 GiB9.25 GiB0.05 GiB24±30%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Arc B570 10GB run?
1658 of 2118 indexed open-weight models fit a Arc B570 10GB at 16,384 context with q4_0 KV cache, the largest being Devstral-Small-2-24B-Instruct-2512 at UD-IQ2_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Arc B570 10GB actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Arc B570 10GB fast for local AI?
Its memory bandwidth is 380 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.