Intel · workstation

Arc Pro A30M 4GB

Arc Pro A30M 4GB has 4 GB of VRAM at 112 GB/s — about 3.72 GiB usable after driver and compositor overhead. 642 of 2118 indexed models fit at 16K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
4 GB
GDDR6
Bandwidth
112 GB/s
64-bit bus
Tensor FP16
dense
TDP
50 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 511audio asr 36vision language 51embedding 23video 2audio tts 19

What fits at 16K context

largest quantization that fits, per model · 642 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
OpenCoder-1.5B-InstructQ3_K_L1.9B1.18 GiB1.74 GiB3.72 GiB0.00 GiB21±30%
Voxtral-Mini-3B-2507Q2_K_L4.7B1.91 GiB1.00 GiB3.72 GiB0.00 GiB21±30%
CycleGRPO-4BI1-IQ3_XXS4.8B1.71 GiB1.20 GiB3.72 GiB0.00 GiB21±30%
Jan-v3-4B-base-instructIQ3_XXS4.4B1.71 GiB1.20 GiB3.72 GiB0.00 GiB21±30%
G9v3-3BQ6_K_L3.0B2.49 GiB0.43 GiB3.72 GiB0.00 GiB21±30%
Darwin-4B-ChimeraQ4_K_M4.0B2.43 GiB0.48 GiB3.72 GiB0.00 GiB21±30%
alduin-4b-it-baseI1-Q5_K_M4.3B2.64 GiB0.26 GiB3.72 GiB0.00 GiB21±30%
gemma-4-E2B-it-abliteratedI1-IQ3_XS5.1B2.85 GiB0.07 GiB3.72 GiB0.00 GiB21±30%
gemma-4-E2B-it-qat-q4_0-unquantized-hereticI1-IQ3_XS5.1B2.85 GiB0.07 GiB3.72 GiB0.00 GiB21±30%
Huihui-gemma-4-E2B-it-qat-q4_0-unquantized-abliteratedI1-IQ3_XS5.1B2.85 GiB0.07 GiB3.72 GiB0.00 GiB21±30%
Gemma4_E2B_Abliterated_Baked_HF_ReadyI1-IQ3_XS5.1B2.85 GiB0.07 GiB3.72 GiB0.00 GiB21±30%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseI1-IQ2_XXS8.1B2.65 GiB0.23 GiB3.72 GiB0.00 GiB21±30%
deepseek-coder-1.3b-instructQ8_01.3B1.33 GiB1.59 GiB3.72 GiB0.00 GiB21±30%
deepseek-coder-1.3b-baseQ8_01.3B1.33 GiB1.59 GiB3.72 GiB0.00 GiB21±30%
granite-8b-code-instruct-4kI1-IQ1_S8.1B1.68 GiB1.20 GiB3.71 GiB0.01 GiB21±30%
granite-8b-code-base-4kI1-IQ1_S8.1B1.68 GiB1.20 GiB3.71 GiB0.01 GiB21±30%
Qwen2.5-3BQ6_K3.1B2.60 GiB0.30 GiB3.71 GiB0.01 GiB21±30%
GRM-Kerlin-3bI1-Q6_K3.4B2.60 GiB0.30 GiB3.71 GiB0.01 GiB21±30%
Qwen2.5-Coder-3B-InstructQ6_K3.1B2.60 GiB0.30 GiB3.71 GiB0.01 GiB21±30%
Qwen2.5-3B-InstructQ6_K3.1B2.60 GiB0.30 GiB3.71 GiB0.01 GiB21±30%
LCO-Embedding-Omni-3B-2605Q6_K4.7B2.60 GiB0.30 GiB3.71 GiB0.01 GiB21±30%
Garnet-OCR-3B-0422I1-Q6_K4.1B2.60 GiB0.30 GiB3.71 GiB0.01 GiB21±30%
gemma-3-4b-it-roleplay-tuned-v1I1-Q5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
Gemma-3-R1984-4BQ5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-ThinkingQ5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
gemma-3-4b-it-roleplay-tuned-v2I1-Q5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
medgemma-1.5-4b-itQ5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
gemma-3-4b-it-heretic-uncensored-abliterated-ExtremeI1-Q5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
medgemma-4b-itQ5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
gemma-3-4b-it-abliteratedQ5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
amoral-gemma3-4B-v1Q5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
Gemma3-4B-CodeCenturionI1-Q5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
ArrowMint-Gemma3-4B-YUKI-v0.1I1-Q5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
gemma-3-4b-itQ5_K_M4.3B2.64 GiB0.26 GiB3.71 GiB0.01 GiB21±30%
Dolphin3.0-Llama3.2-3BQ4_K_L3.2B1.97 GiB0.93 GiB3.71 GiB0.01 GiB21±30%
Llama-Song-Stream-3B-InstructQ4_K_L3.2B1.97 GiB0.93 GiB3.71 GiB0.01 GiB21±30%
Llama-Doctor-3.2-3B-InstructQ4_K_L3.2B1.97 GiB0.93 GiB3.71 GiB0.01 GiB21±30%
llama-3.2-Korean-Bllossom-3BQ4_K_L3.2B1.97 GiB0.93 GiB3.71 GiB0.01 GiB21±30%
Llama-3.2-3B-InstructQ4_K_L3.2B1.97 GiB0.93 GiB3.71 GiB0.01 GiB21±30%
Hermes-3-Llama-3.2-3BQ4_K_L3.2B1.97 GiB0.93 GiB3.71 GiB0.01 GiB21±30%
Jan-code-4bIQ3_XXS4.4B1.70 GiB1.20 GiB3.71 GiB0.01 GiB21±30%
SmolLM2-1.7B-InstructQ6_K1.7B1.31 GiB1.59 GiB3.70 GiB0.02 GiB21±30%
gemma-4-E2B-itQ4_K_S5.1B2.83 GiB0.07 GiB3.70 GiB0.02 GiB21±30%
Nanbeige4.2-3BQ3_K_L4.2B2.15 GiB0.73 GiB3.70 GiB0.02 GiB21±30%
umt5-xxlQ3_K_M5.7B2.85 GiB0.00 GiB3.70 GiB0.02 GiB21±30%
Qwen3-VL-4B-Instruct-Unredacted-MAXI1-IQ3_XS4.4B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Qwen3-VL-4B-Thinking-Unredacted-MAXI1-IQ3_XS4.4B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Zubr1.0-VL-4BI1-IQ3_XS4.4B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Huihui-Qwen3-VL-4B-Instruct-abliteratedI1-IQ3_XS4.4B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Qwen3-VL-4B-Instruct-UncensoredI1-IQ3_XS4.4B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
OpenCaption-4B-VL-SFT-v1.0I1-IQ3_XS4.4B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Parable-Qwen3-4B-Claude-Fable-5I1-IQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Jan-v1-4BIQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Qwen3-4b-Z-Image-Turbo-AbliteratedV1I1-IQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Jan-nano-128kIQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Qwen3-4B-Instruct-2507IQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Qwen3-4B-Thinking-2507IQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Jan-nanoIQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Qwen3-4B-abliteratedIQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
Neuron-4B-InstructI1-IQ3_XS4.0B1.69 GiB1.20 GiB3.70 GiB0.02 GiB21±30%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Arc Pro A30M 4GB run?
642 of 2118 indexed open-weight models fit a Arc Pro A30M 4GB at 16,384 context with q8_0 KV cache, the largest being OpenCoder-1.5B-Instruct at Q3_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a Arc Pro A30M 4GB actually have?
Its nameplate is 4 GB, but about 3.72 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Arc Pro A30M 4GB fast for local AI?
Its memory bandwidth is 112 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.