Intel · workstation

Arc Pro A30M 4GB

Arc Pro A30M 4GB has 4 GB of VRAM at 112 GB/s — about 3.72 GiB usable after driver and compositor overhead. 802 of 2118 indexed models fit at 16K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
4 GB
GDDR6
Bandwidth
112 GB/s
64-bit bus
Tensor FP16
dense
TDP
50 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 59text 660audio asr 37embedding 25video 2audio tts 19

What fits at 16K context

largest quantization that fits, per model · 802 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
gemma-4-E2B-itIQ3_XS5.1B2.89 GiB0.04 GiB3.72 GiB0.00 GiB21±30%
NuExtract-1.5IQ2_M3.8B1.23 GiB1.69 GiB3.72 GiB0.00 GiB21±30%
Phi-3.5-mini-instructIQ2_M3.8B1.23 GiB1.69 GiB3.72 GiB0.00 GiB21±30%
Phi-3-mini-128k-instructIQ2_M3.8B1.23 GiB1.69 GiB3.72 GiB0.00 GiB21±30%
Phi-3.5-mini-instruct_UncensoredIQ2_M3.8B1.23 GiB1.69 GiB3.72 GiB0.00 GiB21±30%
Phi-3-mini-4k-instructIQ2_M3.8B1.23 GiB1.69 GiB3.72 GiB0.00 GiB21±30%
Darwin-4B-ChimeraQ5_K_S4.0B2.65 GiB0.25 GiB3.72 GiB0.00 GiB21±30%
Holo-3.1-4BI1-IQ4_NL5.2B2.76 GiB0.14 GiB3.72 GiB0.00 GiB21±30%
AfriqueQwen3.5-4BI1-IQ4_NL5.2B2.76 GiB0.14 GiB3.72 GiB0.00 GiB21±30%
TimeOmni-1-4BI1-IQ4_NL5.2B2.76 GiB0.14 GiB3.72 GiB0.00 GiB21±30%
LFM2-8B-A1BMoEQ2_K8.3B2.87 GiB0.05 GiB3.72 GiB0.00 GiB56±37%
Llama-3.2-3B-Instruct-abliteratedI1-Q5_K_M3.6B2.41 GiB0.49 GiB3.72 GiB0.00 GiB21±30%
Llama-3.2-3B-Instruct-uncensoredQ5_K_M3.6B2.41 GiB0.49 GiB3.72 GiB0.00 GiB21±30%
stable-code-3bI1-Q4_K_S2.8B1.51 GiB1.41 GiB3.71 GiB0.01 GiB21±30%
rocket-3BQ4_K_S2.8B1.51 GiB1.41 GiB3.71 GiB0.01 GiB21±30%
phi-2Q3_K_L2.8B1.49 GiB1.41 GiB3.71 GiB0.01 GiB21±30%
Miril-Drone-2B-1IQ3_XS5.1B2.88 GiB0.04 GiB3.71 GiB0.01 GiB21±30%
Marco-Nano-InstructMoEI1-IQ1_S8.0B2.44 GiB0.49 GiB3.71 GiB0.01 GiB45±37%
MiniCPM-V-4Q6_K4.1B2.76 GiB0.14 GiB3.71 GiB0.01 GiB21±30%
deepseek-coder-5.7bmqa-baseQ3_K_L5.7B2.81 GiB0.07 GiB3.71 GiB0.01 GiB21±30%
NVIDIA-Nemotron-3-Nano-4B-BF16UD-IQ2_M4.0B2.14 GiB0.74 GiB3.70 GiB0.02 GiB21±30%
Qwen3-VL-8B-InstructUD-IQ1_M8.8B2.24 GiB0.63 GiB3.70 GiB0.02 GiB21±30%
GrammarCoder-7B-BaseI1-IQ2_M7.6B2.60 GiB0.25 GiB3.70 GiB0.02 GiB21±30%
Qwen3-VL-8B-ThinkingUD-IQ1_M8.8B2.23 GiB0.63 GiB3.70 GiB0.02 GiB21±30%
Qwen3-8BUD-IQ1_M8.2B2.23 GiB0.63 GiB3.70 GiB0.02 GiB21±30%
Teuken-7B-instruct-research-v0.4I1-IQ2_XXS7.5B2.72 GiB0.14 GiB3.70 GiB0.02 GiB21±30%
umt5-xxlQ3_K_M5.7B2.85 GiB0.00 GiB3.70 GiB0.02 GiB21±30%
Vikhr-Gemma-2B-instructQ8_02.6B2.59 GiB0.29 GiB3.70 GiB0.02 GiB21±30%
gemma-2-2b-it-abliteratedQ8_02.6B2.59 GiB0.29 GiB3.70 GiB0.02 GiB21±30%
gemma-2-2b-itQ8_02.6B2.59 GiB0.29 GiB3.70 GiB0.02 GiB21±30%
Gemmasutra-Mini-2B-v1Q8_02.6B2.59 GiB0.29 GiB3.70 GiB0.02 GiB21±30%
Phi-4-mini-instruct-abliteratedQ4_K_M3.8B2.32 GiB0.56 GiB3.69 GiB0.03 GiB21±30%
Phi-4-mini-reasoningQ4_K_M3.8B2.32 GiB0.56 GiB3.69 GiB0.03 GiB21±30%
Phi-4-mini-instructQ4_K_M3.8B2.32 GiB0.56 GiB3.69 GiB0.03 GiB21±30%
whisper-mediumF32764M2.85 GiB0.00 GiB3.69 GiB0.03 GiB21±30%
whisper-medium.enF32764M2.85 GiB0.00 GiB3.69 GiB0.03 GiB21±30%
granite-4.0-h-tinyMoEQ3_K_S6.9B2.89 GiB0.04 GiB3.69 GiB0.03 GiB67±37%
DeepSeek-R1-0528-Qwen3-8BUD-IQ1_M8.2B2.23 GiB0.63 GiB3.69 GiB0.03 GiB21±30%
granite-3.3-8b-instructUD-IQ2_XXS8.2B2.16 GiB0.70 GiB3.69 GiB0.03 GiB21±30%
DeepHat-V1-7B-Heretic-AbliteratedI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
ShizhenGPT-7B-VLI1-IQ2_M8.3B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
DeepHat-V1-7BIQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
HuatuoGPT-o1-7BI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
MathSmith-DS-Qwen-7B-LongCoTI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
AstraGPTCoder-7BI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Qwen2.5-Coder-7B-Instruct-Ghidra-v2I1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
EsDrac-v1-7BI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Hemlock-Apothecary-7B-GRPO-e3I1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
openhands-lm-7b-v0.1I1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Hemlock2-Coder-7B-GRPOI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
shellwhiz-7bI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Qwen2.5-Coder-7B-Instruct-abliteratedI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Qwen2.5-Coder-7B-Instruct-OBLITERATED-advancedI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Qwen-STEM-Specialist-7BI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
VulnLLM-R-7BI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Garnet-OCR-7B-0422I1-IQ2_M8.3B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
UwU-7B-InstructI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Video-R1-7BI1-IQ2_M8.3B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
HARC-Qwen2.5-7B-InstructI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
Qwen2.5-Coder-7B-AbliteratedI1-IQ2_M7.6B2.59 GiB0.25 GiB3.69 GiB0.03 GiB21±30%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Arc Pro A30M 4GB run?
802 of 2118 indexed open-weight models fit a Arc Pro A30M 4GB at 16,384 context with q4_0 KV cache, the largest being gemma-4-E2B-it at IQ3_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Arc Pro A30M 4GB actually have?
Its nameplate is 4 GB, but about 3.72 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Arc Pro A30M 4GB fast for local AI?
Its memory bandwidth is 112 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.