8 GB
No registry-backed coding recommendation is currently published for this tier.
Choose a local coding model by hardware tier, not by model family hype. Each pick lists the exact artifact, runtime, quantization or precision, memory assumption, usable context, and whether the hardware fit is tested, official, or estimated.
No registry-backed coding recommendation is currently published for this tier.
No registry-backed coding recommendation is currently published for this tier.
Mistral documents single-RTX-4090 or 32 GB Mac deployment; usable context still depends on quantization and memory headroom.
Use a smaller context if latency or memory pressure is high.
Use hosted or server deployment for practical coding-agent latency.
The full checkpoint is a server-class deployment; use a hosted route or reviewed quantization for smaller hardware.
The official model card recommends vLLM, SGLang, or TokenSpeed; this is not a consumer-hardware recommendation.
Confirm access, license, precision, and memory requirements for the exact artifact before self-hosting.
Poolside documents running quantized builds on a single NVIDIA DGX Spark; confirm exact memory requirements per quantization before self-hosting.
NVIDIA GPUs usually give the best latency for vLLM or SGLang. Apple Silicon unified memory can be practical for quantized models when the OS and IDE have headroom. CPU or heavy offload is best treated as experimentation unless latency is acceptable.
Code chat can work with smaller context and slower tokens. Tool-using agents need more context, better instruction following, lower latency, and memory headroom for repeated file reads and command output.
Server-class open-weight artifacts are usually better consumed through Kilo hosted models or your own on-prem endpoint. Local does not mean free; it moves cost to hardware, power, and operations.
No. Local fit depends on the exact artifact, precision or quantization, runtime, context length, memory headroom, and workload. This guide labels estimates and avoids deriving fit from active parameters alone.
Local inference avoids hosted token bills, but you still pay hardware, power, setup, maintenance, and latency costs.
Evidence-backed coding ranking and use-case picks.
Live zero-price hosted models and tested free winners.
Chronological verified release tracker.
Definitions, licensing, and procurement checklist.
Production data from 10,643 AI code reviews across 13 models.
Local guidance covers 7 curated artifacts. Last verified 2026-07-30.