Skip to main content
Local models matched to your machine

Best Local LLM for Coding by Hardware

A local model has to fit in memory. A 7B model runs on modest hardware, a 30B model needs a serious GPU, and an 80B model is beyond most laptops. Start with your VRAM or unified memory, then choose the strongest coding model your machine can run with useful context and enough headroom for your tools.

Quick answer

8 GB VRAM

Best overall fit: Falcon H1R 7B at Q4.

16 GB VRAM

Best agent-oriented fit: Gemma 4 12B at Q4.

24 GB VRAM

Best overall local coding model: Qwen 3.6 27B.

48 GB+ VRAM

Best high-end local option: Qwen3-Coder-Next.

Hardware tier

Best Local Coding Models for 8 GB VRAM

Use a 7B to 9B model at Q4 and leave memory for the KV cache and coding tools.

Best overall fit: Falcon H1R 7B at Q4.

Falcon H1R 7B

About 4.6 GB at Q4; roughly 32K context on an 8 GB GPU

A compact hybrid reasoning model that genuinely fits an 8 GB card and performs well on reasoning and math tasks.

Caveat: Treat it as a reasoning assistant rather than a long-horizon coding agent; every layer still carries KV-cache cost.

Qwen 3.5 9B

8 GB GPU at Q4 with a constrained context window

Strong reasoning for its size when the prompt supplies enough repository context. It is useful for code explanation, small edits, and focused generation.

Caveat: Its limited stored knowledge becomes visible on unfamiliar frameworks.

Hardware tier

Best Local Coding Models for 16 GB VRAM

This tier can run capable 9B to 12B models with useful coding context at Q4.

Best agent-oriented fit: Gemma 4 12B at Q4.

Gemma 4 12B

12–16 GB VRAM; up to about 128K context on 16 GB

Native tool calling and reliable output formatting make it a practical small model for coding agents even when a similarly sized model scores higher on raw code generation.

Qwen 3.5 9B

16 GB VRAM can hold the model with substantially more context headroom

A strong choice for precise, bounded tasks when you value reasoning quality over broad framework recall.

Falcon H1R 7B

16 GB supports its full published context more comfortably

The extra memory removes the tightest context constraint, but it does not turn the model into a reliable autonomous agent.

Hardware tier

Best Local Coding Models for 24 GB VRAM

A single 3090, 4090, or 5090 reaches the local-coding sweet spot around 24B to 35B.

Best overall local coding model: Qwen 3.6 27B.

Qwen 3.6 27B

Q4 with roughly 64K context on a 24 GB card

The strongest general recommendation for a single 24 GB GPU: capable enough for real offline coding work without requiring multi-GPU hardware.

Devstral Small 2

24 GB at about 32K context

Designed for structured instructions and tool use, with an Apache 2.0 license and strong published SWE-bench performance.

Caveat: Its dense architecture makes long context expensive; the weights fit, but the full KV cache does not fit a consumer card.

Qwen 3.6 35B-A3B

24 GB natively, or 12–16 GB VRAM plus 32 GB system RAM with expert offload

A sparse option for machines that can trade speed for system-memory offload. Only a small share of parameters is active per token.

Caveat: The 27B dense model remains the stronger performer, and PCIe offload reduces speed.

Cohere North Mini Code 1.0

24 GB at up to about 128K context

A 30B MoE model with 3B active parameters, built for repository work and agentic coding rather than general chat.

Caveat: Watch for verbose output that increases generation time and context use.

NVIDIA Nemotron Cascade 2 30B-A3B

24 GB through its full 262K context window

The long-context specialist in this group. Its Mamba-2/MoE-heavy architecture keeps KV-cache growth much lower than a dense model.

Hardware tier

Best Local Coding Models for 48 GB+ VRAM

Two 24 GB GPUs make larger agent-focused models practical, though context still needs headroom.

Best high-end local option: Qwen3-Coder-Next.

Qwen3-Coder-Next

Two 3090s at IQ4_XS for about 128K context; 64 GB VRAM for the full 262K window

An approximately 80B agent-focused coding model trained across coding scaffolds for multi-step tool interactions and long-context repository work.

Caveat: This is workstation hardware: leave room for context, the runtime, and the rest of the coding environment.

Unified memory

Best Local Models for Apple Silicon RAM Tiers

Apple Silicon shares memory between the CPU and GPU. Do not allocate the entire headline RAM figure to model weights: macOS, the runtime, Kilo, your IDE, and the KV cache all need room.

MemoryModels to considerPractical guidance
16 GB unified memoryQwen 3.5 9B, Gemma 4 12B, or Falcon H1R 7BUse Q4 and a conservative context. Gemma 4 12B is the best fit when tool calling matters.
32 GB unified memoryQwen 3.6 27B, Qwen 3.6 35B-A3B, North Mini Code, or Devstral Small 2Qwen 3.6 27B is the default recommendation. Dense Devstral needs a shorter context than the sparse alternatives.
48 GB unified memoryQwen 3.6 27B with its full context targetThe additional headroom is most useful for KV cache and the rest of your development environment.
64 GB+ unified memoryQwen3-Coder-Next at Q4_K_M or preferably IQ4_XSThe larger quantization leaves little headroom at 64 GB, so the smaller artifact is the safer operational choice.

Best Small Coding Models

The best small coding model depends on whether you need tool use, reasoning, or the smallest possible memory footprint. Small models are most reliable for autocomplete, code explanation, commit messages, focused edits, and single-function generation.

Gemma 4 12B

12–16 GB VRAM; up to about 128K context on 16 GB

Native tool calling and reliable output formatting make it a practical small model for coding agents even when a similarly sized model scores higher on raw code generation.

Falcon H1R 7B

About 4.6 GB at Q4; roughly 32K context on an 8 GB GPU

A compact hybrid reasoning model that genuinely fits an 8 GB card and performs well on reasoning and math tasks.

Caveat: Treat it as a reasoning assistant rather than a long-horizon coding agent; every layer still carries KV-cache cost.

Qwen 3.5 9B

8 GB GPU at Q4 with a constrained context window

Strong reasoning for its size when the prompt supplies enough repository context. It is useful for code explanation, small edits, and focused generation.

Caveat: Its limited stored knowledge becomes visible on unfamiliar frameworks.

Connect Kilo

Set Up Ollama or LM Studio

The runtime downloads and serves the model. Kilo supplies the coding-agent workflow, repository context, file edits, and terminal tools.

Ollama setup

Use Ollama for scriptable, terminal-first model management and a lightweight local background service. Pull a quantization that fits your tier, confirm the context setting, then select Ollama as the provider in Kilo.

LM Studio setup

Use LM Studio for a visual model browser, quantization controls, memory settings, and a local OpenAI-compatible server. Download a fitting model, start the server, and point Kilo at the local endpoint.

Know the privacy boundary

Prompts sent to your local model stay on the endpoint you configure. That does not automatically cover sign-in, updates, remote MCP servers, telemetry choices, browser tools, Cloud Agents, or hosted review services. Audit every enabled integration if you need an isolated workflow.

Local, hosted, or hybrid

Local models work well for bounded tasks, sensitive repositories, and predictable workloads. Hosted models often provide better speed, context, and tool reliability for demanding agents. Kilo lets you switch between both approaches per task.

Not Enough Hardware?

You do not need a workstation to use a capable coding model. Start with the best free models for coding or use the Kilo Gateway as a managed cloud fallback when a model will not fit locally.

Local Model FAQ

What is the best local model for coding?

The answer depends on memory. Falcon H1R 7B fits an 8 GB GPU, Gemma 4 12B fits 12–16 GB, Qwen 3.6 27B is the strongest overall choice for 24 GB, and Qwen3-Coder-Next is the high-end choice for 48 GB or more.

What is the best small coding model?

Gemma 4 12B is the strongest small-model choice for tool-calling workflows with 12–16 GB available. Falcon H1R 7B is the safer fit for an 8 GB GPU, while Qwen 3.5 9B is a strong reasoning option for bounded coding tasks.

Does a local model keep all Kilo data offline?

Local inference keeps prompts sent to that model on the endpoint you configure. Sign-in, updates, remote MCP servers, telemetry choices, Cloud Agents, hosted code review, and other integrations have separate network paths.

Is local inference free?

It avoids hosted token charges, but you still pay for hardware, power, storage, setup, maintenance, and the developer time lost when inference is slow.

Hardware figures assume quantized models and are practical estimates, not guarantees. Available context varies with runtime, KV-cache precision, operating-system use, and the rest of your workload. Last reviewed .