Skip to main content

Best AI Models in Kilo

Compare live model rankings from real Kilo Code usage. See which models developers choose for coding, planning, debugging, and agent workflows across 500+ hosted options.

Top Coding Models in Kilo This Week

Our picks based on real-world testing • View usage stats

GPT-5.6 Sol

#1

A flagship model for long-horizon agents and complex problem solving

KiloBench completion76.2%

Claude Opus 4.8

#2

Built for complex planning, orchestration, and long-running professional tasks

Coding Index74.3

Grok 4.6

#3

Frontier coding and STEM performance tuned for long-running agents

KiloBench completion73.0%

Step 3.7 Flash

#4

A fast, open-weight model designed for efficient agentic workflows

context window256k

Kilo Benchmark

Cost vs performance across the best coding models

Bubble size indicates Kilo usage over the last 7 days

Top 10 Most Capable Models

RankModelCompletionCost per attempt
179.3%$107.29
276.2%$87.41
376.2%$91.41
475.3%$112.27
574.2%$72.63
673.0%$33.83
772.8%$48.38
871.5%$113.54
971.0%$87.52
1070.8%$27.29

Official Kilo eval results on Terminal Bench 2.0. Cost and token usage are averaged per complete benchmark attempt.

Top Models by Mode

See which models lead in Code, Plan, Debug, Ask, and Orchestrator

Code

RankModelUsage
01
step-3.7-flash
25.7%
0225.6%
0315.9%
04
longcat-2.0-free
11.5%
054.2%
06
deepseek-v4-flash-0731
2.4%
071.4%
08
glm-5.3-flash
1.2%
091.2%
10
deepseek-v4-flash-latest
0.8%

Plan

RankModelUsage
0117.8%
0217.2%
03
step-3.7-flash
16.0%
04
longcat-2.0-free
13.7%
054.9%
063.1%
07
deepseek-v4-flash-0731
2.6%
081.9%
091.8%
10
deepseek-v4-pro-0813
1.7%

Ask

RankModelUsage
01
step-3.7-flash
22.9%
0217.7%
0316.9%
04
longcat-2.0-free
12.7%
05
deepseek-v4-flash-0731
3.3%
063.3%
071.9%
081.6%
091.6%
101.4%

Debug

RankModelUsage
01
step-3.7-flash
25.3%
0219.9%
0313.7%
04
longcat-2.0-free
12.5%
0511.2%
06
glm-5.3-flash
3.6%
072.1%
081.7%
09
free
1.3%
100.8%

Review

RankModelUsage
0120.2%
02
qwen3.8-flash
15.8%
03
deepseek-v4-pro-0813
11.7%
04
deepseek-v4-flash-latest
10.5%
05
longcat-2.0-free
6.3%
064.0%
073.9%
083.5%
093.3%
103.1%

kiloclaw

RankModelUsage
0132.2%
02
deepseek-v4-flash-0731
18.8%
03
step-3.7-flash
10.3%
04
longcat-2.0-free
8.1%
056.5%
064.7%
07
hf
4.5%
083.8%
091.7%
101.1%

Top Models Today

Most-used models across Kilo Code in the last 24 hours

RankModelUsage
01159.5B
02
step-3.7-flash
154.4B
0385.1B
0432.5B
0513.0B
06
free
12.0B
07
deepseek-v4-flash-0731
11.7B
0810.7B
0910.5B
10
glm-5.3-flash
8.4B

Daily Top Models

Token usage by model over time, stacked daily

AI Provider Token RaceWeeklySee which AI labs are gaining groundReplay weekly token usage across leading proprietary and open-weight labs.Provider usage over timeWatch the race

All Models

Browse and compare all available AI coding models

DeepSeek V4 Flash 0423

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Coding index52.0
$0.14/1M
View

Nemotron 3 Ultra (free)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Code mode rank#10
$0.00/1M
View

DeepSeek V4 Pro 0423

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Kilo Bench completion44.0%
$1.60/1M$15.91/attempt
View

GLM 5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

Code mode rank#23
$0.60/1M
View

Claude Sonnet 5

Recommended

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Kilo Bench completion59.6%
$2.00/1M$36.19/attempt
View

Claude Sonnet 4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

Coding index52.1
$3.00/1M
View

Claude Opus 4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

Code mode rank#27
$5.00/1M
View

Hy3 (free)

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

Kilo Bench completion47.6%
$0.00/1M$0.00/attempt
View

GLM 5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Kilo Bench completion53.0%
$1.40/1M$26.21/attempt
View

GPT-5.6 Terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Code mode rank#31
$2.00/1M
View

kilo-auto/efficient

Model details and benchmarks are being added.

Kilo Bench completion46.7%
$19.60/attempt
View

Claude Opus 4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Kilo Bench completion67.6%
$5.00/1M$85.19/attempt
View

Claude Opus 5

Recommended

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Kilo Bench completion71.5%
$5.00/1M$113.54/attempt
View

Ling-3.0-flash (free)

Recommended

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

Code mode rank#51
$0.00/1M
View

Qwen3.7 Plus (20% off)

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Code mode rank#54
$0.32/1M
View

MiniMax M3

Recommended

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Kilo Bench completion47.6%
$0.30/1M$10.35/attempt
View

Claude Sonnet 4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

Kilo Bench completion55.1%
$3.00/1M$53.37/attempt
View

Ling-2.6-1T (free)

Recommended

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

Code mode rank#60
$0.00/1M
View

Gemini 2.5 Flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

Code mode rank#72
$0.30/1M
View

Kimi K3

Recommended

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Kilo Bench completion72.8%
$3.00/1M$48.38/attempt
View

Laguna S 2.1 (free)

Recommended

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Code mode rank#76
$0.00/1M
View

Gemini 3.1 Pro Preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

Code mode rank#78
$1.00/1M
View

MiMo-V2.5-Pro

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

Kilo Bench completion47.6%
$0.43/1M$4.92/attempt
View

GPT-5.2

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...

Code mode rank#86
$1.75/1M
View

GPT-5.3-Codex

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results...

Code mode rank#86
$1.75/1M
View

GPT-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Kilo Bench completion74.2%
$5.00/1M$72.63/attempt
View

Grok Code Fast 1

Grok Code Fast 1 is a speedy and economical reasoning model that excels at agentic coding. With reasoning traces visible in the response, developers can steer Grok Code for high-quality...

Code mode rank#93
$0.20/1M
View

Qwen3.6 Plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

Coding index54.5
$0.33/1M
View

MiniMax M2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

Code mode rank#100
$0.30/1M
View

GPT-5.4

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

Code mode rank#100
$2.50/1M
View

Claude Fable 5 ($$$$)

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Kilo Bench completion71.0%
$10.00/1M$87.52/attempt
View

Claude Fable 5.1 ($$$$)

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Kilo Bench completion76.2%
$10.00/1M$91.41/attempt
View

Claude Haiku 4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

AvailabilityIn Kilo
$1.00/1M
View

Claude Opus 4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

AvailabilityIn Kilo
$5.00/1M
View

Claude Opus 4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

Kilo Bench completion70.1%
$5.00/1M$100.51/attempt
View

baidu/cobuddy:free

Model details and benchmarks are being added.

AvailabilityIn Kilo
View

Command R+ (08-2024)

command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as compared to the previous Command R+ version, while keeping the hardware footprint...

AvailabilityIn Kilo
$2.50/1M
View

Command R7B (12-2024)

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning...

AvailabilityIn Kilo
$0.04/1M
View

DeepSeek V3

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations...

AvailabilityIn Kilo
$0.32/1M
View

DeepSeek V3.1 Terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

Coding index43.5
$0.27/1M
View

DeepSeek V3.2 Exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

AvailabilityIn Kilo
$0.27/1M
View

R1 0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

AvailabilityIn Kilo
$0.70/1M
View
D

Dots3-Note Preview (free)

Recommended

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...

AvailabilityIn Kilo
$0.00/1M
View

Gemini 2.5 Pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

AvailabilityIn Kilo
$1.25/1M
View

Gemini 3 Flash Preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

AvailabilityIn Kilo
$0.25/1M
View

Gemini 3 Pro Preview

Gemini 3 Pro is Google’s flagship frontier model for high-precision multimodal reasoning, combining strong performance across text, image, video, audio, and code with a 1M-token context window. Reasoning Details must be preserved when using multi-turn tool calling, see our docs here: https://openrouter.ai/docs/use-cases/reasoning-tokens#preserving-reasoning-blocks. It delivers state-of-the-art benchmark results in general reasoning, STEM problem solving, factual QA, and multimodal understanding, including leading scores on LMArena, GPQA Diamond, MathArena Apex, MMMU-Pro, and Video-MMMU. Interactions emphasize depth and interpretability: the model is designed to infer intent with minimal prompting and produce direct, insight-focused responses. Built for advanced development and agentic workflows, Gemini 3 Pro provides robust tool-calling, long-horizon planning stability, and strong zero-shot generation for complex UI, visualization, and coding tasks. It excels at agentic coding (SWE-Bench Verified, Terminal-Bench 2.0), multimodal analysis, and structured long-form tasks such as research synthesis, planning, and interactive learning experiences. Suitable applications include autonomous agents, coding assistants, multimodal analytics, scientific reasoning, and high-context information processing.

AvailabilityIn Kilo
$2.00/1M
View

Gemini 3.5 Flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Kilo Bench completion64.7%
$0.75/1M$104.49/attempt
View

Gemini 3.6 Flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Kilo Bench completion47.9%
$0.38/1M$80.31/attempt
View

Gemini 3.8 Flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Kilo Bench completion75.3%
$0.75/1M$112.27/attempt
View

Gemma 2 27B

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-suited for a variety of...

AvailabilityIn Kilo
$0.65/1M
View

Gemma 3 12B

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

AvailabilityIn Kilo
$0.05/1M
View

Gemma 3 27B

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

AvailabilityIn Kilo
$0.08/1M
View

Gemma 3 4B

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

AvailabilityIn Kilo
$0.05/1M
View

Gemma 4 26B A4B

Recommended

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

AvailabilityIn Kilo
$0.04/1M
View

Gemma 4 26B A4B (free)

Recommended

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

AvailabilityIn Kilo
$0.00/1M
View

Gemma 4 31B

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

AvailabilityIn Kilo
$0.09/1M
View
I

Granite 4.0 Micro

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long...

AvailabilityIn Kilo
$0.02/1M
View

Ling-2.6-1T (retires Aug 24)

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

Kilo Bench completion28.1%
$0.30/1M$30.82/attempt
View

Ling-2.6-flash (free)

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....

AvailabilityIn Kilo
$0.00/1M
View

Ring-2.6-1T (free)

Recommended

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...

AvailabilityIn Kilo
$0.00/1M
View

KAT-Coder-Pro V1

KAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KAT-Coder series. Designed specifically for agentic coding tasks, it excels in real-world software engineering scenarios, achieving 73.4% solve rate on the SWE-Bench Verified benchmark. The model has been optimized for tool-use capability, multi-turn interaction, instruction following, generalization, and comprehensive capabilities through a multi-stage training process, including mid-training, supervised fine-tuning (SFT), reinforcement fine-tuning (RFT), and scalable agentic RL.

AvailabilityIn Kilo
$0.21/1M
View

KAT-Coder-Pro V2.5

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

Kilo Bench completion50.3%
$0.74/1M$36.16/attempt
View

LongCat 2.0

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

Coding index45.3
$0.75/1M
View

Llama 3.1 8B Instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

AvailabilityIn Kilo
$0.02/1M
View

Llama 3.2 1B Instruct

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...

AvailabilityIn Kilo
$0.03/1M
View

Llama 3.2 3B Instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

AvailabilityIn Kilo
$0.05/1M
View

Muse Spark 1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Kilo Bench completion59.8%
$1.25/1M$30.15/attempt
View

Muse Spark 1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

Kilo Bench completion51.5%
$1.25/1M$44.69/attempt
View
M

Phi 4

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...

AvailabilityIn Kilo
$0.07/1M
View

MiniMax M2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

AvailabilityIn Kilo
$0.30/1M
View

MiniMax M2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

AvailabilityIn Kilo
$0.30/1M
View

MiniMax M2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

AvailabilityIn Kilo
$0.30/1M
View

MiniMax M2.7 (free)

Recommended

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

AvailabilityIn Kilo
$0.00/1M
View

MiniMax M3 (free)

Recommended

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

AvailabilityIn Kilo
$0.00/1M
View

Devstral 2 2512

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...

AvailabilityIn Kilo
$0.40/1M
View

Mistral Medium 3.5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Coding index46.9
$1.50/1M
View

Mistral Small 3

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed...

AvailabilityIn Kilo
$0.05/1M
View

Kimi K2 Thinking

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillion-parameter Mixture-of-Experts (MoE) architecture introduced in...

AvailabilityIn Kilo
$0.60/1M
View

Kimi K2.5

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed...

AvailabilityIn Kilo
$0.60/1M
View

Kimi K2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

Kilo Bench completion54.4%
$0.80/1M$24.84/attempt
View

Kimi K2.7 Code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

Kilo Bench completion60.7%
$0.95/1M$32.94/attempt
View

Nex-N2-Pro (free)

Recommended

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

AvailabilityIn Kilo
$0.00/1M
View

Nemotron 3 Super (free)

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Kilo Bench completion15.5%
$0.00/1M$0.00/attempt
View

Nemotron 3 Ultra

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Kilo Bench completion19.1%
$0.50/1M$101.82/attempt
View

Nemotron 3.5 Lightning (free)

Model details and benchmarks are being added.

AvailabilityIn Kilo
View

GPT-3.5 Turbo

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

AvailabilityIn Kilo
$0.50/1M
View

GPT-4 ($$$$)

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning...

AvailabilityIn Kilo
$30.00/1M
View

GPT-4 Turbo ($$$$)

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

AvailabilityIn Kilo
$10.00/1M
View

GPT-4.1

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and...

AvailabilityIn Kilo
$2.00/1M
View

GPT-4.1 Mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard...

AvailabilityIn Kilo
$0.40/1M
View

GPT-4.1 Nano

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million...

AvailabilityIn Kilo
$0.10/1M
View

GPT-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as...

AvailabilityIn Kilo
$2.50/1M
View

GPT-4o (2024-08-06)

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/). GPT-4o ("o" for "omni") is...

AvailabilityIn Kilo
$2.50/1M
View

GPT-4o-mini

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

AvailabilityIn Kilo
$0.15/1M
View

GPT-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

AvailabilityIn Kilo
$1.25/1M
View

GPT-5 Mini

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

AvailabilityIn Kilo
$0.25/1M
View

GPT-5 Nano

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

AvailabilityIn Kilo
$0.05/1M
View

GPT-5.1

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

Coding index49.4
$1.25/1M
View

GPT-5.1-Codex

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....

AvailabilityIn Kilo
$1.25/1M
View

GPT-5.2-Codex

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks....

AvailabilityIn Kilo
$1.75/1M
View

GPT-5.6 Luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

AvailabilityIn Kilo
$0.20/1M
View

GPT-5.6 Sol

Recommended

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Kilo Bench completion76.2%
$4.00/1M$87.41/attempt
View

GPT-6 Astra ($$$$)

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

Kilo Bench completion79.3%
$10.00/1M$107.29/attempt
View

gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

AvailabilityIn Kilo
$0.03/1M
View

gpt-oss-20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

AvailabilityIn Kilo
$0.02/1M
View

gpt-oss-safeguard-20b

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model offers lower latency for safety tasks like content classification, LLM filtering, and trust...

AvailabilityIn Kilo
$0.07/1M
View

o1 ($$$$)

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason...

AvailabilityIn Kilo
$15.00/1M
View

o3

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following....

AvailabilityIn Kilo
$2.00/1M
View

o3 Mini

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding. This model supports the `reasoning_effort` parameter, which can be set to...

AvailabilityIn Kilo
$1.10/1M
View

o4 Mini

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning...

AvailabilityIn Kilo
$1.10/1M
View

Laguna S 2.1

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Kilo Bench completion31.0%
$0.10/1M$1.85/attempt
View

Laguna XS 2.1

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Kilo Bench completion26.7%
$0.10/1M$12.03/attempt
View

Qwen3 235B A22B

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

AvailabilityIn Kilo
$0.46/1M
View

Qwen3 30B A3B

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...

AvailabilityIn Kilo
$0.13/1M
View

Qwen3 8B

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...

AvailabilityIn Kilo
$0.12/1M
View

Qwen3 Coder 480B A35B

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

AvailabilityIn Kilo
$0.97/1M
View

Qwen3 Coder Next

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

AvailabilityIn Kilo
$0.30/1M
View

Qwen3 Coder Plus

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

AvailabilityIn Kilo
$0.65/1M
View

Qwen3 VL 235B A22B Instruct

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...

AvailabilityIn Kilo
$0.26/1M
View

Qwen3.5 397B A17B

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

AvailabilityIn Kilo
$0.39/1M
View

Qwen3.5-122B-A10B

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

AvailabilityIn Kilo
$0.26/1M
View

Qwen3.5-27B

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

AvailabilityIn Kilo
$0.20/1M
View

Qwen3.5-35B-A3B

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

AvailabilityIn Kilo
$0.16/1M
View

Qwen3.7 Max (50% off)

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Kilo Bench completion54.6%
$1.25/1M$20.65/attempt
View

Qwen3.8 Max

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...

Kilo Bench completion54.2%
$2.00/1M$32.71/attempt
View

Grok 4.3

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

Coding index42.2
$1.25/1M
View

Grok 4.5

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

Kilo Bench completion70.8%
$2.00/1M$27.29/attempt
View

Grok 4.6

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Kilo Bench completion73.0%
$2.00/1M$33.83/attempt
View

Grok Build 0.1

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

Kilo Bench completion50.6%
$1.00/1M$30.70/attempt
View

Step 3.5 Flash

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token....

AvailabilityIn Kilo
$0.10/1M
View

Hunyuan A13B Instruct

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...

AvailabilityIn Kilo
$0.14/1M
View

Hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

Coding index58.8
$0.08/1M
View

Hy3 preview (free)

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...

AvailabilityIn Kilo
$0.00/1M
View

Inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Kilo Bench completion43.6%
$0.95/1M$35.69/attempt
View
M

WizardLM-2 8x22B

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is...

AvailabilityIn Kilo
$0.62/1M
View

MiMo-V2-Flash

Model details and benchmarks are being added.

AvailabilityIn Kilo
View

MiMo-V2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Coding index56.8
$0.14/1M
View

GLM 4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

AvailabilityIn Kilo
$0.60/1M
View

GLM 4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

Coding index45.8
$0.43/1M
View

GLM 4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

AvailabilityIn Kilo
$0.40/1M
View

GLM 4.7 Flash (retires Sep 10)

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

AvailabilityIn Kilo
$0.06/1M
View

GLM 5 Turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

AvailabilityIn Kilo
$1.20/1M
View

GLM 5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

Kilo Bench completion49.4%
$1.40/1M$23.98/attempt
View

GLM 5.3

Recommended

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

AvailabilityIn Kilo
$1.40/1M
View

Methodology

Models earn their rank from the developers using them, not from a spec sheet.

Synthetic benchmarks measure one capability at one moment. This leaderboard measures what developers come back to: real coding work, long planning sessions, debugging, review, and agentic tasks across 500+ models.

Modes
Code, Plan, Ask, Debug, Review
Catalog
500+ models

Read the leaderboard like a pro

  1. Start with usage

    The rankings reflect real token usage by Kilo Code developers, not synthetic benchmarks.

  2. Filter by mode

    Use the Top Models by Mode section to see which models lead in Code, Plan, Debug, Ask, and Orchestrator.

  3. Open the model page

    Each model links to a dedicated page with benchmark scores, pricing, context length, and speed data.

  4. Switch in Kilo Code

    All 500+ models are available in Kilo Code. Switch from the model selector at any time.

AI model FAQ

What is the Kilo Code AI Model Leaderboard?

The Kilo Code leaderboard shows live rankings of AI coding models based on real token usage by 5M+ developers. Rankings reflect genuine developer preference and update every 5 minutes.

How are AI coding models ranked?

Models are ranked by total token usage from Kilo Code developers. The ranking reflects real-world developer preference and can be filtered by modes such as Code, Plan, Debug, Ask, and Orchestrator.

Which AI model is best for coding?

The current top-ranked model is MiniMax: MiniMax M3, based on real developer usage. The best AI model for your workflow can vary by coding, planning, debugging, and agent tasks, so compare the live rankings by mode before choosing.

How often does the leaderboard update?

The leaderboard updates every 5 minutes with fresh usage data from real Kilo Code developers.