Skip to main content

Xiaomi: MiMo-V2.5-Pro Coding Benchmark

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

Context1,048,576tokens
Max Output131,072tokens
Inputmodality
Price$0.43/1M input

Try Xiaomi: MiMo-V2.5-Pro in Kilo Code

Experience this model with the most popular open source coding agent. Free to start, pay only for AI usage. Use in popular IDEs like VS Code, JetBrains, command line, or cloud agents.

5M+

Downloads

500+

models supported

Free

to Start

Access 500+ models including Xiaomi: MiMo-V2.5-Pro and many more in Kilo Code

Coding Performance

Coding benchmarks and performance metrics for development tasks

Kilo Bench

% Completion on Terminal Bench 2.0
47.6%
Cost per attempt (USD)
$4.92
Benchmark
Terminal Bench 2.0

Official Kilo eval results. Cost is averaged per complete benchmark attempt.

Security

Enkrypt AI red-team scores for Xiaomi: MiMo-V2.5-Pro. Each value is the share of successful attacks on a 0–100 scale — lower is safer.

Overall risk
27.9

Lower is safer

Safety

35.3/ 100

Composite safety risk from Enkrypt red-team evaluations. Lower is safer.

NIST

28.0/ 100

Average attack success across NIST-mapped tests: bias, harm, toxicity, CBRN, and insecure code.

OWASP

31.0/ 100

Weighted average of the same tests using OWASP Top 10 for LLMs 2025 risk rankings.

Attack categories

Percentage of successful attacks in each Enkrypt red-team category.

Jailbreak24.5/ 100

Share of jailbreak tests that bypassed the model's safety constraints.

Bias71.1/ 100

Share of tests that elicited biased responses.

Harmful content2.2/ 100

Share of tests that produced dangerous, violent, or hateful content.

Toxicity1.8/ 100

Share of tests that produced toxic or abusive content.

CBRN47.0/ 100

Share of tests that elicited chemical, biological, radiological, or nuclear assistance.

Insecure code17.3/ 100

Share of tests that produced vulnerable or malicious code.

Score
Value
Overall risk
27.9 / 100
Safety
35.3 / 100
NIST
28.0 / 100
OWASP
31.0 / 100
Jailbreak
24.5 / 100
Bias
71.1 / 100
Harmful content
2.2 / 100
Toxicity
1.8 / 100
CBRN
47.0 / 100
Insecure code
17.3 / 100

Security scores from the Enkrypt AI Safety Leaderboard · Last checked Oct 5, 2026

Real-World Usage

Real-world usage statistics from the Kilo Code community

Weekly Token Usage

Mode Rankings (Last Week)

Where this model ranks for each built-in mode

Code

Write, modify, and refactor code

No data

Ask

Get answers and explanations

No data

Debug

Diagnose and fix software issues

No data

Orchestrator

Coordinate tasks across multiple modes

#37

Real-world metrics from the Kilo Code Leaderboard

OpenClaw Benchmarks

PinchBench measures how Xiaomi: MiMo-V2.5-Pro performs on real OpenClaw agent tasks: multi-step execution, tool use, recovery, latency, and cost.

PinchBench run

Average score

87.5%

#10 of 50 official models

Average time

251m 19s

6 runs · per OpenClaw task

Average cost

$12.131

Per benchmark run

Category breakdown

Best verified PinchBench v2 run by OpenClaw task family.

Coding97.9% · 10/14 cleared
Csv Analysis94.9% · 3/26 cleared
Productivity94.3% · 4/8 cleared
Writing94.2% · 1/6 cleared

Top task results

Highest-scoring benchmark tasks from the same submission.

Analysis
Access Control Log Anomaly Detection
100.0%
Log Analysis
Apache Error Log - Identify Problematic Client IPs
100.0%
Productivity
Calendar Event Creation
100.0%
Coding
CI/CD Pipeline Debug
100.0%
Skills
Create Project Structure
100.0%
Coding
Dockerfile Optimization
100.0%

Autonomous task execution

Xiaomi: MiMo-V2.5-Pro shows strong average success across OpenClaw-style benchmark runs, useful for recurring research, browser, and file-based automations.

Tool use and recovery

PinchBench tasks stress multi-step planning, tool calls, and judge-verified completion rather than single prompt coding snippets.

Agent workflow fit

Its deliberate average runtime and premium run cost help set expectations for long-running agents and production workflows.

Agentic benchmarks from the PinchBench Leaderboard

Pricing

Cost per 1 million tokens

Input Tokens
$0.43
per 1M tokens
Output Tokens
$0.87
per 1M tokens

Example Cost

Analyzing a 10,000 line codebase (≈40k input tokens, 10k output tokens) costs approximately $0.0261

Coding Capabilities

Features and parameters relevant to coding tasks

Coding Features

Function Calling
Can call external functions/APIs
Tool Choice
Control over function selection
Structured Outputs
JSON schema validation
Reasoning Tokens
Extended thinking for complex problems

Pricing details from OpenRouter

Technical Details

Architecture and implementation specifications

Specifications
Model ID
xiaomi/mimo-v2.5-pro
Artificial Analysis Slug
MiMo-V2.5-Pro
Created
April 22, 2026
Tokenizer
Other
Input Modalities
Text
Context Window
1,048,576 tokens
Max Completion Tokens
131,072 tokens
Input Price
$0.43 per 1M tokens
Output Price
$0.87 per 1M tokens
Cache Read Price
$0.00 per 1M tokens
Content Moderation
Disabled

Ready to try Xiaomi: MiMo-V2.5-Pro?

Install Kilo Code and start using Xiaomi: MiMo-V2.5-Pro for your coding projects today. Choose from 500+ AI models with complete freedom.

  1. Install Kilo Code

    Get the extension from VS Code Marketplace, JetBrains Plugin Repository, or the CLI.

  2. Open the model selector

    Click the model name in the Kilo Code chat panel to open the selector.

  3. Choose your model

    Search or browse to find and select your preferred model.

  4. Start coding

    Use Code, Ask, Debug, or Plan mode — the model is ready immediately.