Skip to main content
Engineering

How We Made Kilo Faster

50x faster recall search, 72% less CLI startup time, and faster agent sessions.

Published

Reading time

8 minutes

Marius Wichtner, one of Kilo's engineers, ships code constantly. He also built two features that make that pace possible: Agent Manager, which runs several agent sessions in parallel in isolated workspaces, and Kilo Swarm, which lets those agents share updates instead of quietly duplicating work.

Using his own tools, Marius opened roughly 350 pull requests over about seven weeks. A meaningful share were about making Kilo faster. We wanted to know exactly how much faster, based not on memory but on the measurements recorded when each change shipped.

How we checked

This didn't start as a planned report. It started as a Slack conversation:

#team-core-agents
Marius Wichtner9:51 AM

@Job Rietbergen (CET) maybe we also need a blog post about it? I can pull some numbers.

Job Rietbergen (CET)9:23 AM

Sure, pass me some data and I can write it up.

Marius Wichtner9:23 AM

@Kilo do you have access to the kilocode repository now?

KiloAGENT9:24 AM

Yes, I have access to Kilo-Org/kilocode. I can analyze code, answer questions about it, or make changes via a Cloud Agent session. What do you need?

Marius Wichtner9:27 AM

@Kilo go through all the PRs I authored over the last few weeks related to performance: tool rendering, session rendering in Agent Manager and VS Code, and streaming performance. Use 5-6 subagents to check the reported numbers, then report back with a table containing only the exact measurements from the PR descriptions.

KiloAGENT9:34 AM

Analysis complete. Full scan of ~350 PRs by marius-kilocode (Aug 1–Sep 24), ~45 perf-related, ~35 with exact numbers in descriptions. Headline results: [table below]

Marius Wichtner9:35 AM

Not bad, huh? 😄

Recreated from a real conversation.

Kilo scanned all 350 pull requests and reported back only the numbers that were actually written down at the time — nothing estimated, nothing rounded up. About 45 of them were performance-related. Roughly 35 of those had an exact before-and-after measurement attached. Here are the clearest ones.

Measured improvements

Nine before-and-after measurements reported in the linked pull requests

CLI startup

kilo --version

3.6x

Before
1,482ms
After
413-421ms

-72% elapsed time

Windows launch

Warm-cache --version

6.3x

Before
3,925ms
After
625ms

-84% elapsed time

Recall search

Query on 14GB index

50x

Before
30.2s
After
0.6s

50x faster

Cold recall search

Cold query, 63GB index

7.6x

Before
7.1s
After
0.93s

7.6x faster

Long session load

Opening a 1,693-message session

2.4x

Before
580.9ms
After
238.0ms

-59% elapsed time

Terminal connect

Ready to connected on restart

48x

Before
4,097ms
After
85ms

-98% elapsed time

Worktree create

Warm-cache claim

8.0x

Before
2,167ms
After
272ms

-87% elapsed time

Pre-LLM snapshot

First track after prepare

13.8x

Before
1,269ms
After
92ms

-93% elapsed time

Worktree delete

Cold-cache backend

10x

Before
1,110ms
After
111ms

-90% elapsed time

Bars are normalized within each benchmark so improvements across millisecond and second-scale workloads can be compared without hiding smaller paths.

The impact

What this means in practice

Turned into plain terms, the table above breaks down into four everyday moments:

Starting Kilo

Running a basic command used to take close to a second and a half. It now takes under half a second — the difference between a pause you notice and one you don't. A separate, Windows-only slowdown (an extra check the CLI was re-running on every single launch) got fixed the same way: do it once, remember the answer.

CLI startup

kilo --version

-72%

elapsed time

Sources: #12682, #13412, #13555

Searching your own code

Kilo's local search used to be slow enough that a 30-second query was normal on a large codebase. That same search now finishes in about half a second. On an even bigger codebase, a second fix shaved further time off the slower, uncached version of that same search.

Starting a new agent session

Every time you kick off a new agent task, Kilo sets up an isolated workspace for it behind the scenes. That setup used to take a couple of seconds end to end. Across a handful of fixes to different steps in that process, it now takes a fraction of a second — closer to instant than to a loading screen.

Where first-prompt time went

Create click to LLM request

Snapshot work Other startup work

Create click to LLM request: 5.5-5.9s before, 2.0-2.2s after. Snapshot work fell from 4.5-4.8s to 1.1-1.3s. Three runs per build in isolated VS Code with a warm worktree pool and snapshots enabled. Bars use range midpoints; gray segments are the derived remainder, not separately measured phases.

Source: #14115

Working with an agent in the chat view

This is the one area that isn't really about speed at all — it's about glitches. The chat view used to visibly jump or flicker while an agent was mid-response: the scroll position would snap around, finished tool calls would redraw themselves, and long sessions would gradually slow down the longer they stayed open. None of that was a single big number to fix; it was roughly a dozen small, specific bugs, each one caught, measured, and closed. The net effect is that the chat view just holds still now, even during long or busy sessions.

A few of the ~45 performance fixes only affect the team building Kilo rather than anyone using it — faster test runs and faster builds, mainly. Those don't change what you experience day to day, so they're not broken out above, but they're part of why the team can keep shipping fixes like these at this pace.

Feel the difference yourself.

Install Kilo in VS Code, JetBrains, or your terminal and put the faster workflow to work.

Explore more engineering stories