Local inference

Run models on your own machine through a managed local model server, so you can work with a model without sending your prompts to a hosted provider.

ℹ️Info

The local model server requires a Kilo account. Sign in to use it.

Set up a local model

Sign in, turn the server on, and import a model from a GGUF file from the Local Model Server settings. That section covers the full set of server and per-model options.

Pick it in chat

To use an imported model in a chat, open the model picker, then select Local Model Server. Your imported models appear there. Choose one and use it like any other model.