Before you connect
- The model file, quantization, context length, and offload settings all affect memory use and latency.
- A model that works for chat may not reliably follow tool schemas or complete multi-step coding tasks.
- Local inference removes hosted token charges, but hardware, electricity, storage, and operator time still have costs.
- Keep the local server bound to your machine unless you intentionally secure and operate it for remote access.