When it makes sense
- Personal or confidential data that has to stay in your own infrastructure under the GDPR
- Many requests a day, where the per-token prices of cloud APIs add up
- Workflows that need fast answers, for example in the service desk
The right model server for the job
We do not tie you to one tool. Which model server fits depends on how many people and workflows use it and on the hardware you have:
| Model server | Fits | Why |
|---|---|---|
| vLLM | Production servers with many users and workflows at the same time | Continuous batching and PagedAttention for high throughput, one model across several GPUs, an OpenAI-compatible API |
| llama.cpp | CPUs, smaller GPUs and machines at the edge | Runs quantised models in the GGUF format with little memory, with or without a graphics card |
| Ollama | Pilots, workstations and a quick start | Downloads and starts a model with one command, good for trying models before one goes into production |
All three offer an API in the format of OpenAI. Workflows and tools that work with a cloud model work with the local one as well.
What we set up
- The model server on your hardware or in an EU data centre: vLLM for production, llama.cpp or Ollama where resources are small
- Models chosen for the task: small models for classification and extraction, larger open models such as gpt-oss, Llama, Mistral or Qwen where more reasoning is needed
- n8n workflows that use the local model like a cloud API, with access control and logging
- Monitoring of load and response times
How we choose the hardware
We start from the tasks, not from the largest model. We look at how many requests a day there are, how many run at the same time and how fast the answers have to be. Then we pick model sizes, quantisation and hardware that fit. Our post How to run LLMs locally shows the model classes and hardware checks we use. What the difference in cost looks like in a real workflow is in our cost comparison of cloud and local LLMs.
Cloud and local together
Not every task has to run locally. Tasks with sensitive data run on your own model, others can use a cloud model. n8n decides per workflow, so you do not have to pick one for everything.
Ready to get started?
Tell us which process you want to improve. We look at it with you and suggest how to start.
