Model fit
Choosing a model for private, local workloads
What changes when a request stays on hardware you control instead of a cloud provider, and how ClawAI’s Local-Only and Privacy-First routing modes fit private workloads. Confirm the live catalog before choosing a plan.
All tasks · Last reviewed:
What running a model locally actually changes
A cloud provider elsewhere in ClawAI’s catalog runs a model on its own infrastructure and charges per request; Ollama and llama.cpp instead load an open-weight model onto hardware you control, so the request never leaves it. That changes who can see the request, not what any given model is capable of — see local AI, on the model providers page, for the full mechanism rather than repeating it here.
ClawAI’s Local-Only and Privacy-First routing modes
Local-Only routing keeps every request on hardware you control, using Ollama or llama.cpp rather than any cloud provider. Privacy-First routing is a separate mode with its own priorities; both exist specifically because not every workload should default to Auto routing. Choosing between them, or pinning a specific local model under Manual Model mode, is a workload decision worth making deliberately rather than leaving to a general-purpose default.
Choosing which open-weight model to run
This page deliberately names no specific open-weight model, for the same reason the local AI provider page does not: the field moves faster than a static page can track, and a stale recommendation is worse than none. See what is local-first AI, linked below, for how to think about the trade-off between an open-weight model you run yourself and a cloud provider.
No specific model is named here on purpose — open-weight models and their capabilities change quickly, and you choose which ones to run. Confirm plan behaviour for local workloads on the pricing page. Confirm the live catalog on the pricing page
Questions people ask
- Which open-weight model should I run for a private workload?
- This page does not recommend one — see what is local-first AI, linked below, for how to think about the choice, since the right model depends on your hardware and task in a way a static page cannot track responsibly.
- What is the difference between Local-Only and Privacy-First routing?
- Local-Only keeps every request on hardware you control via Ollama or llama.cpp; Privacy-First is a separate routing mode with its own priorities. Both exist because not every workload should default to Auto routing.
- Does running a model locally cost anything through ClawAI?
- ClawAI does not charge a per-token rate for a locally run model the way it does for a cloud provider, since there is no cloud provider being billed — the cost is the hardware you already run it on. Confirm current plan behaviour on the pricing page.
Try it rather than take our word for it
ClawAI’s Local-Only routing mode keeps every request on hardware you control via Ollama or llama.cpp — a real, shipped connector, not a roadmap item.