Skip to content

Models

Utter uses up to three local models, and works with none of them. Rules, window state and accessibility handle the deterministic commands. Each model you add unlocks something:

TierWhat it doesWithout it
Speech recognitionturns your voice into textthe push-to-talk keys have nothing to transcribe; text commands and the dry run still work
Decision headchooses among prepared actions when no rule matchesonly rule-matched phrasing works
Screen visionfinds a described element on a screenshot“click …” stops after the accessibility attempt

Models are your choice. The installer never downloads one, and the settings app downloads nothing without you pressing the button.

assistant recommend (and the Recommended for your computer section of the Models page) looks at your GPU vendor, its VRAM and your RAM, and suggests one model per tier. These are the rules it applies:

HardwareBackendModelDeviceFootprint
NVIDIA GPU with 8 GB VRAM or morefaster-whisperlarge-v3-turboCUDAabout 2 GB VRAM
AMD or Intel GPU (Vulkan)whisper.cppsmallVulkanabout 1 GB VRAM
No usable GPUwhisper.cppbase.enCPU, int8about 1 GB RAM
VRAMModel classQuantisationFootprint
24 GB or more7 to 8 B parametersAWQabout 6 GB VRAM
8 to 24 GB4 B parametersAWQabout 3 GB VRAM
8 GB or less1.5 to 3 B parameters on CPU, or an existing endpointq4about 3 GB RAM

The shipped default serving script uses a 4-bit AWQ build of Qwen3-4B-Instruct under the served name qwen3-4b.

VRAMModelFootprint
16 GB or moreUI-TARS-7Babout 8 GB VRAM
6 to 16 GBUI-TARS-2Babout 4 GB VRAM
Less than 6 GBnone; accessibility only

On the Models page, Get it pre-fills a ready-to-pull source for a recommendation when one is known (for UI-TARS-7B that is hf:ByteDance-Seed/UI-TARS-1.5-7B). Applying a recommendation also updates the matching config keys; for the decision model, match the name to what your server actually serves.

$XDG_DATA_HOME/utter/models/ # override with UTTER_MODELS
manifests/<host>/<ns>/<name>/<tag>.json
blobs/sha256-<hex>

The layout is Ollama-style: manifests name a model and tag; blobs are content-addressed by SHA-256, so two models that share a file share one blob.

FormResolves to
hf:org/repothe Hugging Face repository
hf:org/repo:fileone file from that repository (.../resolve/main/file)
https://...a direct download
file://... or a bare local patha file already on disk
Terminal window
assistant models list
assistant models show <name>
assistant models pull hf:org/repo[:file] [--tag latest]
assistant models rm <name>
assistant models prune

Add --json to any of them for machine-readable output; the settings app uses the same JSON and shows live progress while a pull runs.

  • Resumable. A partial file is kept and continued with curl -C - (or an HTTP Range request when curl is absent).
  • Retries with backoff and stall detection.
  • Verified. The SHA-256 of the finished file is checked before it becomes a blob. When the server exposes a digest up front it is compared too. A mismatch fails the pull.
  • Atomic. The blob is renamed into place only after verification.
  • One pull at a time per model (a lock file), with a disk-space preflight before starting.

rm drops the manifest and any blob no other manifest references. prune collects orphaned blobs and unfinished downloads; the Models page calls this Clean up.

The store holds the files. Serving them is a separate step, done by vLLM (or any OpenAI-compatible server such as Ollama or llama.cpp) and pointed to from ~/.config/utter/config.toml:

[router]
llm_base_url = "http://127.0.0.1:8001/v1"
llm_model = "qwen3-4b"
[vision]
base_url = "http://127.0.0.1:8000/v1"
model = "uitars"

scripts/serve_planner.sh and scripts/serve_vision.sh start vLLM with sensible defaults; see Configuration. If you point either endpoint at another machine, the settings app marks it in amber: that is the one case where your data leaves the computer.

The shipped serving scripts assume one NVIDIA GPU shared by both vLLM servers. Per-component footprint:

ComponentModelPrecisionGPU memory settingOn disk
Speech recognition (in process)distil-small.enfaster-whisper, float16~0.5 GB—
Decision head / plannerQwen3-4B-Instruct-2507-AWQ-4bitW4A16 (4-bit AWQ)--gpu-memory-utilization 0.303.3 GB
Screen visionUI-TARS-2B-SFTbf16--gpu-memory-utilization 0.559.2 GB

0.30 + 0.55 = 0.85, so the default pair fits one ~24 GB GPU.

TierWhat runsStatus
24 GBFull stack; planner ~7.2 GB (0.30), vision ~13 GB (0.55); together ~0.85 of the cardFits the shipped defaults — the latency numbers below were measured with both models on a 24 GB card
16 GBSame models with lower UTTER_VISION_GPU_MEM_UTIL and UTTER_PLANNER_GPU_MEM_UTIL (sum below ~0.9)Expected; untested
8 GB2B vision + 4B AWQ planner at lower utilisation (assistant recommend estimates 4B AWQ ≈ 3 GB, UI-TARS-2B ≈ 4 GB)Expected; untested
No GPU / CPU-onlyVision disabled (accessibility-only), smaller STTExpected; untested

Only the 24 GB row is what the shipped defaults target; the other rows have not been tested. Use assistant recommend (or the Recommended for your computer section of the Models page) to see what fits your machine.

Latency was measured on NVIDIA RTX 3090 Ti (24 GB) with the models above, 2026-10-02 (30 warm calls and 1 cold call per path). It will differ per machine.

PathCold (first call)Warm p50Warm p95
Rules (layer 1, no model)9.2 ms<1 ms<1 ms
Decision head (layer 2, local LLM)108.8 ms9.2 ms11.6 ms
Vision (UI-TARS screenshot grounding)676.5 ms91.0 ms140.6 ms
End-to-end utter assistant --dry-run114 ms113 ms114 ms
Sleep → wake (planner reload to ready)~21 s——
  • “Cold” is the first call after the servers are up but idle (cold CUDA kernels/caches), not model loading.
  • The end-to-end time is dominated by Python interpreter startup (~113 ms), not the decision head (about 9 ms warm).
  • The vision numbers include a synthetic 1344×756 image.

Saying your sleep phrase stops the model services listed in [sleep] services and unloads the speech model, freeing the GPU for something else. Utter also does this by itself after [sleep] idle_minutes (15 by default) without a push-to-talk key or a spoken command; typing and clicking in other apps do not count. Holding a push-to-talk key wakes everything: speech comes back first, and the bigger models reload in the background. See Getting started and the [sleep] keys in Configuration.