Install
このコンテンツはまだ日本語訳がありません。
Utter runs on Linux on Wayland — niri and KDE Plasma (KWin) are first-class, other compositors get partial support (dictation, typing, launching; no window actions) — with PipeWire and Python 3.12+. macOS is experimental (see macOS). An NVIDIA GPU is recommended for the larger models but not required. Building the settings app needs Node + pnpm and a Rust toolchain. Only x86_64 release assets are published today.
The installer is an interactive wizard. It walks each component (the runner and assistant
CLI, your spoken language, the settings app, the background services, speech models and the
optional Noctalia widget) and asks whether you want it. In a terminal, Enter accepts the
recommended default. --yes accepts them all non-interactively. Every download is verified against
the release’s sha256sums.txt.
English ships inline — the language step defaults to English and downloads nothing, so the default install needs no model fetch. Other languages are opt-in: the step offers the matching multilingual speech model and voice (with sizes) and only downloads what you accept.
Pick the path that suits you. On a Mac none of these apply: see
macOS (experimental) for the .dmg and macos/setup.sh.
1. Clone and install
Section titled “1. Clone and install”Keep the repository and install from your own checkout:
git clone https://github.com/sujaisubbanna/utter-assistant.gitcd utter-assistant
./install.sh --dry-run # walk the wizard, print the plan, change nothing./install.sh # install the components you chooseWant the smallest possible install, distro packages and the background service and nothing else? Use the developer installer:
install/install.sh --dry-runinstall/install.sh --yesUndo either one with ./install.sh --uninstall. Details, including what the developer installer
does step by step, are in Clone and build from source.
2. Remote install
Section titled “2. Remote install”No clone needed:
curl -fsSL https://utter.sujaisubbanna.com/install.sh | bash# recommended defaults, no promptscurl -fsSL https://utter.sujaisubbanna.com/install.sh | bash -s -- --yes
# native package (.deb/.rpm) through your package manager instead of the AppImage (needs sudo)curl -fsSL https://utter.sujaisubbanna.com/install.sh | bash -s -- --package --yes
# only these componentscurl -fsSL https://utter.sujaisubbanna.com/install.sh | bash -s -- --only core,gui --yesOther flags: --skip <csv>, --with-noctalia, --dry-run, --uninstall. The full flag and
environment matrix, the ten wizard steps and where everything ends up are in
Install from the web.
3. Build from source
Section titled “3. Build from source”The assistant core is stdlib-only Python; the settings app is Tauri v2 + React + Tailwind CSS v4.
# the assistant core, straight from the checkoutpython3 -m utter.daemon --text "open youtube" --dry-runscripts/verify.sh # unit + e2e + conformance
# the settings appcd gui-tauripnpm installpnpm tauri build # release binary (and a .deb) in src-tauri/target/releasepnpm tauri dev # ...or run it with hot reloadOn Wayland with a dual-NVIDIA setup, launch the built app with the DMABUF renderer disabled.
The utter-gui launcher does this for you:
WEBKIT_DISABLE_DMABUF_RENDERER=1 ./gui-tauri/src-tauri/target/release/utterSee Clone and build from source for services, the assistant CLI and
the model store.
GPU requirements and latency
Section titled “GPU requirements and latency”The shipped serving scripts assume one NVIDIA GPU shared by both vLLM servers. Per-component footprint:
| Component | Model | Precision | GPU memory setting | On disk |
|---|---|---|---|---|
| Speech recognition (in process) | distil-small.en | faster-whisper, float16 | ~0.5 GB | — |
| Decision head / planner | Qwen3-4B-Instruct-2507-AWQ-4bit | W4A16 (4-bit AWQ) | --gpu-memory-utilization 0.30 | 3.3 GB |
| Screen vision | UI-TARS-2B-SFT | bf16 | --gpu-memory-utilization 0.55 | 9.2 GB |
0.30 + 0.55 = 0.85, so the default pair fits one ~24 GB GPU.
| Tier | What runs | Status |
|---|---|---|
| 24 GB | Full stack; planner ~7.2 GB (0.30), vision ~13 GB (0.55); together ~0.85 of the card | Fits the shipped defaults — the latency numbers below were measured with both models on a 24 GB card |
| 16 GB | Same models with lower UTTER_VISION_GPU_MEM_UTIL and UTTER_PLANNER_GPU_MEM_UTIL (sum below ~0.9) | Expected; untested |
| 8 GB | 2B vision + 4B AWQ planner at lower utilisation (assistant recommend estimates 4B AWQ ≈ 3 GB, UI-TARS-2B ≈ 4 GB) | Expected; untested |
| No GPU / CPU-only | Vision disabled (accessibility-only), smaller STT | Expected; untested |
Only the 24 GB row is what the shipped defaults target; the other rows have not been tested.
Use assistant recommend to see what fits your machine.
Latency was measured on NVIDIA RTX 3090 Ti (24 GB) with the models above, 2026-10-02 (30 warm calls and 1 cold call per path). It will differ per machine.
| Path | Cold (first call) | Warm p50 | Warm p95 |
|---|---|---|---|
| Rules (layer 1, no model) | 9.2 ms | <1 ms | <1 ms |
| Decision head (layer 2, local LLM) | 108.8 ms | 9.2 ms | 11.6 ms |
| Vision (UI-TARS screenshot grounding) | 676.5 ms | 91.0 ms | 140.6 ms |
End-to-end utter assistant --dry-run | 114 ms | 113 ms | 114 ms |
| Sleep → wake (planner reload to ready) | ~21 s | — | — |
- “Cold” is the first call after the servers are up but idle (cold CUDA kernels/caches), not model loading.
- The end-to-end time is dominated by Python interpreter startup (~113 ms), not the decision head (about 9 ms warm).
- The vision numbers include a synthetic 1344×756 image.
See Models for the recommendation rules and the model store.
After installing
Section titled “After installing”Open the Utter settings app, set your push-to-talk keys on the Voice page, and pick a recommended model on the Models page. Then follow Getting started.
systemctl --user enable --now utter-runner.service # start the runnerassistant doctor --json # verify deps + pluginsutter-gui # open the settings window