Configuration
このコンテンツはまだ日本語訳がありません。
Utter is configured by a single TOML file. The settings app edits the same file (it preserves your comments), so you can use whichever you prefer.
Config files and precedence
Section titled “Config files and precedence”~/.config/utter/config.toml(honoursXDG_CONFIG_HOME) is your override.- When that file does not exist, the shipped
config.default.tomlfrom the repository is used.
Start from the shipped default:
mkdir -p ~/.config/uttercp config.default.toml ~/.config/utter/config.tomlUnknown keys are ignored, so a typo silently does nothing. Check the key names against the tables below.
Separately, the modular runner has its own TOML (runner/config.example.toml in the repo)
with [[plugin]] entries and [runner], [socket], [security] and [policy] sections. That
file decides which plugins run and what they may do; see Plugin protocol and
Trust and safety.
Config sections
Section titled “Config sections”Defaults are those shipped in config.default.toml.
[general]
Section titled “[general]”| Key | Default | Meaning |
|---|---|---|
trigger | hotkey | hotkey runs Utter’s own evdev push-to-talk listener |
[hotkey]
Section titled “[hotkey]”| Key | Default | Meaning |
|---|---|---|
key | KEY_RIGHTCTRL | evdev key held for the standalone push-to-talk path |
| Key | Default | Meaning |
|---|---|---|
dictation_key | KEY_F13 | held: the transcript is typed into the focused field |
assistant_key | KEY_INSERT | held: the transcript is executed as a command |
[audio]
Section titled “[audio]”| Key | Default | Meaning |
|---|---|---|
sample_rate | 16000 | speech-to-text input rate |
channels | 1 | mono |
device | "" | empty = default input |
| Key | Default | Meaning |
|---|---|---|
backend | faster_whisper | faster_whisper, whisper_cpp or none |
model | distil-small.en | backend-specific model name or path |
device | cuda | faster-whisper device |
compute_type | float16 | faster-whisper compute type |
language | auto | spoken language: auto, en, en-GB, de-DE, … (independent of the UI language) |
Linux spoken replies (macOS uses [macos] tts_backend / tts_voice / tts_rate):
| Key | Default | Meaning |
|---|---|---|
enabled | true | Linux spoken replies on/off (macOS still uses [macos]) |
engine | auto | auto (piper → espeak-ng → espeak → spd-say), none, or an engine name to force |
language | auto | spoken language, same resolution rules as [stt] language |
voice | "" | engine voice: espeak name (en-gb, de) or a Piper .onnx model/path; empty derives from language |
[router]
Section titled “[router]”| Key | Default | Meaning |
|---|---|---|
llm_fallback | true | allow the free-form fallback planner |
llm_base_url | http://127.0.0.1:8001/v1 | OpenAI-compatible endpoint for the decision head and planner |
llm_model | qwen3-4b | served model name |
decide_threshold | 0.5 | minimum probability for the decision head to act |
[vision]
Section titled “[vision]”| Key | Default | Meaning |
|---|---|---|
enabled | true | enable the vision tier (T3) |
base_url | http://127.0.0.1:8000/v1 | UI-TARS vLLM endpoint |
model | uitars | served model name |
target_width | 1344 | screenshot resize width; the biggest latency lever |
cuda_visible_devices | "1" | GPU pin for the vision server |
[actions]
Section titled “[actions]”| Key | Default | Meaning |
|---|---|---|
require_confirm | ["send","submit","delete","purchase","pay","confirm order"] | substrings in a step’s args that force confirmation |
click_duration_ms | 40 | how long a simulated click is held |
preferred_browser | "" | profile or app id for web actions; empty = the browser you are looking at, else any open one, else the desktop default |
[sleep]
Section titled “[sleep]”| Key | Default | Meaning |
|---|---|---|
enabled | true | allow the sleep phrase |
trigger | ["go to sleep"] | phrases that put Utter to sleep (assistant mode only) |
services | ["utter-vision", "utter-planner"] | user units stopped on sleep |
unload_speech | true | also unload the speech model |
on_idle | true | also fall asleep after idle_minutes without Utter activity (independent of enabled) |
idle_minutes | 15 | minutes without a push-to-talk key, a spoken command or a wake before Utter sleeps by itself; desktop input elsewhere does not count, and the timer waits while Utter is listening or running a command |
| Key | Default | Meaning |
|---|---|---|
enabled | true | the optional Noctalia on-screen display; UTTER_OSD=0 disables it at runtime |
position | bottom_center | where the panel appears |
dismiss_ms | 1200 | how long the result stays visible |
stream | true | best-effort live transcript while you speak |
stream_interval_ms | 700 | how often the live transcript updates |
window_s | 6 | audio window for live decoding |
[daemon]
Section titled “[daemon]”| Key | Default | Meaning |
|---|---|---|
log_level | INFO | daemon log level |
Hotkeys and push-to-talk
Section titled “Hotkeys and push-to-talk”Utter runs its own standalone evdev push-to-talk. The daemon runs its own key
listener and transcribes with the configured speech backend. The two dedicated
[ptt] keys are handled by the same listener: the dictation key types text
normally, and the assistant key routes the text to Utter and never types it.
Keys are evdev names (KEY_F13, KEY_INSERT, KEY_RIGHTCTRL, …). The listener only needs
your user to be in the input group. It does not grab the device, so the key keeps working
for everything else.
keyd remap
Section titled “keyd remap”The shipped defaults (KEY_F13 and KEY_INSERT) assume a kernel-level remap with
keyd that turns two awkward physical keys into those codes.
This is one way to do it, in /etc/keyd/default.conf:
[ids]*
[main]rightalt = insertcapslock = f13
[shift]capslock = capslock- Physical Caps Lock becomes F13 (dictation key).
- Shift + Caps Lock is the real Caps Lock, thanks to the
[shift]layer. - Physical Right Alt becomes Insert (assistant key). Hold Right Alt to issue a command without typing it.
Reload keyd after editing (needs root):
sudo keyd reload# or: sudo systemctl restart keydThen make sure [ptt] matches the remapped names. If you do not want a remap, pick any key your
keyboard already has and change [ptt] on the Voice page instead.
Speech-to-text backends
Section titled “Speech-to-text backends”Three backends are supported:
whisper_cpp, viapywhispercpp. No GPU required; the model stays resident.faster_whisper, an optional dependency with GPU support throughdeviceandcompute_type. Falls back to CPU / int8 if the requested device fails.none, transcription disabled.
Model lookup for whisper.cpp checks, in order: an explicit path or $UTTER_WHISPER_MODEL,
$UTTER_MODELS_DIR, <repo>/models/whisper, then ~/.cache/whisper. The default whisper.cpp
filename is ggml-small.en.bin. If no local file exists and the configured name is a valid
whisper.cpp model, pywhispercpp downloads it.
Spoken language (STT and TTS)
Section titled “Spoken language (STT and TTS)”[stt] language and [tts] language are a separate axis from the settings-app UI
language (English/Español strings). Resolution is shared (utter/locale.py): a concrete
value is normalised (en_GB/en-US.UTF-8 → en-GB), while auto (the default) reads
LC_ALL → LC_MESSAGES → LANG. If nothing resolves, the language stays unknown: Whisper
auto-detects and TTS uses the engine default — Utter never silently guesses English.
Whisper .en checkpoints (distil-small.en, ggml-base.en.bin, …) are English-only.
Pairing one with a non-English language logs a warning, and Utter does not switch models or
download anything by itself. The Voice page offers an explicit switch instead:
| Model | Approx. download | Notes |
|---|---|---|
small | ~480 MB | multilingual, modest GPU/CPU cost |
large-v3-turbo | ~1.6 GB | multilingual, fastest large variant |
On Linux, [tts] engine = "auto" probes piper (only with a local voice model), then
espeak-ng, espeak, spd-say; empty [tts] voice derives an espeak voice from the
language (de-DE → de, en-GB → en-gb). No engine installed means a warning and spoken
replies disabled, never a crash. English and Español are the only UI locales shipped inline;
more UI locales are not implemented yet. On macOS, a resolved [stt] language becomes the
SFSpeechRecognizer locale, with [macos] speech_locale as the fallback.
The decision head and the planner
Section titled “The decision head and the planner”Two things use the local language-model endpoint on port 8001:
- The constrained decision head asks the model to pick one letter from a list of fully
resolved candidates.
[router] decide_thresholdgates the choice. It tries vLLM’sstructured_outputs.choice, then the legacyguided_choice, then a plain call, and fails open to rules on any error. - The free-form planner is the JSON fallback used only when rules and the decision head both
fail. Disable it with
[router] llm_fallback = false.
Serve the planner with vLLM (GPU 1, port 8001, served name qwen3-4b):
scripts/serve_planner.sh# backgrounded with logs:setsid bash -c 'scripts/serve_planner.sh > /tmp/vllm-planner.log 2>&1 &'Override model, port and GPU with UTTER_PLANNER_MODEL_PATH, UTTER_PLANNER_PORT,
UTTER_PLANNER_SERVED_NAME, UTTER_PLANNER_GPU_MEM_UTIL and UTTER_CUDA_VISIBLE_DEVICES. The
default checkpoint is a 4-bit AWQ build of Qwen3-4B-Instruct. Any OpenAI-compatible server
(Ollama, llama.cpp) works too; point llm_base_url and llm_model at it from the LLM page.
Vision
Section titled “Vision”The vision tier grounds a description to a click point using UI-TARS served by vLLM:
scripts/serve_vision.sh# backgrounded with logs:setsid bash -c 'scripts/serve_vision.sh > /tmp/vllm-serve.log 2>&1 &'- Endpoint
http://127.0.0.1:8000/v1, model nameuitars. - Overrides:
UTTER_VISION_MODEL_PATH,UTTER_VISION_PORT,UTTER_VISION_GPU_MEM_UTIL,UTTER_CUDA_VISIBLE_DEVICES. - The planner and vision servers can share one GPU through vLLM’s memory utilisation flags (vision 0.55, planner 0.30 by default).
On the client side, Utter captures with grim, resizes the screenshot to target_width
(rounded to a multiple of 28 for the model) and converts the model’s 0 to 1000 normalised
coordinates back to screen pixels. Client-side environment overrides: UTTER_VISION_URL,
UTTER_VISION_MODEL, UTTER_VISION_TARGET_WIDTH, UTTER_VISION_TIMEOUT.
Disable vision entirely with [vision] enabled = false. “Click …” steps then stop after the
accessibility attempt and report that vision is disabled.
Utter uses PipeWire: sounddevice for capture, pactl and pw-play for routing and sounds.
It does not ship a PipeWire configuration.
- Capture opens an input stream at
[audio] sample_rateandchannels. Keep a denoised source (for example an RNNoise filter chain) as the default input for clean dictation. - Defaults keeper:
scripts/utter-audio-defaults.shloops every 10 seconds and re-asserts the default sink and source after PipeWire restarts. It is idempotent, never forces volume, and only unmutes when it (re)sets the sink. Edit theSINKandSRCnames at the top of the script for your hardware.
pactl list short sinkspactl list short sourcespactl get-default-sinkpactl get-default-sourceSounds
Section titled “Sounds”Three UI sounds are synthesised on first use (standard library only) and played without blocking
with pw-play (fallback paplay):
| Name | When |
|---|---|
start | listening started (one tone per mode) |
detected | a command was understood |
not_detected | nothing matched |
Disable them with UTTER_SOUNDS=0.
Services
Section titled “Services”The repository ships these unit files:
| Unit | File | Role |
|---|---|---|
utter-runner.service | install/utter-runner.service | the modular runner (plugin supervisor) |
utter.service | systemd/utter.service | the legacy daemon (python -m utter.daemon) |
ydotoold.service | systemd/ydotoold.service | input-injection daemon used by ydotool |
All are user units bound to graphical-session.target. The legacy daemon unit has placeholder
paths you substitute for your checkout:
mkdir -p ~/.config/systemd/usersed "s|@REPO@|$PWD|g" systemd/utter.service > ~/.config/systemd/user/utter.servicecp systemd/ydotoold.service ~/.config/systemd/user/systemctl --user daemon-reloadsystemctl --user enable --now ydotoold.service utter.service
systemctl --user status utter.servicejournalctl --user -u utter.service -fydotoold.service runs ydotoold with a per-user socket; Utter sets YDOTOOL_SOCKET for its
clients automatically.
Profiles and apps
Section titled “Profiles and apps”Adding or editing applications (profiles, aliases, compositor app ids, catalogs, niri phrases) is
covered in Apps and actions. The short version: create
~/.config/utter/profiles/<kind>_<id>.yaml or edit the app on the App actions page, then
regenerate the catalog with scripts/gen_app_catalog.py.
Safety knobs
Section titled “Safety knobs”Read Trust and safety in full. The configuration points are:
-
Confirmation list.
[actions] require_confirmis a list of substrings; if any appears in a step’s args, confirmation is requested. -
Dangerous ops are off by default in the runner. Terminal commands and raw input need an explicit opt-in in the runner config and still require confirmation:
[policy]enabled_ops = ["action.terminal"] # or "action.input" -
Never weaken policy by letting screen, accessibility, OCR or clipboard content author action arguments. The runner enforces this invariant and rejects such requests.
-
Dry run. The
utter_pyplugin defaultsUTTER_DRY_RUNto on; onlyUTTER_DRY_RUN=0touches the desktop. Route-only testing:python -m utter.daemon --text "<cmd>" --dry-run.
App-targeted actions and background input
Section titled “App-targeted actions and background input”An utterance can name its target up front instead of acting on the focused window. The leading
<app> must resolve to a real profile or a CLI agent (utter/data/cli_agents.json); generic words
(“media”, “editor”, “music”, “terminal”) are never claimed, so type ok, press enter and
pause keep their existing focused behaviour.
| Utterance | Plan |
|---|---|
codex type ok | TYPE_TEXT{text="ok", app="codex"} |
codex press enter | KEY{chord="Return", app="codex"} |
spotify pause | MEDIA{command="pause", app="spotify"} |
close steam | CLOSE_APP{app="steam"} |
Target-first type/write, press/hit/send and <app> <media-command> are matched before the
generic focused type/key rules. close <app> closes one of the app’s windows. A window_id is
never set at plan time: the executor resolves the target from the live compositor window list, which
may only select a window.
Mechanism. Wayland has no background key injection — wtype/ydotool emit to the focused
surface only, and niri/KWin expose no per-window injection. So a targeted key/type_text is a
focus round-trip: focus the target, inject, then restore the previous focus. close_app and
media are genuinely focus-free; they never focus.
Latency (measured on one machine — niri, RTX 3090 Ti, 2026-10-02):
- same workspace: ~38 ms, invisible (no viewport movement);
- cross-workspace, compositor animations off: ~41 ms;
- cross-workspace, animations on: ~250 ms of visible viewport scroll
(
horizontal-view-movement) — this is why it is gated.
[target]mode = "round_trip" # round_trip | leave | offcross_workspace = "ask" # ask | allow | refuse (fallback; see [wayland])restore = "if_unchanged" # if_unchanged | always | neverfocus_timeout_ms = 500
[wayland]cross_workspace = "auto" # auto | ask | allow | refuseassume_animations_off = false[target].mode:round_trip= focus, act, restore;leave= focus, act, stay on the target;off= only focus/close/media, refuse targeted input.[target].cross_workspaceis the fallback policy.[wayland].cross_workspacesupersedes it:autoallows cross-workspace without asking only whenassume_animations_offis true, otherwise it falls back to[target].cross_workspace; an explicitask/allow/refusewins.restore = "if_unchanged"only restores focus if the user did not move away during the round-trip.- Same-workspace single-window input runs without asking; cross-workspace or a fullscreen previous window requires confirmation (or refuses when there is no confirmation channel, e.g. headless).
The app name comes from the user (trusted intent) and profile app_ids are precomputed; the
compositor window list is untrusted and may only select a window — a window title is never turned
into text or a command. screen-provenance args are still rejected -32006, and close_app
requires confirmation. See Trust and safety for the full rules.
| Platform | Mechanism | Notes |
|---|---|---|
| Linux/Wayland (niri, KWin) | Focus round-trip | Visible only when the target is off-screen or on another workspace with animations on. |
| Hyprland | Focus round-trip | sendshortcut exists but is unreliable for native-Wayland Electron/Chromium apps and can silently do nothing; the round-trip is the safer path. |
| macOS | CGEventPostToPid | Native key event posted directly to a target process with no focus change (keyboard only; the mouse cannot target a background window). Experimental, untested on real Apple hardware. |
| Windows | Not supported | — |