Zum Inhalt springen

Configuration

Dieser Inhalt ist noch nicht in deiner Sprache verfügbar.

Utter is configured by a single TOML file. The settings app edits the same file (it preserves your comments), so you can use whichever you prefer.

  1. ~/.config/utter/config.toml (honours XDG_CONFIG_HOME) is your override.
  2. When that file does not exist, the shipped config.default.toml from the repository is used.

Start from the shipped default:

Terminal window
mkdir -p ~/.config/utter
cp config.default.toml ~/.config/utter/config.toml

Unknown keys are ignored, so a typo silently does nothing. Check the key names against the tables below.

Separately, the modular runner has its own TOML (runner/config.example.toml in the repo) with [[plugin]] entries and [runner], [socket], [security] and [policy] sections. That file decides which plugins run and what they may do; see Plugin protocol and Trust and safety.

Defaults are those shipped in config.default.toml.

KeyDefaultMeaning
triggerhotkeyhotkey runs Utter’s own evdev push-to-talk listener
KeyDefaultMeaning
keyKEY_RIGHTCTRLevdev key held for the standalone push-to-talk path
KeyDefaultMeaning
dictation_keyKEY_F13held: the transcript is typed into the focused field
assistant_keyKEY_INSERTheld: the transcript is executed as a command
KeyDefaultMeaning
sample_rate16000speech-to-text input rate
channels1mono
device""empty = default input
KeyDefaultMeaning
backendfaster_whisperfaster_whisper, whisper_cpp or none
modeldistil-small.enbackend-specific model name or path
devicecudafaster-whisper device
compute_typefloat16faster-whisper compute type
languageautospoken language: auto, en, en-GB, de-DE, … (independent of the UI language)

Linux spoken replies (macOS uses [macos] tts_backend / tts_voice / tts_rate):

KeyDefaultMeaning
enabledtrueLinux spoken replies on/off (macOS still uses [macos])
engineautoauto (piper → espeak-ng → espeak → spd-say), none, or an engine name to force
languageautospoken language, same resolution rules as [stt] language
voice""engine voice: espeak name (en-gb, de) or a Piper .onnx model/path; empty derives from language
KeyDefaultMeaning
llm_fallbacktrueallow the free-form fallback planner
llm_base_urlhttp://127.0.0.1:8001/v1OpenAI-compatible endpoint for the decision head and planner
llm_modelqwen3-4bserved model name
decide_threshold0.5minimum probability for the decision head to act
KeyDefaultMeaning
enabledtrueenable the vision tier (T3)
base_urlhttp://127.0.0.1:8000/v1UI-TARS vLLM endpoint
modeluitarsserved model name
target_width1344screenshot resize width; the biggest latency lever
cuda_visible_devices"1"GPU pin for the vision server
KeyDefaultMeaning
require_confirm["send","submit","delete","purchase","pay","confirm order"]substrings in a step’s args that force confirmation
click_duration_ms40how long a simulated click is held
preferred_browser""profile or app id for web actions; empty = the browser you are looking at, else any open one, else the desktop default
KeyDefaultMeaning
enabledtrueallow the sleep phrase
trigger["go to sleep"]phrases that put Utter to sleep (assistant mode only)
services["utter-vision", "utter-planner"]user units stopped on sleep
unload_speechtruealso unload the speech model
on_idletruealso fall asleep after idle_minutes without Utter activity (independent of enabled)
idle_minutes15minutes without a push-to-talk key, a spoken command or a wake before Utter sleeps by itself; desktop input elsewhere does not count, and the timer waits while Utter is listening or running a command
KeyDefaultMeaning
enabledtruethe optional Noctalia on-screen display; UTTER_OSD=0 disables it at runtime
positionbottom_centerwhere the panel appears
dismiss_ms1200how long the result stays visible
streamtruebest-effort live transcript while you speak
stream_interval_ms700how often the live transcript updates
window_s6audio window for live decoding

See Noctalia widget and OSD.

KeyDefaultMeaning
log_levelINFOdaemon log level

Utter runs its own standalone evdev push-to-talk. The daemon runs its own key listener and transcribes with the configured speech backend. The two dedicated [ptt] keys are handled by the same listener: the dictation key types text normally, and the assistant key routes the text to Utter and never types it.

Keys are evdev names (KEY_F13, KEY_INSERT, KEY_RIGHTCTRL, …). The listener only needs your user to be in the input group. It does not grab the device, so the key keeps working for everything else.

The shipped defaults (KEY_F13 and KEY_INSERT) assume a kernel-level remap with keyd that turns two awkward physical keys into those codes. This is one way to do it, in /etc/keyd/default.conf:

[ids]
*
[main]
rightalt = insert
capslock = f13
[shift]
capslock = capslock
  • Physical Caps Lock becomes F13 (dictation key).
  • Shift + Caps Lock is the real Caps Lock, thanks to the [shift] layer.
  • Physical Right Alt becomes Insert (assistant key). Hold Right Alt to issue a command without typing it.

Reload keyd after editing (needs root):

Terminal window
sudo keyd reload
# or: sudo systemctl restart keyd

Then make sure [ptt] matches the remapped names. If you do not want a remap, pick any key your keyboard already has and change [ptt] on the Voice page instead.

Three backends are supported:

  • whisper_cpp, via pywhispercpp. No GPU required; the model stays resident.
  • faster_whisper, an optional dependency with GPU support through device and compute_type. Falls back to CPU / int8 if the requested device fails.
  • none, transcription disabled.

Model lookup for whisper.cpp checks, in order: an explicit path or $UTTER_WHISPER_MODEL, $UTTER_MODELS_DIR, <repo>/models/whisper, then ~/.cache/whisper. The default whisper.cpp filename is ggml-small.en.bin. If no local file exists and the configured name is a valid whisper.cpp model, pywhispercpp downloads it.

[stt] language and [tts] language are a separate axis from the settings-app UI language (English/Español strings). Resolution is shared (utter/locale.py): a concrete value is normalised (en_GB/en-US.UTF-8 → en-GB), while auto (the default) reads LC_ALL → LC_MESSAGES → LANG. If nothing resolves, the language stays unknown: Whisper auto-detects and TTS uses the engine default — Utter never silently guesses English.

Whisper .en checkpoints (distil-small.en, ggml-base.en.bin, …) are English-only. Pairing one with a non-English language logs a warning, and Utter does not switch models or download anything by itself. The Voice page offers an explicit switch instead:

ModelApprox. downloadNotes
small~480 MBmultilingual, modest GPU/CPU cost
large-v3-turbo~1.6 GBmultilingual, fastest large variant

On Linux, [tts] engine = "auto" probes piper (only with a local voice model), then espeak-ng, espeak, spd-say; empty [tts] voice derives an espeak voice from the language (de-DE → de, en-GB → en-gb). No engine installed means a warning and spoken replies disabled, never a crash. English and Español are the only UI locales shipped inline; more UI locales are not implemented yet. On macOS, a resolved [stt] language becomes the SFSpeechRecognizer locale, with [macos] speech_locale as the fallback.

Two things use the local language-model endpoint on port 8001:

  • The constrained decision head asks the model to pick one letter from a list of fully resolved candidates. [router] decide_threshold gates the choice. It tries vLLM’s structured_outputs.choice, then the legacy guided_choice, then a plain call, and fails open to rules on any error.
  • The free-form planner is the JSON fallback used only when rules and the decision head both fail. Disable it with [router] llm_fallback = false.

Serve the planner with vLLM (GPU 1, port 8001, served name qwen3-4b):

Terminal window
scripts/serve_planner.sh
# backgrounded with logs:
setsid bash -c 'scripts/serve_planner.sh > /tmp/vllm-planner.log 2>&1 &'

Override model, port and GPU with UTTER_PLANNER_MODEL_PATH, UTTER_PLANNER_PORT, UTTER_PLANNER_SERVED_NAME, UTTER_PLANNER_GPU_MEM_UTIL and UTTER_CUDA_VISIBLE_DEVICES. The default checkpoint is a 4-bit AWQ build of Qwen3-4B-Instruct. Any OpenAI-compatible server (Ollama, llama.cpp) works too; point llm_base_url and llm_model at it from the LLM page.

The vision tier grounds a description to a click point using UI-TARS served by vLLM:

Terminal window
scripts/serve_vision.sh
# backgrounded with logs:
setsid bash -c 'scripts/serve_vision.sh > /tmp/vllm-serve.log 2>&1 &'
  • Endpoint http://127.0.0.1:8000/v1, model name uitars.
  • Overrides: UTTER_VISION_MODEL_PATH, UTTER_VISION_PORT, UTTER_VISION_GPU_MEM_UTIL, UTTER_CUDA_VISIBLE_DEVICES.
  • The planner and vision servers can share one GPU through vLLM’s memory utilisation flags (vision 0.55, planner 0.30 by default).

On the client side, Utter captures with grim, resizes the screenshot to target_width (rounded to a multiple of 28 for the model) and converts the model’s 0 to 1000 normalised coordinates back to screen pixels. Client-side environment overrides: UTTER_VISION_URL, UTTER_VISION_MODEL, UTTER_VISION_TARGET_WIDTH, UTTER_VISION_TIMEOUT.

Disable vision entirely with [vision] enabled = false. “Click …” steps then stop after the accessibility attempt and report that vision is disabled.

Utter uses PipeWire: sounddevice for capture, pactl and pw-play for routing and sounds. It does not ship a PipeWire configuration.

  • Capture opens an input stream at [audio] sample_rate and channels. Keep a denoised source (for example an RNNoise filter chain) as the default input for clean dictation.
  • Defaults keeper: scripts/utter-audio-defaults.sh loops every 10 seconds and re-asserts the default sink and source after PipeWire restarts. It is idempotent, never forces volume, and only unmutes when it (re)sets the sink. Edit the SINK and SRC names at the top of the script for your hardware.
Terminal window
pactl list short sinks
pactl list short sources
pactl get-default-sink
pactl get-default-source

Three UI sounds are synthesised on first use (standard library only) and played without blocking with pw-play (fallback paplay):

NameWhen
startlistening started (one tone per mode)
detecteda command was understood
not_detectednothing matched

Disable them with UTTER_SOUNDS=0.

The repository ships these unit files:

UnitFileRole
utter-runner.serviceinstall/utter-runner.servicethe modular runner (plugin supervisor)
utter.servicesystemd/utter.servicethe legacy daemon (python -m utter.daemon)
ydotoold.servicesystemd/ydotoold.serviceinput-injection daemon used by ydotool

All are user units bound to graphical-session.target. The legacy daemon unit has placeholder paths you substitute for your checkout:

Terminal window
mkdir -p ~/.config/systemd/user
sed "s|@REPO@|$PWD|g" systemd/utter.service > ~/.config/systemd/user/utter.service
cp systemd/ydotoold.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now ydotoold.service utter.service
systemctl --user status utter.service
journalctl --user -u utter.service -f

ydotoold.service runs ydotoold with a per-user socket; Utter sets YDOTOOL_SOCKET for its clients automatically.

Adding or editing applications (profiles, aliases, compositor app ids, catalogs, niri phrases) is covered in Apps and actions. The short version: create ~/.config/utter/profiles/<kind>_<id>.yaml or edit the app on the App actions page, then regenerate the catalog with scripts/gen_app_catalog.py.

Read Trust and safety in full. The configuration points are:

  • Confirmation list. [actions] require_confirm is a list of substrings; if any appears in a step’s args, confirmation is requested.

  • Dangerous ops are off by default in the runner. Terminal commands and raw input need an explicit opt-in in the runner config and still require confirmation:

    [policy]
    enabled_ops = ["action.terminal"] # or "action.input"
  • Never weaken policy by letting screen, accessibility, OCR or clipboard content author action arguments. The runner enforces this invariant and rejects such requests.

  • Dry run. The utter_py plugin defaults UTTER_DRY_RUN to on; only UTTER_DRY_RUN=0 touches the desktop. Route-only testing: python -m utter.daemon --text "<cmd>" --dry-run.

An utterance can name its target up front instead of acting on the focused window. The leading <app> must resolve to a real profile or a CLI agent (utter/data/cli_agents.json); generic words (“media”, “editor”, “music”, “terminal”) are never claimed, so type ok, press enter and pause keep their existing focused behaviour.

UtterancePlan
codex type okTYPE_TEXT{text="ok", app="codex"}
codex press enterKEY{chord="Return", app="codex"}
spotify pauseMEDIA{command="pause", app="spotify"}
close steamCLOSE_APP{app="steam"}

Target-first type/write, press/hit/send and <app> <media-command> are matched before the generic focused type/key rules. close <app> closes one of the app’s windows. A window_id is never set at plan time: the executor resolves the target from the live compositor window list, which may only select a window.

Mechanism. Wayland has no background key injection — wtype/ydotool emit to the focused surface only, and niri/KWin expose no per-window injection. So a targeted key/type_text is a focus round-trip: focus the target, inject, then restore the previous focus. close_app and media are genuinely focus-free; they never focus.

Latency (measured on one machine — niri, RTX 3090 Ti, 2026-10-02):

  • same workspace: ~38 ms, invisible (no viewport movement);
  • cross-workspace, compositor animations off: ~41 ms;
  • cross-workspace, animations on: ~250 ms of visible viewport scroll (horizontal-view-movement) — this is why it is gated.
[target]
mode = "round_trip" # round_trip | leave | off
cross_workspace = "ask" # ask | allow | refuse (fallback; see [wayland])
restore = "if_unchanged" # if_unchanged | always | never
focus_timeout_ms = 500
[wayland]
cross_workspace = "auto" # auto | ask | allow | refuse
assume_animations_off = false
  • [target].mode: round_trip = focus, act, restore; leave = focus, act, stay on the target; off = only focus/close/media, refuse targeted input.
  • [target].cross_workspace is the fallback policy. [wayland].cross_workspace supersedes it: auto allows cross-workspace without asking only when assume_animations_off is true, otherwise it falls back to [target].cross_workspace; an explicit ask/allow/refuse wins.
  • restore = "if_unchanged" only restores focus if the user did not move away during the round-trip.
  • Same-workspace single-window input runs without asking; cross-workspace or a fullscreen previous window requires confirmation (or refuses when there is no confirmation channel, e.g. headless).

The app name comes from the user (trusted intent) and profile app_ids are precomputed; the compositor window list is untrusted and may only select a window — a window title is never turned into text or a command. screen-provenance args are still rejected -32006, and close_app requires confirmation. See Trust and safety for the full rules.

PlatformMechanismNotes
Linux/Wayland (niri, KWin)Focus round-tripVisible only when the target is off-screen or on another workspace with animations on.
HyprlandFocus round-tripsendshortcut exists but is unreliable for native-Wayland Electron/Chromium apps and can silently do nothing; the round-trip is the safer path.
macOSCGEventPostToPidNative key event posted directly to a target process with no focus change (keyboard only; the mouse cannot target a background window). Experimental, untested on real Apple hardware.
WindowsNot supported—