Skip to content

Getting started

This page assumes Utter is installed. If it is not, start with the install overview.

  • Linux on Wayland. niri and KDE Plasma (KWin) are first-class; other compositors get partial support (dictation, typing, launching; no window actions). macOS is experimental.
  • PipeWire for audio.
  • Python 3.12 or newer.
  • An NVIDIA GPU is recommended for the larger models, but not required. Utter runs with no models at all, and with CPU-only speech recognition.

1. Start the runner and open the settings app

Section titled “1. Start the runner and open the settings app”
Terminal window
systemctl --user enable --now utter-runner.service # start the background runner
assistant doctor # check tools, plugins and permissions
utter-gui # open the settings window

If utter-gui is not found, the installer’s prefix is probably not on your PATH:

Terminal window
export PATH="$HOME/.local/bin:$PATH"

assistant doctor tells you which command-line tools are missing and whether your user can read input devices. The most common fixes are covered in Troubleshooting.

Utter has two push-to-talk keys, one per mode. Both are plain evdev key names and both are set on the Voice page of the settings app, or under [ptt] in ~/.config/utter/config.toml:

[ptt]
dictation_key = "KEY_F13"
assistant_key = "KEY_INSERT"
HoldWhat happens to your wordsOn screen
Assistantthe assistant key (Insert by default)Turned into an action: open apps and sites, click things, press shortcutsA yellow waveform
Dictationthe dictation key (F13 by default)Typed into whatever field has focusA blue waveform with “Dictation · typing”

Each mode has its own start sound, so you can tell them apart without looking.

The defaults assume a keyboard remap that turns an awkward physical key into Insert or F13. Most people will want keys that are actually on their keyboard. Any evdev name works (KEY_RIGHTCTRL, KEY_PAUSE, KEY_SCROLLLOCK, …). If you use keyd, the configuration guide shows how to map Caps Lock and Right Alt onto the defaults.

The key listener only needs your user to be in the input group. It never grabs the keyboard, so the key keeps working for other apps too.

Open the Models page. The Recommended for your computer section looks at your processor, memory and graphics card and suggests one model per tier:

TierWhat it doesNeeded for
Speech recognitionturns your voice into texteverything you say
Decision modelchooses among prepared actions when no rule matchesfuzzier phrasing (“pull up youtube”)
Screen visionfinds a described element on a screenshot“click the search box” when accessibility fails

Nothing is downloaded without you. Start with speech recognition: without it, the assistant key has nothing to transcribe. The decision and vision models are optional and can be added later.

Press Get it next to a recommendation, or paste a source such as hf:org/name into Add a model. Downloads resume if interrupted and are checked against their SHA-256 digest. The Models guide lists what is recommended for each GPU class and how the model store works.

Hold the assistant key, say “open youtube”, release. A browser opens or an existing YouTube tab is focused. Then try:

  • “close this”, “fullscreen”, “workspace 3”, “focus right” (window and workspace control)
  • “pause”, “next track” (any MPRIS media player)
  • “new tab”, “find” (the focused app’s own shortcuts)
  • “click the search box” (accessibility first, then vision if you installed it)

Hold the dictation key and talk to type into the focused field. Dictation never runs commands.

To see what Utter would do without doing it, use the dry run from a source checkout:

Terminal window
python3 -m utter.daemon --text "open youtube" --dry-run

Say “go to sleep” in assistant mode and Utter frees your graphics card. The AI model services stop, the speech model is unloaded, and only a tiny listener stays alive. Hold either push-to-talk key to wake it: speech is back in a fraction of a second, and the bigger models reload in the background.

Utter also falls asleep by itself after 15 minutes without being used. Only Utter activity counts (a push-to-talk key, a spoken command, waking up), never typing or clicking in other apps, so the graphics card is freed while you work. Waking is the same key press.

The phrase and the timer are yours to choose, on the General page or in the config file:

[sleep]
enabled = true
trigger = ["go to sleep", "take a break"]
services = ["utter-vision", "utter-planner"] # what gets unloaded
unload_speech = true
on_idle = true # sleep by itself when unused…
idle_minutes = 15 # …after this long
  • Configuration for every config section, speech backends, the decision head, vision and audio.
  • Apps and actions to edit an app’s shortcuts or teach Utter a new application.
  • Trust and safety before you enable terminal commands or raw input.