Zum Inhalt springen

macOS (experimental)

Dieser Inhalt ist noch nicht in deiner Sprache verfügbar.

Utter is a Linux-first project. A macOS port exists behind a small platform switch, using native Apple frameworks instead of the Wayland tooling. It is experimental and has not yet been run on real Apple hardware: the code was written and unit-tested on Linux with the Apple frameworks stubbed out. Expect rough edges and please report what breaks.

ConcernmacOS backend
Push-to-talkQuartz CGEventTap (PyObjC), pynput fallback
Speech-to-textApple Speech.framework on-device with pause chunking (> 50s), then local whisper.cpp fallback
Spoken repliessay or AVSpeechSynthesizer
Decision router / LLMOllama (http://127.0.0.1:11434/v1, Metal) default; LM Studio / llama.cpp
Vision / GroundingOllama VLM (llama3.2-vision:11b, Metal) default; LM Studio
Typing and key chordsQuartz CGEventPost, AppleScript fallback
Focused app and windowNSWorkspace + Accessibility kAXRaiseAction for per-window raise
Screenshots, clipboard, soundsscreencapture, pbpaste, afplay
Background servicelaunchd user agents (com.utter.assistant, com.utter.runner)
HomebrewFormula/utter.rb (CLI + service) and Casks/utter.rb (app)

On Linux, Utter runs background services assuming NVIDIA GPUs and CUDA. Macs have Apple Silicon with Metal and unified memory. Utter integrates with ready-made macOS tools rather than shipping a bespoke inference server:

  • LLM / Decision Router: Ollama (Metal-accelerated out of the box, standard OpenAI-compatible /v1 endpoint on port 11434). LM Studio and llama.cpp (llama-server) are supported drop-ins.
  • Vision / Screen Grounding: Ollama hosting multimodal models (e.g. llama3.2-vision:11b).
  • STT: Apple Speech.framework (zero-model, on-device) or whisper.cpp (Metal).
  • TTS: macOS say.
Terminal window
brew install ollama
brew services start ollama
# Pull recommended router and vision models
ollama pull qwen2.5:3b
ollama pull llama3.2-vision:11b

Utter probes the local runtime before requests and reports status in utter doctor and utter capabilities. If Ollama is stopped or a model is missing, Utter returns a clean, structured status rather than hanging or crashing.

Why not Voca? VocaHQ’s macOS app is VocaMac: a compiled Swift menu-bar app with no socket or API for live transcripts, so it cannot be driven from Python. Utter is standalone and therefore uses Apple’s own speech recognition with a whisper.cpp fallback. If you prefer VocaMac’s models, stt_backend = "vocamac" drives its file-transcription CLI (untested).

  1. Download utter-gui_<ver>_aarch64.dmg (Apple Silicon) or utter-gui_<ver>_x86_64.dmg (Intel) from the release page and drag utter to Applications. First launch: clear Gatekeeper once via xattr -cr /Applications/utter.app (or right-click → Open on macOS < 15).
  2. Open it. The Set up page unpacks the Python runtime and the assistant that ship inside the app into ~/Library/Application Support/utter/, starts the two launchd agents and then walks you through the permissions. No Homebrew, no Python, no terminal.

Updates work the same way: drop in the new app and Set up offers Update.

  • Install CLI + launchd service: brew install --build-from-source Formula/utter.rb
  • Start background daemon: brew services start utter
  • Install GUI app: brew install --cask Casks/utter.rb

Developers can run from a checkout instead: macos/setup.sh creates .venv-macos, installs the [macos] extras and the launchd agents. The Linux installer (install.sh) refuses to run on macOS and points you here.

On a Mac the settings app opens on Set up the first time. Like Raycast’s onboarding it lists each permission with a one-line reason, a Grant access button that triggers the macOS prompt, an Open System Settings button that jumps to the exact pane, and a status badge that re-checks live until everything is granted. The two launchd agents (plugin runner and voice assistant) are shown underneath with a Start button.

The prompts come from the assistant’s own Python, not from the app window, because macOS grants permissions to the process that asks. That is the name you will see in System Settings. The same check works from a terminal:

Terminal window
.venv-macos/bin/python -m assistant macos-permissions --request all

The rest of the app adapts too: Voice edits the [macos] keys and engines, Spoken replies offers the system voices, General controls the launchd agents, and Troubleshooting tails ~/Library/Logs/utter/.

[macos]
stt_backend = "apple_speech" # apple_speech | whisper_cpp | faster_whisper | vocamac
stt_fallback = "whisper_cpp"
speech_locale = "en-US" # fallback when [stt] language does not resolve
on_device_only = true # never send audio to Apple's servers
tts_backend = "say" # say | avspeech | none
tts_voice = "" # `say -v ?` lists voices
tts_rate = 0
hotkey_backend = "quartz" # quartz | pynput
dictation_key = "right_option" # transcript is typed
assistant_key = "right_command" # transcript is run as an action
injection = "quartz" # quartz | applescript
notifications = true
[macos.runtime]
provider = "ollama" # ollama | lm_studio | llama_cpp | custom
base_url = "http://127.0.0.1:11434/v1"
model = "qwen2.5:3b"
vision_provider = "ollama" # ollama | lm_studio | llama_cpp | custom
vision_base_url = "http://127.0.0.1:11434/v1"
vision_model = "llama3.2-vision:11b"

Linux ignores this section entirely. On macOS, [router] and [vision] automatically resolve to the Metal-native endpoints configured in [macos.runtime] (defaulting to local Ollama) rather than the Linux CUDA/vLLM endpoints.

  • Verified by CI: automated matrix packaging of both Apple Silicon (aarch64) and Intel (x86_64) .dmg/.app bundles, fail-safe codesigning/notarization, platform detection tests.
  • Implemented against documented APIs (untested on real hardware): push-to-talk with two keys, Metal runtime resolution and graceful degradation, native speech recognition with pause-aware audio chunking (> 50s) and whisper.cpp runtime fallback, dictation typing, app and URL launching, window focus with AX window raising (kAXRaiseAction), clipboard, spoken replies, notifications, screenshots for the vision tier with Retina geometry handling, the plugin runner socket with peer credentials (LOCAL_PEERCRED / LOCAL_PEERPID and proc_pidpath), Homebrew formulas (Formula/utter.rb and Casks/utter.rb).
  • Linux-only: niri compositor actions and workspace-aware focusing, AT-SPI accessibility clicks (macOS goes straight to vision), MPRIS media keys, the Noctalia widget and OSD, the sandbox wrapper, the Linux installer wizard, and the settings app’s service controls (they call systemctl).
  • Needs real Mac hardware to confirm: physical audio input, hardware key-tap edge detection, actual AX synthetic keystrokes, real Screen Recording permission capture, and physical Metal GPU inference under load with Ollama.

The complete matrix, with the per-module status, is in docs/MACOS.md.