Speech to text for Wayland Linux and Windows
Press a key. Speak. It’s already typed.
A daemon-friendly dictation CLI. Realtime transcription lands straight in the focused window — no GUI, no cloud account, no waiting for a model to load.
How it works
One key, held open by a warm daemon
Bind a key
One shortcut, registered system-wide.
Speak
A warm daemon holds the model and socket. No cold start.
Press again
The same key stops the take and finalises it.
Text lands
Typed, pasted, copied, or written to stdout.
Providers
Pick your trade-off, not ours
| Provider | Model | Realtime | Audio |
|---|---|---|---|
| Mistral | voxtral-mini-transcribe-realtime | streaming | cloud |
| Deepgram | nova-3 | streaming | cloud |
| Groq | whisper-large-v3-turbo | on pause | cloud |
| Whisper | ggml, on-device | on pause | on device |
One dictation: speak, stop, and the take is inserted once — polished, with self-corrections and spoken formatting applied. Clipboard commands run on the same shortcut when you copied text first.
Details
A UNIX citizen that happens to hear
Stdout first
--pipe-to hands text to any command — wl-copy, ydotool, sed, your own script.
Two platforms, one config
Same keys on both. PipeWire/ydotool on Wayland, WASAPI/Win32 on Windows.
Idle costs nothing
The daemon sleeps until a shortcut arrives. No polling, no background work.
Polish is separate
A text model cleans up after recognition — OpenCode Zen, Mistral, or local Ollama.
Your vocabulary
A dictionary of names and mishearings is applied before text reaches the screen.
Fully offline option
Local Whisper keeps audio on the machine. Nothing uploaded, no key needed.
Install
Running in about a minute
dictate doctor checks the whole chain — keys, mic, permissions, daemons.
curl -fsSL https://dictate.adityamer.dev/install.sh | shDetects your distro, installs PipeWire deps, and pulls the latest release binary. Then run dictate setup.
Read https://dictate.adityamer.dev/INSTALL.md and follow it step by step to install and configure dictate on this machine. Ask me the setup questions first, then execute everything non-interactively using 'dictate config set'.
Or hand that prompt to Claude Code, Cursor, Copilot, Windsurf, or Gemini CLI.