Setting it up
1. Give it a hotkey. Open Settings, then Hotkeys under App, and set the Voice Agent Hotkey. It’s empty by default, so until you set one the only way to reach the agent is by saying its name. 2. Check it’s enabled. Open Settings, choose Language Models under AI Models, then the Voice Agent tab. Enable voice agent is on by default.Using it
Press the hotkey and speak. You don’t need to say the agent’s name first — the hotkey already means “this is an instruction”. The app’s examples:- “Translate to Italian: I talk faster than I type”
- “Write a short email to my boss telling him I want a promotion, his name is Mark”
- “What is 5 × 5?”
Editing text you’ve already written
Select some text first, then press the hotkey and say what you want changed — “make this more formal”, “turn this into bullet points”, “fix the typos”. The agent rewrites the selection in place instead of adding something new below it. With nothing selected, the same hotkey behaves as described above: the result is inserted at your cursor.Sharing your screen as context
The agent can also look at what’s on your screen, so a command can refer to what you’re looking at — “reply to this email”, “explain the error on screen”, “summarize this page”. This is off by default. To turn it on, open Settings, choose Language Models under AI Models, then the Voice Agent tab, and enable Share screen context. On macOS you’ll be asked for Screen Recording permission the first time. When it’s on, pressing the voice agent hotkey captures the display your cursor is on and sends that image with your command. A few things worth knowing:- The screenshot is used for that one request only. It’s never saved to disk, never added to your notes or history, and never written to logs.
- The dictation panel itself is excluded from the capture.
- Only the voice agent hotkey captures. Ordinary dictation and saying your agent’s name never do.
- Not every model can read images. If yours can’t, the command still runs — just without the screenshot.
Screen context isn’t available on Linux under Wayland, which doesn’t allow an
app to capture the screen without a system prompt each time. The toggle is
disabled there and the agent works normally without it.
Choosing where it runs
On the Voice Agent tab you choose the model, independently of the ones used for transcription and cleanup:
Cloud and self-hosted modes work without naming a model. The others need one
chosen explicitly, and the agent won’t run until you do.