Skip to main content
The voice agent turns speech into an instruction rather than text. Press its hotkey, say what you want, and what appears at your cursor is the result — a drafted email, a rewritten paragraph, an answer. The app sums it up as “Speak a request — your AI agent types the result, not your words.”

Setting it up

1. Give it a hotkey. Open Settings, then Hotkeys under App, and set the Voice Agent Hotkey. It’s empty by default, so until you set one the only way to reach the agent is by saying its name. 2. Check it’s enabled. Open Settings, choose Language Models under AI Models, then the Voice Agent tab. Enable voice agent is on by default.

Using it

Press the hotkey and speak. You don’t need to say the agent’s name first — the hotkey already means “this is an instruction”. The app’s examples:
  • “Translate to Italian: I talk faster than I type”
  • “Write a short email to my boss telling him I want a promotion, his name is Mark”
  • “What is 5 × 5?”
The result is typed wherever your cursor is, exactly like dictation — you can use it inside an email, a document, a chat box.

Editing text you’ve already written

Select some text first, then press the hotkey and say what you want changed — “make this more formal”, “turn this into bullet points”, “fix the typos”. The agent rewrites the selection in place instead of adding something new below it. With nothing selected, the same hotkey behaves as described above: the result is inserted at your cursor.

Sharing your screen as context

The agent can also look at what’s on your screen, so a command can refer to what you’re looking at — “reply to this email”, “explain the error on screen”, “summarize this page”. This is off by default. To turn it on, open Settings, choose Language Models under AI Models, then the Voice Agent tab, and enable Share screen context. On macOS you’ll be asked for Screen Recording permission the first time. When it’s on, pressing the voice agent hotkey captures the display your cursor is on and sends that image with your command. A few things worth knowing:
  • The screenshot is used for that one request only. It’s never saved to disk, never added to your notes or history, and never written to logs.
  • The dictation panel itself is excluded from the capture.
  • Only the voice agent hotkey captures. Ordinary dictation and saying your agent’s name never do.
  • Not every model can read images. If yours can’t, the command still runs — just without the screenshot.
Using a different model for screenshots. Some people want a cheap, fast model for everyday commands and a stronger one when an image is involved. Turn on Separate vision model underneath and pick that second model; it’s used only for commands that carry a screenshot.
Screen context isn’t available on Linux under Wayland, which doesn’t allow an app to capture the screen without a system prompt each time. The toggle is disabled there and the agent works normally without it.

Choosing where it runs

On the Voice Agent tab you choose the model, independently of the ones used for transcription and cleanup: Cloud and self-hosted modes work without naming a model. The others need one chosen explicitly, and the agent won’t run until you do.

Its instructions

Agent prompt, on the same tab, is the system prompt used when the agent runs. Leave it empty and a built-in default is used. Set it if you want a consistent tone or format across everything the agent writes.

If nothing happens

If the agent can’t run — no model chosen in a mode that needs one, or it’s switched off — a dictation started with the voice agent hotkey comes back as the plain transcript, with no cleanup applied. So getting your own raw words back from that shortcut is the signal that the agent isn’t reachable, rather than that it misunderstood you. Check, in order: Enable voice agent is on, and a model is selected if your mode requires one. If the agent was reachable but the request failed, you’ll see an Agent Unavailable notice alongside the raw transcript, so you can tell a genuine failure apart from the agent not being set up. And if you have screen context on but the screenshot couldn’t be sent, you’ll see a Screen Context Skipped notice — the command still ran, just without the image.