OpenAI just gave its desktop app a voice. On Thursday, the company updated the ChatGPT desktop app to support ChatGPT Voice, letting you talk to the app to control AI agents and get things done on your computer, according to TechCrunch AI. This is the same tech behind OpenAI’s new ChatGPT-Live voice models, which launched earlier this month. What stands out here is the shift from talking to a chatbot toward talking to something that actually does the work for you.
Here’s what the desktop update brings to the table.
- It talks and takes action. The smartphone version of ChatGPT Voice launched with smoother conversations and better interruption handling, but it wasn’t built to act on your phone. TechCrunch AI reports the desktop version goes further, letting you dictate complex, multi-step commands and respond when ChatGPT needs your input mid-task.
- It plugs into Work and Codex. ChatGPT Voice works with both ChatGPT Work and Codex, OpenAI’s coding agent. It can also tap computer-use skills to look up websites and apps on its own, so the voice command becomes the whole workflow rather than the first step of one.
- It can see your screen on macOS. With a feature called Appshots, Mac users can let the app access what’s on their screen, including alt-text. That screen awareness is what lets a spoken command connect to whatever you’re actually looking at.
- One command, many steps. In a demo video, OpenAI showed a developer asking ChatGPT to create a new thread, open a pull request, and find the root cause of a bug, all from a single spoken instruction. That’s the pitch in a nutshell: say the goal out loud, let the agent chain the steps.
- You can reach it from your phone. OpenAI says users can run ChatGPT Voice in Codex from the iOS app through remote access. So the heavy lifting happens on your desktop while you kick it off from wherever you are.
Why this matters
Voice used to be a novelty layer on top of a chatbot. Ask a question, hear an answer. This is different. OpenAI is positioning voice as the control surface for agents that operate your machine, which is a much bigger claim. For developers especially, describing a bug fix out loud and having Codex chase it down is a real change in how the work gets triggered.
OpenAI isn’t alone in this race. TechCrunch AI notes that Anthropic has also updated its voice mode for Claude, which can tap the company’s Opus, Sonnet, and Haiku models to complete tasks in apps like Gmail, Calendar, Slack, Notion, and Canva. So the two biggest labs are now both betting that voice plus agents is where assistants are headed. The competition is squarely on who can make spoken commands reliably turn into finished work across your apps.
A few caveats worth keeping in mind. The Appshots screen-access feature is macOS-only for now, so Windows and Linux users don’t get that piece yet. And the original article doesn’t spell out pricing or which subscription tiers unlock the desktop voice features, so it’s not clear whether this is available to free users or gated behind paid plans. If you’re deciding whether to lean on it, that’s the detail to watch for.
The direction is clear either way. We’re moving from typing prompts to speaking intentions, and from assistants that answer to agents that act. Whether OpenAI or Anthropic nails the reliability first is the open question. You can find the full details at the original source.