Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

OpenAI brings ChatGPT Voice to the desktop app, enabling hands-free agent control

OpenAI on Thursday updated its ChatGPT desktop app with voice mode support, letting users speak commands to control AI agents and perform computer tasks. The feature uses the ChatGPT-Live voice model family launched earlier this month and works with ChatGPT Work and Codex for multi-step operations.

Published

OpenAI on Thursday announced that it has updated its ChatGPT desktop application with ChatGPT Voice support, allowing users to talk to the app to control AI agents and perform tasks on their computers.

The new feature taps OpenAI's new family of voice models called ChatGPT-Live, which the company launched earlier this month. These models are designed for more natural live conversations and provide the underlying voice technology for the desktop experience.

OpenAI said ChatGPT Voice works with both ChatGPT Work and Codex, and can also tap computer use skills to look up websites and apps. On macOS, with Appshots, users can let the voice assistant access what's on their screen, including alt-text.

Unlike the smartphone version, which offered smoother conversations with better interruption handling but was not built to take action on phones, Thursday's desktop update is more capable. It allows users to dictate complex commands involving many steps, and respond when ChatGPT needs their input.

In a demo video, OpenAI showed a developer asking ChatGPT to create a new thread, make a pull request, and find the root cause of a bug — all with a single voice command. The company said users can also utilize ChatGPT Voice in Codex from the iOS app through remote access.

This update marks a significant shift for AI voice interaction, moving from simple Q&A toward genuine computer control. Anthropic recently updated Claude's voice mode as well, enabling it to use Opus, Sonnet, and Haiku models to complete tasks in apps like Gmail, Calendar, Slack, Notion, and Canva.

Voice-controlled agents on the desktop mean users could soon manage development workflows, data analysis tasks, and office applications through natural speech, substantially lowering the barrier between humans and AI interaction. The key question going forward is whether voice-plus-agent combinations will find their first large-scale adoption in enterprise productivity scenarios.

Why it matters

OpenAI's desktop voice mode extends AI voice interaction from conversation to action, enabling agent control through speech. This evolution could significantly lower the adoption barrier for AI agents and shift enterprise AI interaction from text to voice.

OpenAIChatGPTVoice ModeDesktop
Back to realtime news

Nearby Updates

All

07/24, 21:46

Xiaomi new phone passes certification with a dozen AI models including DeepSeek, ERNIE Bot, and Tongyi Qianwen

A Xiaomi phone model 2608BPX34C has passed network access certification in China, filing support for a dozen large language models including DeepSeek, ERNIE Bot, Tongyi Qianwen, Zhipu AI, and Xiaomi's own models. Industry watchers speculate the filings are preparation for the upcoming HyperOS 4 system.

07/24, 21:05

Brown & Brown Partners with Anthropic, McKinsey and Accenture to Build AI Model

Insurance brokerage giant Brown & Brown has partnered with Anthropic, McKinsey, and Accenture to jointly build an AI model for its business. The collaboration signals accelerating enterprise AI adoption in the tightly regulated insurance sector.

07/24, 20:55

Google's Gemini 3.5 Pro Slips Behind Schedule as Compute Constraints Ripple Across the AI Industry

Google's next-generation flagship model Gemini 3.5 Pro has fallen behind schedule due to compute constraints, reflecting a broader infrastructure bottleneck gripping the entire AI industry. The delay highlights how the frontier model arms race is increasingly constrained by hardware availability rather than algorithmic advances.

07/24, 22:19

UniWorld-View Tops Fei-Fei Li Team's World Model Leaderboard, Fully Open-Sourced with Ascend NPU Support

UniWorld-View, developed by Toozhan AI in collaboration with Peking University and Peng Cheng Laboratory, has topped the world model leaderboard maintained by Fei-Fei Li's team. The model supports novel view synthesis from a single image or monocular video, is compatible with domestic Ascend NPUs, and has released all code and weights as open source.