Speak the request, and the work runs in the background — this is now the default
On September 23 (local time), OpenAI said on its official X account that ChatGPT's voice mode can now use plugins such as email, calendar and Slack. A plugin is a part that connects an outside service to an AI assistant. Voice mode can also run on GPT-6 Astra, Sol and Luna, and in ChatGPT Work on web and mobile you can create documents, slides, sites and spreadsheets or hand off browser tasks just by talking. It is rolling out globally in the latest version of the app.
The developer side points the same way. According to OpenAI's developer documentation, the voice model GPT-Live handles only the conversation and hands actual work, such as looking up information or using tools, to an agent behind it. An agent is an AI program that takes a goal and works through several steps on its own. Because it can listen while speaking, the conversation keeps going even if the user adds a detail while an order status is being checked. Voice sessions are billed per second, and the backend model is billed separately.
Meta said at Connect 2026 on September 23 (local time) that it will bring its personal agent Muse to its AI glasses in the coming months. Say its name while wearing the glasses and it can act on a product or flier in front of you; in voice mode it keeps working in the background while you talk. For shopping it connects to Walmart, Best Buy, Sephora and others plus Shop Pay and PayPal; for work, to Notion, GitHub and Box. Muse also gets its own email address.
Google's Live Avatar likewise calls tools asynchronously in the background while the conversation continues. All four companies have split the speaking window from the working engine. The user talks; the agent does the work behind the scenes.
Making a voice or a face got cheaper

On September 23 (local time), Google released the speech synthesis model Gemini 3.8 Flash TTS and Flash-Lite TTS for high-volume work. TTS is technology that reads text aloud. You can pick from more than 2,000 ready-made voices or describe a new one in words, and attach acting cues such as emotion, pace or laughter to each line. Google said it ranked first on Hume AI's voice design evaluation with 71.4.
The part to watch is replication. It clones a voice from a 30-second recording, but only when a consent recording made by the voice owner matches the reference voice. Generated audio carries a SynthID watermark that human ears cannot hear, plus C2PA provenance information. Voice replication in AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland or India.
The next day, September 24 (local time), Live Avatar, which gives Gemini 3.8 Live a talking face, arrived in Gemini Enterprise. It matches lip movement and expression while switching across 97 languages, and custom avatars built from a reference image are open only to allowlisted enterprises. On the same 23rd, Google let anyone with a Google account use the new video model Omni 1.1 Flash in Google Vids to make 1080p video at no cost. Making more requires a paid plan.
YouTube, in the feature list from its Made on YouTube event, previewed Auto Dubbing for Live, which turns a live stream into speech in other languages in real time. It opens in limited form early next year, starting with English and Spanish. The same list includes Likeness detection, which finds and removes AI videos that use your face and voice without permission. Tools for making voices and faces and tools for protecting them arrived in the same week.
At the office: Copilot builds apps and waits for you

On September 25 (local time), Microsoft announced on its official blog that Copilot gets Home, Code and Autopilot. Home brings quick conversation (Chat) and end-to-end delegated work (Cowork) into one place and lets you create and edit Word, Excel and PowerPoint documents right there.
Code builds small apps such as dashboards, trackers and automations from a spoken or written description. It uses the same technology as GitHub Copilot, runs in an isolated environment, and can be hosted inside the company tenant — the private space each company gets. Autopilot, an agent previously called Scout, takes a name, a role and a goal, then watches channels, follows up and picks work back up days later.
The pricing structure changed with it. Everyday answers, drafts and summaries run within a fixed per-user fee, while Cowork, Code, Autopilot and top models such as Astra and Fable run on usage-based billing. Home and Code roll out through the early-access Frontier program in the coming weeks, and Autopilot goes to private preview at the end of the month. The longer an agent works, the longer its costs accumulate.
Devices and infrastructure: glasses become the entrance, and power is sought in space
At Connect 2026, Meta unveiled Meta VR Glasses, weighing about 100 grams. Meta compared the weight to a deck of cards and said it is the first IMAX Enhanced certified VR device, with more than 100 immersive live sports events a year. Price and release date were not in the announcement. Ray-Ban Meta Audio, with earbud features, and Ray-Ban Meta (Gen 3) were also introduced.
Meta's AI glasses are getting FDA-cleared hearing enhancement software. Adults who perceive mild to moderate hearing loss can set it up at home in minutes without a clinic visit or prescription, and it launches in the US later this year for $149.99 as an add-on or with a Meta One subscription. Meta said it will have more than 100 styles of AI glasses by the end of the year.
On September 24 (local time), Google said it will send a test satellite carrying its TPU AI chips up on SpaceX's Transporter-18 rideshare mission. It is the first experiment of Project Suncatcher, which asks whether AI compute could be placed in space, and it was built with the satellite company Planet. Satellites in low Earth orbit can generate up to eight times more solar power than on Earth. In ground radiation tests, Trillium TPUs withstood more radiation than a five-year mission would deliver, and in 2027 Google will test linking two satellites by laser.
In recommendations and judgments too, words became the input. On September 23 (local time), Spotify opened Taste Profile in beta to Premium listeners aged 18 and over in the US; it shows a summary of your taste and lets you correct it in writing. On September 18, TechCrunch reported that Jev, a decision model that returns probabilities instead of sentences, is drawing developer interest, citing Vercel getting results five to 18 times faster after switching a command safety classifier to Jev.
Three things users should check
First, the permissions you opened to voice. The moment you send mail or change a schedule by speaking, one misheard word becomes an action. Check which plugins are connected to voice mode and whether confirmation before sending is turned on.
Second, consent for voices and faces. Google required a consent recording from the owner for replication, and YouTube released a tool to find and remove unauthorized use. If you plan to use an executive's or employee's voice in company promotion or training videos, first keep a record of who consented and when.
Third, where charges start. Microsoft moved long-running agent work to usage-based billing, and OpenAI's GPT-Live bills voice by the second and the backend model separately. Before adopting, test with a small month of usage and set a budget cap first.
| Check | Evidence from this week's announcements | First step |
|---|---|---|
| Permissions opened to voice | Email, calendar and Slack plugins in ChatGPT voice mode; Muse on glasses | Review connected plugins and confirm-before-action settings |
| Consent for voice and face | Gemini TTS consent recording; YouTube Likeness detection | Record who consented and when |
| Pay-as-you-go charges | Copilot Cowork, Code, Autopilot; GPT-Live per-second billing | Small-scale trial and a budget cap |
