
Summary
> "AI you can talk to like a phone call" has finally arrived.
"AI you can talk to like a phone call" has finally arrived.
A voice AI that can hold a conversation just like a human — you can interrupt, give backchannel cues, or stay silent. GPT-Live, released by OpenAI on July 8, 2026, has evolved voice AI from "walkie-talkie" to "phone call."
What You'll Learn in This Article
- What GPT-Live is — a 3-minute overview
- How it differs from the previous Advanced Voice Mode (AVM)
- How the full-duplex architecture works
- New experiences enabled by delegating reasoning to GPT-5.5
- Feature comparison by pricing plan
- Japanese language support status and actual quality
- How to get the most out of GPT-Live
What Is GPT-Live?
GPT-Live is OpenAI's third-generation voice architecture. It's a new-generation product that completely replaces ChatGPT's voice mode.
Its defining feature is the adoption of the full-duplex method. This means you can "talk while listening" — fundamentally different from the traditional approach of "waiting for the other person to finish speaking."
Three generations of ChatGPT voice mode:
| Generation | Name | Release | Method |
|---|---|---|---|
| 1st Gen | Cascade Method | 2023 | Sequential processing of STT→LLM→TTS. Slow and unnatural |
| 2nd Gen | Advanced Voice Mode | Sep 2024 | Single model, detects turn-ending via silence |
| **3rd Gen** | **GPT-Live** | **Jul 2026** | **Full-duplex. Can talk while listening** |
With the arrival of GPT-Live, the previous Advanced Voice Mode (AVM) will no longer be the default voice mode. After updating, users automatically migrate to GPT-Live.
Comparison with Advanced Voice Mode (AVM)
| Feature | AVM (Old) | GPT-Live (New) |
|---|---|---|
| Architecture | Half-duplex (turn-based) | Full-duplex (simultaneous bidirectional) |
| Turn-taking | Detects silence, then responds | Natural interruptions, backchannels, and pauses |
| Web Search | Not available during voice sessions | Delegates to GPT-5.5 for concurrent execution |
| Interruptions | May break the session | Mid-sentence interruptions handled naturally |
| Visual Output | None | Displays weather, stocks, images, etc. alongside voice |
| Real-time Translation | None | Supported (simultaneous interpretation) |
| BrowseComp Accuracy | 0.7% | 75.2% |
Key differences worth noting:
The reason BrowseComp jumped from 0.7% to 75.2% is that GPT-Live separated the voice layer from the reasoning layer. Voice interactions are handled by a lightweight model, while complex web searches and reasoning are delegated to GPT-5.5. This makes "researching while talking" possible.
How the Full-Duplex Architecture Works
The technical core of GPT-Live lies in making conversation decisions multiple times per second.
Traditional voice AI was a serial pipeline of "input → process → output." GPT-Live constantly processes input and output simultaneously, making decisions — "speak," "listen," "pause," "interrupt," "call a tool" — many times per second.
What it can do in practice:
- Give backchannel cues — responds with "uh-huh," "I see" while the user is still speaking
- Hold silence — waits silently while the user is thinking
- Mid-sentence interruption — stops immediately when the user says "hold on"
- Research in the background — when asked "check the current stock price," delegates the search to GPT-5.5 while continuing the conversation
OpenAI's Head of Product, Atty Eleti, cited a 30-40 minute conversation during a walk as an example, saying "you no longer need to take turns speaking."
Feature Comparison by Pricing Plan
GPT-Live has two models, and available features vary by plan.
| Plan | Monthly Price | Model Used | Available Intelligence Levels |
|---|---|---|---|
| Free | $0 | GPT-Live-1 mini | Instant only |
| Go | ~$8 | GPT-Live-1 | Instant + Medium |
| Plus | $20 | GPT-Live-1 | Instant + Medium |
| Pro | $100–$200 | GPT-Live-1 | Instant + Medium + High |
| Team | Enterprise | GPT-Live-1 | Full access |
Differences between Intelligence levels:
- Instant — GPT-5.5 Instant. Fast responses. Everyday conversation level
- Medium — GPT-5.5 Thinking (medium reasoning effort). For complex questions
- High — GPT-5.5 Thinking (high effort). Pro plan only. Highest quality reasoning
Intelligence level is selectable from Settings → Voice → Intelligence. However, GPT-Live-1 mini (Free users) is locked to Instant.
Japanese Language Support Status
OpenAI has announced that GPT-Live supports "most major languages," but a specific language list has not been published.
The Hindi simultaneous interpretation demo shown at the briefing was rated as having "a strong accent and slightly unnatural," so it's best to assume that non-English quality is still a work in progress.
Since Japan is a major market, there's a good chance support will be strengthened early on, but for now, the realistic stance is: "English is best, Japanese probably works but isn't perfect."
How to Use GPT-Live and Use Cases
Getting Started
- Update the ChatGPT app to the latest version (iOS / Android / Web)
- Tap the headphone icon voice button
- GPT-Live launches automatically (switches from the old AVM)
- Select Intelligence level from Settings → Voice → Intelligence
Recommended Use Cases
1. Hands-free productivity Gather information and brainstorm ideas just by talking — while commuting or cooking.
2. Research while conversing "Check OpenAI's current stock price" → GPT-Live searches in the background while talking → returns results by voice. The conversation never pauses.
3. Real-time translation Have GPT-Live provide simultaneous interpretation while listening to an English webinar. The quality isn't perfect, but it's sufficient for "getting the gist."
4. Meeting preparation Brainstorm agenda items by voice while on the move → GPT-Live organizes the notes → review as text when you arrive.
Current Limitations
- No video or screen sharing — these still use the legacy voice mode (AVM)
- API not yet available — as of July 2026, ChatGPT app only. API has a notification waitlist
- Uneven non-English quality — languages other than English need testing before production use
Frequently Asked Questions (FAQ)
Q1: Is GPT-Live free to use?
GPT-Live-1 mini is available on the Free plan. However, the Intelligence level is locked to Instant, and advanced delegation to GPT-5.5 is limited to paid plans.
Q2: What happens to Advanced Voice Mode?
The existing AVM is replaced by GPT-Live. However, if you need video or screen sharing, you can still select the legacy voice mode from settings.
Q3: How good is the Japanese quality?
OpenAI has announced support for "most major languages," but specific Japanese quality evaluations have not been published. Consider non-English quality to still be a work in progress.
Q4: Can I use it via API?
As of July 2026, the API is not yet available. Developers can sign up for notifications.
Q5: Which plan do you recommend?
For everyday voice conversations, Plus ($20/month) is sufficient. If you need advanced reasoning (High effort), consider Pro ($100–$200/month).
Q6: How is GPT-Live different from GPT-5.6?
GPT-Live is a voice interaction engine, while GPT-5.6 is a text-based foundation model. GPT-Live uses GPT-5.5 (one generation before GPT-5.6) as its reasoning engine behind the scenes.
Q7: Is there a time limit on calls?
No time limit has been announced for GPT-Live itself. Existing per-plan message limits and rate limits apply.
Q8: Can I go back to the legacy voice mode?
You can select the previous Advanced Voice Mode from settings. However, GPT-Live becomes the default.
Summary
GPT-Live is OpenAI's first step toward a long-term vision of making "voice the primary interface for computing."
Key takeaways from this article:
- Full-duplex method enables natural conversation where you can "talk while listening"
- Separation of voice and reasoning layers dramatically improves BrowseComp from 0.7% to 75.2%
- GPT-Live-1 mini is available on Free, but advanced reasoning is limited to paid plans
- Japanese support is at a "probably works" level. Expect quality improvements ahead
- If you want to experience the future of voice AI, update ChatGPT now and give it a try
Voice AI has evolved from "walkie-talkie" to "phone call." The next step is "face-to-face conversation" — and OpenAI is steadily moving in that direction.
This article is based on information as of July 10, 2026. For the latest details on GPT-Live, check the OpenAI official site.
Related Reading
- Claude Fable 5 Financial Asset Protection Guide 2026: AI Agent for Asset Management, Monitoring & Optimization
- Cloudflare Monetization Gateway Complete Guide 2026: How to Add One-Time Billing to Web Pages, APIs & MCP Tools
- A Fable of Codexes Complete Guide 2026: How to Build a Claude-Commanded AI Worker Army
- How to Drastically Improve AI UI Generation Instructions Using component.gallery 2026: A Practical Guide to the Component Terminology Encyclopedia
- ChatGPT vs Claude vs Gemini 2026 Full Comparison: Complete Guide from Free to Paid Plans
この記事をシェアする
Related articles

2026年7月19日
[2026] How to Dramatically Improve AI UI Generation with component.gallery! A Practical Guide to the Component Terminology Encyclopedia

2026年6月15日
ChatGPT vs Claude vs Gemini 2026: Ultimate Comparison! From Free to Paid — Complete Guide

2026年6月18日
Free AI Models Guide 2026: 8 Ways to Use Claude Opus 4.8, GPT-5.5 & Gemini 2.5 Pro for $0

2026年6月18日
Accio Work Complete Guide 2026: Alibaba-Partnered AI Agent Automates Sourcing, Store Building, and Sales

2026年6月19日
【2026】Ollama Complete Setup Guide: Running Local AI on a Mini PC

2026年6月23日
Blueprint.am Complete Guide 2026: "Claude for Hardware" Auto-Generates Wiring Diagrams, BOMs, and Assembly Instructions