Google DeepMind 发布 Gemini 3.8 Live 和 3.8 Live Extended Thinking
原文给出两个语音对话模型的定位分工和多项基准数字,读者可以据此评估其推理、成本与工具调用能力。
Google DeepMind 发布 Gemini 3.8 Live 与 Gemini 3.8 Live Extended Thinking 两个近实时语音对话模型,主打语音智能体和复杂任务执行。
Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking 是我们迄今最先进的实时对话模型。智能与并行推理能力的重大升级,使它们在与用户协作以及通过语音执行复杂任务时更加直观。
Tom Ouyang
Malini Jaganathan
技术团队成员,代表 Gemini Audio 团队
今天,我们推出两款新模型,在近实时推理方面带来进步,从而更有效地赋能语音智能体,并让与 AI 的对话感觉更直观、更智能。
对于开发者和企业而言,这些模型提供了构建可靠、可投入生产的语音智能体的基础模块。它们还让用户在 Gemini 应用、Google Workspace 和 Search 中与 Gemini 对话更加流畅、更具协作性——帮助你仅用语音就能应对复杂任务。
Gemini 3.8 Live Extended Thinking 提供企业级的任务完成能力与智能水平,在 Artificial Analysis 的语音到语音质量指数(Speech to Speech Quality Index)上以 82.6 分拿下总榜第一,并在智能体任务完成方面领先,在 τ-Voice 上达到 68.6%,在 Sierra 的 τ-Voice-banking 基准上达到 35.1%。它还具备强大的推理能力,在 Big Bench Audio 上得分 97.7%,同时相较于其他前沿模型保持了极具竞争力的价格。
Gemini 3.8 Live 在用户中表现出很高的偏好度,在 Speech Agent Arena 中位列第二。除这一表现外,它依然极具成本效益——为开发者和企业提供了一个能力强、效率高、专为规模化而打造的模型。
在 ServiceNow 的 EVA-Bench(一个用于评估语音智能体的基准)上,我们的模型通过成功平衡准确性与对话质量,推动了复杂工作流的帕累托前沿。
注:该测试运行于 Gemini Enterprise Agent Platform 上的 Live API。
Gemini 3.8 Live 以近乎实时的速度处理视觉输入,为对话补充上下文,从而给出更有帮助的回复。它能在对话过程中自动检测并在 97 种支持语言之间切换。它会在后台执行工具调用和 API 调用,同时继续对话,因此模型可以在确认请求后继续聊天,而任务则在后台完成。
Gemini 3.8 Live 实时指导员工入职,利用视觉上下文回答现场提问。
观看 Gemini 3.8 Live 借助视觉上下文、推理和自然的对话流,近乎实时地下棋。
对于需要更深层推理的任务,3.8 Live Extended Thinking 能够边推理边说话。它为复杂工作流带来更强的智能,同时保持不间断的对话流——使用诸如“让我查一下……”这样的早期口头提示来自然地回应提示词,并通过实时进度叙述,让用户随时了解多步骤后台任务的进展。
观看 Gemini 3.8 Live Extended Thinking 将原始草图和近乎实时的语音反馈转化为可用的 React 组件。
看看 Gemini 3.8 Live Extended Thinking 如何协调多步骤预订和异步函数调用——全程不打断自然的实时对话。
观看 Gemini 3.8 Live 通过自然语音即时构建完整的商业计划和定制营销工具包。
在 Google Workspace 和 Search 中,我们的 Live 模型带来更直观、更具协作性的体验——尤其是在处理你最复杂的任务时。
在 Google Workspace 中通过 Docs Live、Gmail Live 和 Keep Live 试用 Gemini 3.8 Live Extended Thinking。
通过使用 Gemini Live API,Agora、Fishjam、LiveKit、Pipecat、Vercel 和 Vision Agents 等开发者平台,开发者可以轻松构建和部署高性能的语音驱动界面。这些平台在后台管理复杂的实时媒体流基础设施,让开发者能够完全专注于打造用户体验。
我们还与 Salesforce、Genspark 和 Lumeris 等公司合作,它们对 3.8 Live 和 3.8 Live Extended Thinking 充满期待,并强调了其令人印象深刻的延迟、流畅性和工具调用能力。
我们的 AI 产品生成的所有音频都使用 SynthID 添加水印。这种不可感知的水印直接编织进音频输出中,确保 AI 生成的内容可被检测,从而帮助防止错误信息传播。有关我们安全与责任方法的详细信息,请查阅模型卡。
3.8 Live 从今天开始逐步推出:
3.8 Live Extended Thinking 从今天开始逐步推出:
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet. Major upgrades in intelligence and parallel reasoning make them more intuitive to collaborate with and use to execute complex tasks using your voice.
Principal Engineer
Member of Technical Staff, on behalf of the Gemini Audio Team
Today, we’re introducing two new models that bring advancements in near real-time reasoning to more effectively enable voice agents and make conversing with AI feel more intuitive and intelligent.
For developers and enterprises, these models deliver the building blocks for reliable, production-ready voice agents. They also make speaking with Gemini across the Gemini app, Google Workspace, and Search more fluid and collaborative — helping you tackle complex tasks using just your voice.
Gemini 3.8 Live Extended Thinking provides enterprise-grade task completion and intelligence, capturing the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index (82.6), and leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. It also provides strong reasoning capabilities, scoring 97.7% on Big Bench Audio, while maintaining a highly competitive price point compared to other frontier models.
Gemini 3.8 Live has shown a high preference among users, securing a second place in the Speech Agent Arena. In addition to this performance, it remains highly cost-effective — providing developers and enterprises with a capable and efficient model built for scale.
On ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents, our models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality.
Note: This was run on the Live API on Gemini Enterprise Agent Platform.
Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses. It automatically detects and transitions between 97 supported languages mid-conversation. It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background.
Gemini 3.8 Live guides employee onboarding in real time, using visual context to answer live questions.
Watch Gemini 3.8 Live play chess in near real-time using visual context, reasoning, and natural conversational flow.
For tasks that require deeper reasoning, 3.8 Live Extended Thinking reasons and speaks simultaneously. It delivers increased intelligence for complex workflows while maintaining an uninterrupted conversational flow — using early verbal cues like “Let me check that…” to acknowledge prompts naturally, and live progress narration to walk users through multi-step background tasks as they progress.
Watch Gemini 3.8 Live Extended Thinking transform raw sketches and near real-time voice feedback into functional React components.
See Gemini 3.8 Live Extended Thinking coordinate multi-step bookings and asynchronous function calls — all without interrupting natural live conversation.
Watch Gemini 3.8 Live build complete business plans and custom marketing toolkits on the fly through natural speech.
Across Google Workspace and Search, our Live models deliver more intuitive, collaborative experiences — especially when tackling your most complex tasks.
Try Gemini 3.8 Live Extended Thinking in Google Workspace with Docs Live, Gmail Live, and Keep Live.
By using the Gemini Live API, developer platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.
We’re also partnering with companies like Salesforce, Genspark, and Lumeris who are excited about 3.8 Live and 3.8 Live Extended Thinking, highlighting its impressive latency, fluidity, and tool-calling capabilities.
All audio generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on our approach to safety and responsibility, review the model card.
3.8 Live is rolling out starting today:
3.8 Live Extended Thinking is rolling out starting today:
来源:Google DeepMind:Blog(RSS)· deepmind.google
配图