返回新闻列表
2026-09-11 03:25
AI123

OpenAI发布AI语音模型GPT-Live-1,客服任务完成率达83.6%

OpenAI于2026年9月11日发布新款语音模型GPT-Live-1,目前已通过API向开发者提供。根据OpenAI公布的评测数据,GPT-Live-1与GPT-6 Astra搭配、在中等等推理强度下,于Tau3基准测试中首次任务完成率达到83.6%,而GPT-Realtime-2.1的成绩为45.7%。

Tau3基准测试用于评估语音智能体在航空、零售和电信等场景中完成客户支持任务的能力。同一组合在TauBanking测试中获得38.1%的得分,该测试要求智能体在银行文档中查找信息,并借助工具解决客户请求。

在对话交互方面,GPT-Live-1在Artificial Analysis的Conversational Dynamics基准测试中得分97.3%,该测试评估模型在对话中判断何时等待、接话、回应打断或通过简短确认继续说话的能力。在Full Duplex Bench v1.5上,模型在重叠语音、背景人声、旁白和打断等场景中的交互性得分为80.1%。Full Duplex Bench v1显示,GPT-Live-1在用户讲话结束后平均0.798秒开始回应,而GPT-Realtime-2.1为1.41秒。

在工具调用层面,由后端模型处理工具调用的GPT-Live-1方案在Full Duplex Bench v3上实现了87%的工具调用成功率,并在独立的响应质量评估中获得90%。OpenAI已在发布博客中提供更多演示和说明。

来源
  1. 2026-09-11 02:35 | X:Re For the GPT-Live-1's intelligence, Tau3 tests whether a voice agent can actually complete customer-support tasks across airline, retail, and teleco...
    阅读原文

    Re For the GPT-Live-1's intelligence, Tau3 tests whether a voice agent can actually complete customer-support tasks across airline, retail, and telecom scenarios.Paired with GPT-6 Astra at medium reasoning effort, GPT-Live-1 completed 83.6% of tasks on the first attempt, compared with 45.7% for GPT-...

其他新闻

查看全部