Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, calling them its most advanced audio models yet.

The two models talk, reason, and handle tasks. But how do they rank with independent analysts? The Extended Thinking version topped Artificial Analysis’ Speech-to-Speech Index at 82.6, ahead of GPT-Live-1 and Grok Voice.

Follow us on X to get the latest news as it happens

Say “hi” to our most advanced audio models from @GoogleDeepMind yet built for natural, production-ready voice applications.

🔷 Gemini 3.8 Live
🔷 Gemini 3.8 Live Extended Thinking

With these models, you can speak naturally, collaborate easily, and tackle complex tasks using… pic.twitter.com/TCcuANcIWl

— Google (@Google) September 15, 2026

How the Benchmarks Rank Google’s Newest Conversational AI Model

Artificial Analysis’s Index averages speech reasoning, agentic performance, arena preference, and task success rate.

Gemini 3.8 Live Extended Thinking, tested at high reasoning effort, debuted in first place. GPT-Live-1 Astra followed at 81.5 and Grok Voice Think Fast 2.0 High at 81.3.

How Gemini 3.8 Live Ranked on The Speech-to-Speech Index
How Gemini 3.8 Live Ranked on The Speech-to-Speech Index. Source: Artificial Analysis

The standard Gemini 3.8 Live placed fifth with a score of 76.0. Both variants beat Gemini 3.1 Flash Live High, which scored 71.5.

The agentic gap is wider. Extended Thinking reached 68.6% on the Tau Voice benchmark, against 37.7% for the previous generation.

On the speech reasoning benchmark, Extended Thinking scored 97.7%, edging Grok Voice at 97.2%. However, it trailed Qwen Audio 3.0 Realtime Plus, which scored 99.2%.

Human Testers Still Reach for the Older Gemini Model

Price is where the distance opens up. The standard model runs $0.84 per hour of input audio. That is roughly half its predecessor’s $1.75 and the Index’s cheapest rate. 

Extended Thinking runs at $3.50 per hour. That undercuts GPT-Live-1 Sol at $4.47 and Grok Voice Think Fast 2.0 High at $4.80.

Latency fell as well. Average time to first audio dropped to 1.18 seconds, compared with 2.99 seconds for the older Gemini model.

Listeners, however, are not fully sold. Gemini 3.1 Flash Live still leads in preference in blind Speech Agent Arena conversations, with an Elo of 1096.

Gemini 3.8 Live sits second at 1083. The Extended Thinking variant trails at 990, despite completing 89.1% of its tasks.

Voice AI is now being graded on two scales that point in different directions. Benchmarks reward the model that reasons hardest. Preference rewards the one who talks best. Google is currently leading both, with different models.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

The post Google Just Released Its Most Advanced Audio Model. Here Is How It Ranks appeared first on BeInCrypto.

Read Original Source