Back to the edition
AIAnalysis

Can an AI get things done while you’re still talking to it?

Google’s new Gemini Live models bring voice, camera context and background work together. The interesting test is what happens after the assistant says it can help.

An illustrated phone’s copper voice waveform links a calendar, a document and a doorway of light.
A conversation becomes a path toward completed work. AI-generated conceptual illustration by The Daybreak—not a Gemini screenshot, product demonstration or reported event. The Daybreak / Original Daybreak graphic
Image details

Displayed without generative alteration.

The revealing moment in a conversation with an AI is often the pause. You have asked it to do something. It has understood enough to answer. Now it has to go away, check another system and come back with something more useful than a reassuring sentence.

Google’s latest voice models are built to make that gap less awkward. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, announced September 15 in a post updated September 17, can use visual context and keep a conversation moving while work happens in the background, Google says. Its published demonstrations include coordinating bookings and turning sketches plus spoken feedback into React components. These are company demonstrations, not tests conducted by The Daybreak. Google’s announcement[1]

The attraction is easy to understand. Instead of translating a messy problem into a perfectly formed prompt, you could explain, point, interrupt and revise. An assistant that can stay with that conversation while checking the relevant information would remove some of the work of operating software.

But two separate achievements sit inside that promise. One is sounding like a capable collaborator. The other is completing the right action. Independent measurements of the new Gemini models offer a particularly clear reason to keep both in view.

Close view of the rear camera assembly on a Google Pixel phone.
A Pixel phone’s camera, photographed September 4, 2024. File photograph; not a demonstration of Gemini 3.8. Camera context is one part of Google’s voice-assistant ambitions. Kyu3a / Wikimedia Commons · resized · CC BY-SA 4.0 · Source. Kyu3a / CC BY-SA 4.0
Image details

Displayed without generative alteration.

Where you can actually use it

Google says standard 3.8 Live is rolling out through Search Live and its developer tools. Extended Thinking is rolling out through Gemini Live and the developer tools, with different subscription requirements across Workspace products. Enterprise access includes private previews and features described as coming soon. A model announcement therefore does not mean every account has every demonstration available. Rollout details[1]

For a reader who wants a concrete place to start, Search Live is the clearer path. Google’s support documentation says it works in the Google app on Android and iOS, in regions and languages where AI Mode is available. It supports spoken follow-up questions, web links and optional camera sharing. The camera stops when the app goes into the background or the phone locks, even if the voice conversation continues. Search Live help[2]

That last distinction is useful. A voice session continuing in another app is not the same thing as the assistant continuing to see your surroundings. If you move the phone away from a label or close the camera, the conversation has lost that live source of information. A good response should reflect the change rather than act as though the view is still present.

This is also why a model’s capabilities and a finished product’s capabilities must be kept separate. A developer can connect a voice model to a scheduling service. That does not establish that a consumer’s Search Live session can make the same booking. The application has to provide the connection, the relevant access and the rules for using it.

Watch Google’s demonstration of a conversation continuing during restaurant-booking tasks. This is a vendor demonstration, not an independent test.[3]

What happens behind the voice

Google’s Live API documentation describes a continuous connection that accepts sound, images and text and returns spoken responses. Developers can build around that connection rather than assemble every interaction as a fresh text exchange. The documentation lists interruptions, transcripts and tool use among its supported features. Live API overview[4]

The important part for getting something done is the tool. In Google’s developer guide, the application defines functions the model may request, then handles their results. A request to turn on a light, for instance, is represented by a function call; the application has to report what happened. The model’s speech is only one part of the chain. Tool-use documentation[5]

The same guide explains how background functions can avoid stopping the conversation. Their results can interrupt the dialogue, wait for a quiet moment or become information the assistant uses later. Support varies by model, and some tables in the guide still describe earlier Live versions. The principle is nevertheless straightforward: conversation and external work do not have to take turns occupying the whole interaction. Background functions[5]

Two parallel lanes show a conversation continuing while a background task checks availability, requests an action and returns evidence.
An illustrative workflow by The Daybreak, not a recording or product screenshot. The useful connection is between what the assistant says and what the external service confirms. The Daybreak / Original editorial graphic
Image details

Displayed without generative alteration.

Consider a hypothetical appointment change. You ask for Thursday afternoon. While the scheduling service checks, you remember that you cannot arrive before three. A useful assistant incorporates the correction, finds an appropriate opening and confirms the result. If the original request has already gone through, it needs to say so and resolve the conflict.

That is where the apparently small advance becomes consequential. The assistant must remember which version of your request is current, distinguish a search from a reservation and avoid describing a pending action as completed. Filling the silence is the easiest part to notice. Maintaining the correct state of the task is the part that determines whether you have an appointment.

The stronger thinker is not better at every part

Artificial Analysis’s current measurements show Gemini 3.8 Live Extended Thinking, in its High configuration, completing 68.6% of the simulated customer-service tasks in its τ-Voice evaluation. Standard 3.8 Live scores 30.1%. On a separate conversational-dynamics measure, the order reverses: standard Live scores 96.1%, versus 91.9% for Extended Thinking. These are separate tests, observed September 21, not predictions of any individual customer’s experience. Independent results[6]

Bar chart comparing the two Gemini models: Extended Thinking scores higher on simulated task completion, while standard Live scores higher on conversational dynamics.
The Daybreak graphic using Artificial Analysis’s published results, checked September 21, 2026. Each pair uses its own test; the two measures should not be combined into a real-world success rate. Data. The Daybreak / Original editorial graphic
Image details

Displayed without generative alteration.

Artificial Analysis’s methodology explains why those distinctions matter. Its task evaluation gives an agent tools and policies in simulated airline, retail and telecom settings, then checks the final database state. Its conversational measure examines behaviors such as pausing, handling interruptions and continuing appropriately when a listener offers a brief acknowledgment. Being pleasant to talk to and leaving the database correct answer different questions. Evaluation methodology[7]

Nor is a benchmark a live call center. The original τ-Voice research uses simulated callers and controlled conditions. Its March preprint found that noise and varied accents made the then-tested voice agents less successful. Those older results do not measure Gemini 3.8, but they explain why judging a voice assistant from a clean demonstration misses part of the problem. τ-Voice research[8]

The new scores make the release worth examining without settling the argument. A substantial gain in a demanding task test is meaningful. It does not tell a customer whether a particular airline has supplied the assistant with working tools, whether a booking service is available or whether the system can recover gracefully when a call drops.

The competition is also learning to listen

The week brought another piece of the voice-AI race. On September 18, SpaceXAI released Grok Voice Transcribe 2.0, a speech-to-text model. Its announcement emphasizes difficult recordings: overlapping voices, phone connections, accents and spoken account details. That is a different product from Gemini’s conversational models; accurate transcription does not itself complete a task. Grok’s release[9]

The distinction matters because an assistant can fail before reasoning even begins. A wrong digit in an address or a missed correction changes the problem it is trying to solve. Conversely, a perfect transcript cannot rescue an assistant that misunderstands the policy or sends the wrong request. Progress has to travel through the whole chain: hearing, interpreting, acting and confirming.

Rows of cubicles in a large call center in Lakeland, Florida.
A call center in Lakeland, Florida, in a 2006 file photograph. This is historical context, not a Google deployment or evidence of jobs replaced by AI. Petiatil / Wikimedia Commons · public domain · Source and rights. Petiatil / Public domain by author
Image details

Displayed without generative alteration.

For the person seeking help, the most valuable improvement might be less repetition. Imagine an assistant that can carry the details of a problem through the entire exchange, accept a correction once and preserve it when the task reaches another service. That is a more useful standard than whether the voice sounds convincingly human.

It also gives companies a better test than call length alone. A short conversation can still leave a customer with the wrong result. The relevant outcome is whether the requested problem was resolved, whether the person understood what changed and whether another call was necessary. Those are proposed measures of usefulness, not outcomes established by this release.

What to watch after the demonstration

Google’s model card is explicit about remaining limitations. Gemini 3.8 Audio can hallucinate and can encounter slow responses or timeouts. It accepts audio, images, video and text, but those inputs do not make every answer a verified observation. The card lists a January 2025 knowledge cutoff, which further distinguishes built-in knowledge from current information retrieved during a session. Model card[10]

For a practical trial, choose something whose outcome you can inspect. Show the assistant a non-sensitive object and ask a question with a checkable answer. In an application that supports actions, inspect the resulting document, appointment or other record. Add a correction mid-conversation and see whether it survives into the final result. That tests the promise of a collaborative interface more directly than an open-ended chat.

The release is interesting because it moves several ingredients closer together: a voice that can stay in the conversation, visual information that reduces the need to explain everything and tools that can change something outside the chat. The independent results suggest there is still a choice between strengths, rather than one universally superior conversational model.

The moment that matters comes at the end. An assistant says the task is done. The record agrees. You did not have to repeat the request, rebuild the context or discover later that a confident sentence was all it had produced. That is the version of voice AI worth wanting—and the standard this new generation now has to meet.

Sources & further reading

Original reporting and research behind this article.

  1. Google’s announcementAccessed 2026-09-21
  2. Search Live helpAccessed 2026-09-21
  3. Watch Google’s demonstration of a conversation continuing during restaurant-booking tasksAccessed 2026-09-21
  4. Live API overviewAccessed 2026-09-21
  5. Tool-use documentationAccessed 2026-09-21
  6. Independent resultsAccessed 2026-09-21
  7. Evaluation methodologyAccessed 2026-09-21
  8. τ-Voice researchAccessed 2026-09-21
  9. Grok’s releaseAccessed 2026-09-21
  10. Model cardAccessed 2026-09-21
Return to the edition