Beyond the Chat Box: The Coming Wave of Interactive Agents

General

Matt Wyman

,

CEO/Co-Founder

For the last three years, "agent" has mostly meant a text box. You type, it types back, maybe it calls a tool. That model is about to stop being the whole story. Chat will not go away, any more than the command line did when the GUI arrived. The terminal is having a renaissance right now, and some of us never left. But chat is about to become one mode among several, not the interface itself.

We are at the very early stage of a new type of agent. It listens, it speaks, it sees what you see, and most importantly it builds the experience as the conversation unfolds. The screen is no longer a fixed set of pages the user navigates. It is something the agent composes in response to what the user says and does.

This matters to anyone shipping customer-facing software, because the unit of quality is changing. It is no longer "did the agent answer correctly." It is "did the person get where they were trying to go."

The store you can talk to

E-commerce is where this is showing up first, and the shift is bigger than it looks. For twenty years, online retail has been built by marketers. Category pages, filters, recommendation carousels, retargeting. The store pushes products at you and hopes something sticks.

The new pattern inverts that. You describe what you actually need:

  • "A gift for my father-in-law who just retired and is getting into woodworking, under $150."

  • "Something to stop my bathroom mirror fogging up. I rent, so nothing permanent."

  • "Hiking boots for wide feet. My last pair gave me blisters on the outside heel."

The agent asks a clarifying question or two, narrows the field, and then assembles a response: a short set of options, a side-by-side comparison, a fit guide, a bundle. Not a page that existed before the conversation, but a view built for it.

That is a salesperson, not a marketer. The good salesperson listens first, understands the problem behind the request, and is willing to say "that one won't work for you." Retailers have never been able to put that person in front of every digital visitor. Now they can.

Of course anything can be abused. An agent optimized for margin rather than fit is just a more persuasive marketer, and a more persuasive marketer with the customer's trust is worse than a carousel. Which is exactly why the behavior of these agents needs to be measured, not assumed.

Voice moves into the app

Mobile is taking the same technical step toward a very different experience. App teams are adding voice as a first-class input inside the UI, not as a separate assistant bolted on the side.

The reasoning is simple. Mobile lives in the physical world. People use their phones while walking, driving, cooking, holding a child, or standing in a store aisle. In those moments, speaking is the fastest and often the only practical way to tell an app what you want. Tapping through four screens to reschedule a delivery is friction. Saying "move it to Thursday after five" is not.

What makes this interesting is the blend. The user speaks, the app responds by changing what is on screen, the user taps to confirm or corrects by voice, and the agent carries context across all of it. Voice in, visual out, touch to commit. Neither a voice agent nor a chatbot, but a single conversation expressed across modalities.

What actually changes under the hood

The common thread in both cases is that the interface stops being a fixed artifact and becomes an output of the agent. Not of a single model, but of an agent that orchestrates several: speech recognition, a language model, retrieval and ranking, vision, and the tools behind them. That has real engineering consequences.


One conversation, many modalities: the person speaks and taps, the agent orchestrates several models and tools, and composes the view and the spoken reply
  • State lives in three places at once. The conversation history, the UI the user is looking at, and the backend systems the agent calls (catalog, inventory, cart, account) all have to stay consistent. A recommendation the agent spoke about but the screen no longer shows is a bug, even if every individual component behaved correctly.

  • The user's actions are part of the dialogue. A tap, a scroll past three products, a filter toggled, a long pause. These are turns in the conversation, and the agent is expected to interpret them. "3 products shown, user selected none" is as meaningful as anything the user said.

  • Modality errors compound. A speech recognition miss on "wide" versus "white" sends a footwear search down the wrong path, and the visual results then reinforce the error. The failure starts in audio and surfaces on screen.

  • The output space is effectively unbounded. When the agent composes the view, there is no finite set of pages to test. Two users with the same intent will see different screens.

Every one of these is a non-deterministic system interacting with a non-deterministic human. That is the part most teams are not yet ready to handle.

From bleeding edge to early majority in 2027

We have all known multi-modal experiences were coming. Today you can point to a handful of bleeding-edge examples: a few retailers with conversational shopping, a few apps with voice woven into the UI. By 2027 we will be past that edge and into the early majority. The model capabilities are already here. What remains is more than integration. It is conversational design: deciding how the agent listens, what it asks, when it shows instead of tells, and how it recovers when it gets something wrong. That is the work enterprises are funding now. Even the titles are new. Conversational Design and Conversational Engineering are emerging as disciplines of their own, with their own teams and budgets.

That transition is where the business pain shows up. Early adopters accept rough edges because they are experimenting. The early majority does not. They have brand standards, compliance teams, conversion targets, and a board asking why the AI initiative has not moved a metric yet. The questions they ask are practical:

  • Will it recommend something we do not sell, cannot ship, or are not allowed to say?

  • Does it treat the customer with wide feet the same as the one with a gift budget of $40?

  • When the voice layer mishears, does the experience recover or does the customer abandon?

  • Did last week's prompt change quietly make the agent more pushy?

None of these are answered by a demo. They are answered by evidence at scale, before launch and continuously after.

Simulating real people is the path forward

Traditional QA assumes a finite set of paths. Click this, expect that. Scripted test cases and screen-level assertions break the moment the screen is generated rather than designed. Offline evals on single responses help, but they score the answer, not the journey.

The only way to know how an interactive agent behaves is to have realistic people use it, many times, across the situations that matter. At human scale that is a usability study: slow, expensive, and stale by the next release. At machine scale it is simulation.

A useful simulation for this new class of agent has to do a few things well:

  1. Play a person, not a script. A synthetic user with a goal, a constraint, and a temperament. The retiree's son-in-law with a budget, the renter who cannot drill holes, the impatient caller who interrupts.

  2. Use the same modality the customer does. Speak to the agent through audio when the customer will speak. Look at and act on the rendered UI when the customer will tap. Testing through an API alone skips the layer where many failures start.

  3. Judge the outcome, not just the words. Did the person end up with something that fits their need? Did the agent stay within policy? Did what was said match what was shown?

  4. Run continuously. Every prompt, model, or catalog change is a new system. The simulation suite is the regression test.


A synthetic Driver talks, looks and clicks through the Target, and the Run records what was said, shown and done, scored by Checks

This is the direction we are building at Okareo. Our simulations already drive agents over API and over voice with synthetic users, and we are extending the same driver to operate agents through their UI, so a single simulated person can talk, look, and click their way through an experience. The result is a record of what the person said, what the agent did, and what the screen showed at each step, which is what a product team actually needs to judge whether the experience works.

What to do now

If you lead product or engineering for a customer-facing experience, the window to prepare is the next twelve months. A few practical moves:

  • Define success as the customer's outcome. "Found a product that fits" and "completed the change without help" are measurable. "Responded helpfully" is not.

  • Write down your personas and their hard cases now. The wide-footed hiker and the $40 gift buyer are your test suite. So is the customer with a strong accent in a noisy store.

  • Treat every modality as a failure surface. Audio, screen, and tool calls each break differently, and they break each other.

  • Make simulation part of the release gate, not a project you start after the first incident.

The agents coming next will feel less like software and more like the best salesperson or service rep you ever dealt with. The companies that win will be the ones that can prove their agent behaves that way, for every kind of person, before that person ever shows up.

Join the trusted

Future of AI

Get started delivering models your customers can rely on.

Join the trusted

Future of AI

Get started delivering models your customers can rely on.