Artificial intelligence (AI) has entered a new phase in customer experience (CX). It is no longer enough for a virtual agent to route a call or answer a question well. Customers expect it to understand what they are trying to achieve, take the right action and help move the interaction toward a meaningful outcome. 

That was the promise behind Genesys Cloud™ Agentic Virtual Agent: AI that can understand goals, determine next steps, use tools and complete work safely within enterprise guardrails. 

Building on that foundation, the next opportunity is to make agentic experiences easier to design and deploy, more dependable at scale and more natural for customers to engage with. As organizations move from early deployments to broader adoption, those qualities become increasingly important to deliver business returns. 

Genesys is investing across each of those dimensions — making agentic experiences easier to design, test and deploy, strengthening the intelligence and reliability behind them, and making customer interactions feel more natural. Three recent advancements bring that evolution to life: spec-driven development, the APT-2 model upgrade and enhanced voice experiences. Together, they strengthen the full lifecycle of agentic AI, from the first business idea to the live customer conversation. 

Why great ideas get stuck in configuration   

Most organizations don’t lack ideas for improving CX. They know the journeys they want to simplify, the issues they want to resolve and the moments where customers need more help. 

The difficulty comes with turning that intent into a working experience that involves configuring prompts, tools, knowledge, policies, tests, deployment checks and monitoring. 

Each element matters, but the distance between the desired outcome and the final configuration is where good ideas slow down or stall. 

Spec-driven development closes that distance. It gives teams a structured, AI-powered framework to translate that goal, and any related documents and examples, into the platform configuration needed to deliver it. 

This doesn’t require users to become prompt engineers. Behind the scenes, AI uses a structured set of reusable skills, guidance and examples to help create and check the configuration in a consistent, controlled way. 

Different teams can work the way that suits them. Developers may choose AI coding environments and programmatic interfaces, while business and technical users can work inside Genesys Cloud with conversational assistancerecommendations and structured workflows. The interfaces enable collaboration: describe the outcome, provide the right context, let AI help design the path — without losing the testing, governance and oversight enterprise CX demands.

Reasoning customers can trust

Building an agentic experience faster only matters if the experience performs reliably. 

Customers don’t think about the model behind a virtual agent. They judge what happens in the moment. Did it understand the request? Did it remember what they’d already said? Did it take the right action? Did it solve the problem? 

AI services have to earn customers’ trust quickly. One wrong answer can undermine confidence. One false confirmation can create more work. And one inconsistent policy response can result in real costs. 

McKinsey’s 2026 AI Trust Maturity Survey identifies inaccuracy as one of the most frequently cited AI risks. As AI systems become more autonomous, McKinsey warns that organizations must manage not only the risk of AI giving a wrong answer, but also of it taking unintended actions. 

The APT-2 model upgrade behind Agentic Virtual Agent strengthens that foundation. It’s built to improve accuracy and factual grounding, and to reduce the failure modes that erode customer trust. 

This matters because real customer journeys rarely follow a straight path. A billing question can become a payment request. A delivery update can turn into an address change. A support interaction can require troubleshooting, verification and action across several systems. Handling that takes room to adapt: richer instructions, more complete policies, more tools and continuity as the conversation changes direction. 

Updates to the APT-2 model include an expanded context window, giving Agentic Virtual Agent more capacity to support richer agent specifications and longer, more complex workflows. Speed matters for trust too, even when customers can’t pinpoint why. They know when an interaction stalls, when they’re waiting too long, when self-service starts to demand more effort than it’s worth. Evaluations of the updated APT-2 model performance show up to 50% faster time-to-first-response in testing compared to APT-1, so conversations don’t feel like they’re stalling, even under high demand.  

And because enterprises serve customers across countries and languages, consistency has to travel too: the evaluations show multilingual accuracy improving by approximately 20%, along with higher throughput — more requests per second and output tokens per second — so performance holds up at scale, not just in a demo. 

That kind of performance is what turns a promising pilot into a tool that enterprises can depend on at scale.  

Closing the gap between AI and natural conversation

Reliable reasoning and conversational flow is essential, but how an interaction feels matters just as much — and nowhere more than in voice. 

Customers notice when a system responds too slowly, interrupts before they’ve finished speaking or leaves an awkward pause after every sentence. They notice a flat tone that’s disconnected from the conversation. Even when the answer is correct, small moments can make an interaction feel mechanical, impacting engagement. 

Three changes are closing that gap: 

  • Improved end-of-turn detection goes beyond identifying the end of speech — it helps determine when a person has finished their thought, cutting delays without creating awkward interruptions. 
  • AI-native voice engines work closely with the language and action models behind the conversation, producing speech that reflects the intent and structure of the response rather than simply reading generated text aloud. 
  • Contextual tuning uses the conversation itself to improve term handling and make voice more expressive when the moment calls for it, resulting in an interaction that feels more personable.    

Individually, these are technical refinements. Together, they help determine whether a voice interaction feels like talking to a system or being understood – and that distinction often decides whether a customer trusts the rest of the conversation.

The enterprise bar for agentic AI

These three innovations are stronger together than they are individually. 

Spec-driven development changes how fast teams can design and refine  an experience. Stronger reasoning changes how much they can trust it to follow through. And enhanced voice changes how naturally customers experience it in the moment. 

Put together, they answer the three questions every organization scaling agentic AI is asking: 

  • How quickly can we build it?  
  • How confidently can we trust it?  
  • How effectively will customers engage with it? 

That’s what turns agentic AI from an isolated pilot or point-solution into an enterprise capability. It’s a better way to design, deliver and continuously improve customer experiences at scale, from intent to outcome. 

For the technical details behind these enhancements, visit the Developer Center for deep dives on spec-driven development and the APT-2 model upgrade.