Contact center technology vendor UJET has published a seven-part testing framework intended to catch failures in agentic AI systems before they reach live customers, arguing that demos measure capability once while production demands reliability every time.
The report cites two academic benchmarks. Salesforce AI Research’s CRMArena-Pro found leading large language model agents succeeded on about 58% of single-turn service tasks, falling to roughly 35% in multi-turn settings requiring context retention. A separate framework, τ-bench, reported that state-of-the-art agents completed under 50% of retail and airline service tasks on a single attempt, and under 25% when required to succeed on the same task eight consecutive times.
The seven tests are:
- Replaying 20 to 50 real failed conversations drawn from escalation queues;
- Running the same 20 conversations eight times and scoring only cases that succeed every time;
- Multi-turn scenarios with customers who correct themselves;
- Adversarial attempts including prompt injection, with a target attack success rate of zero for unauthorized financial actions, policy invention and personal data disclosure;
- A “stack test” requiring the agent to complete a transaction and produce verifiable writes to every system of record;
- Five escalation scenarios in which the receiving human must inherit full context; and
- A disclosure and audit test confirming the agent identifies itself as AI on the first turn and that any session can be reconstructed within 24 hours.
The article grounds its liability discussion in Moffatt v. Air Canada, a 2024 British Columbia Civil Resolution Tribunal decision ordering the airline to pay $812.02 after its chatbot invented a bereavement fare policy. On disclosure, the piece notes that Article 50 of the EU AI Act took effect Aug. 2, 2026, and points to California’s B.O.T. Act and Utah’s amended AI Policy Act as U.S. requirements.
UJET cites a June 2025 Gartner prediction that more than 40% of agentic AI projects will be canceled by the end of 2027, along with Gartner’s estimate that only about 130 of the thousands of vendors marketing agentic AI offer substantial agentic capability, which Gartner termed “agent washing.”
The article promotes UJET’s own AXO platform and states the framework should be applied to the company’s products as well as competitors’.




