Enterprises that watched an artificial intelligence agent clear internal testing and then fail in production are moving faster than their peers to take humans out of deployment decisions, according to data from VentureBeat Intelligence.
Forty-nine percent of the 108 enterprises surveyed said an AI agent or large language model feature that had passed company testing went on to create a problem visible to customers. That figure was 50% in June. Nearly a quarter of respondents, 24%, said the outcome had occurred more than once. Across 265 responses over the two months, the share reporting at least one test-approved system that disappointed customers stayed within a single percentage point.
Trust in automated evaluation rose over the same period. Thirteen percent of July respondents said they trust automated evaluation, up from 5% in June. The share naming poor alignment between tests and real-world results as their biggest concern fell 10 points, from 29% to 19%.
Confidence tracked closely with experience. Among enterprises reporting a test-approved system that later disappointed a customer, 4% placed complete faith in automated checks. Among those reporting no comparable incident, 24% did, a sixfold difference. In raw counts, that was two of 53 burned enterprises and 10 of 41 unburned ones.
Overall, 67% of respondents said they either already let an agent push code or change a system without a person’s approval in certain low-risk cases, or are modifying their pipelines to support that practice within the coming year, unchanged from June. Thirty-seven percent already permit it in limited cases and 30% are building toward it.
The split by experience was wide. Among enterprises where a test-approved system had disappointed a customer, 85% were pursuing the no-approval model, compared with 61% of the group reporting no such incident. Eleven percent of burned respondents rejected end-to-end deployment automation for the years ahead, versus 24% of unburned respondents.




