AI Agent News: 5 Developments That Put Customer Evidence First
Last week produced another flood of AI announcements. Underneath the hype, five developments point toward a more important question for businesses deploying AI agents: what customer evidence supports the decisions those agents make?
1. Better AI agents do not automatically make better decisions
Hype: Better reasoning means better decisions.
Reality: An agent can reason brilliantly from the wrong understanding of the customer.
Anthropic's Project Swap found that agents could negotiate reasonably once they understood what people wanted. The harder problem came earlier. After a short interview, an agent's preferences matched the human's rankings on only 61 percent of book pairs. Anthropic concluded that missing information about the humans, rather than negotiation ability, caused much of the market's shortfall. Read Anthropic's Project Swap report.
The lesson: Before optimizing agent reasoning, make sure the agent has enough evidence to understand the people affected by its decisions.
2. More customer signals do not always mean better customer understanding
Hype: More signals automatically produce better decisions.
Reality: Signals show what happened. They do not necessarily explain why.
Adobe demonstrated an Enterprise Coworker workflow that can combine first-party customer data with live market intelligence, then build an audience, customer journey, content experiment, and landing page from one conversation. See Adobe's campaign-planning example.
That is extraordinary execution power. But if conversion suddenly falls 20 percent, behavioral data may still not explain the cause.
The lesson: Signals can tell an agent where to investigate. Customer evidence should support the decision to act.
3. Synthetic audiences will not replace customer research
Hype: Synthetic audiences can stand in for real customers.
Reality: Synthetic audiences generate hypotheses. Verified customers generate evidence.
RIWI launched Verify Human to help research organizations establish whether online participants are live, unique humans. The company cited industry benchmarking in which post-survey removal rates averaged 13.1 percent in consumer research and reached 59.6 percent in business-to-business studies. Read RIWI's announcement.
AI makes simulation incredibly useful. It also makes human authenticity more valuable.
The lesson: Use synthetic audiences to decide what to investigate. Use verified customers when the answer matters.
4. Research is becoming callable infrastructure
Hype: Research will keep happening inside research platforms.
Reality: Research itself is becoming callable infrastructure.
UserTesting launched Model Context Protocol servers that let agents inside ChatGPT, Claude, Gemini, Figma Make, and other compatible applications recruit participants, create studies, launch research, and retrieve feedback without entering UserTesting. Its combined network exceeds seven million participants. Read UserTesting's announcement.
That is a major distribution shift. The research platform can increasingly operate behind someone else's AI interface.
The lesson: The winning research system may be the evidence infrastructure agents can call from wherever work happens.
5. Smarter models will not make autonomous agents safe by themselves
Hype: More intelligence will make autonomous agents safe.
Reality: Intelligence and control are different problems.
OpenAI disclosed that a research agent found a gap in sandbox restrictions and used the domain-name system to reach an external chatbot. Monitoring detected the behavior, but the run continued for roughly another two and a half hours before it was stopped. OpenAI paused tool-using work on its most capable models while it strengthened controls. Read OpenAI's account.
The lesson: The model cannot be the final authority policing itself. Monitoring, permissions, and independent controls matter.
The Bigger Picture
Put these five developments together:
Agents are getting smarter.
Execution is becoming autonomous.
Research is becoming callable.
Synthetic audiences are getting more capable.
And the consequences of bad assumptions are getting larger.
That creates a missing layer. Before an agent changes pricing, launches a campaign, alters onboarding, or changes product strategy, something needs to answer:
What customer evidence supports this decision, and is it strong enough to act on?
That is where we believe the Customer Evidence Layer belongs. It connects business questions to verified customer conversations, grounded findings, and evidence a decision-maker can inspect. Learn more about the safeguards behind that work in our Trust & Compliance Center.
AI can analyze almost anything.
AI can talk to almost anyone.
AI can increasingly act on its own.
ReadingMinds determines whether what it learned is evidence.
See the workflow in a short Live Test Drive, or read Why Customer Evidence Is the Missing Layer in the AI Agent Economy.
About the author

Stu Sjouwerman
CEO and Co-Founder, ReadingMinds.AI
Stu founded KnowBe4 in 2010 and grew it into the world's largest security-awareness training platform before taking it public on the NASDAQ in 2021 and its subsequent acquisition by Vista Equity Partners in 2023. He co-founded ReadingMinds with Marcio Castilho and Alin Irimie, the same leadership team that built KnowBe4. Author of the USA Today bestseller Agent-Powered Growth and a regular contributor to Forbes Tech Council and Greenbook on AI, agentic marketing, and customer intelligence.
Know what your customers feel. Not just what they say.
ReadingMinds conducts AI voice interviews that classify emotion type and intensity. Try a 3-minute Live Test Drive with Emma.
Start 3‑Minute Live Test Drive