AI Agent Testing

Custom evals for every AI Agent

As Principal PM for the Enterprise Adoption team, I shaped a bulk-testing feature that would allow every brand to test their AI Agent against custom evaluation criteria for all of their key use cases before launch — now used for every implementation.

Role
Product Manager
Team
1 Product Designer, 1 Eng. Manager, 8 Engineers, plus another Product Manager to complete delivery during my parental leave
Timeframe
  • Jul—Oct 2025: Vision through Milestone 1 build & leave handoff
  • Nov—Feb 2026: Team finished build on-time without me
Scope
Product strategy, requirements & roadmapping, evaluation framework design, cross-functional facilitation, executive alignment

Summary

AI Managers had no safe way to know what a change would do before it hit real customers — testing was manual, reactive, and lived in staff-only tools. As Principal PM, I defined the strategy, requirements, and vision for a testing platform that lets customers simulate and evaluate AI Agent conversations at scale.

TL;DR:

Key Themes

Defining "correct" for a probabilistic system

Testing generative AI isn't like testing deterministic software — the same input can produce different, equally valid outputs. The heart of this project was designing an evaluation framework where customers define quality on their own terms: natural-language success criteria, judged Pass/Fail by an LLM, held to a measurable bar of agreement with human annotators. Test cases became the goalposts that define what a "good" AI Agent meant.

Scoping ambition into a buildable sequence

Everyone agreed testing mattered; nobody agreed on where to start. Simulation fidelity, version control, drafts, deployments — the problem space sprawled. I ran prework alignment, engineering experiments, and UX working sessions to converge fourteen opinionated people (plus four Director+ stakeholders) on a narrow, extensible foundation: single-turn simulation first, evaluation second, dashboard testing third, with other iterations deliberately deferred.

Infrastructure as product strategy

The cheapest version of this feature would have been a bolted-on testing tool. Instead, I aligned the team on building simulation and evaluation as reusable domains — foundations that other product teams could build on, that could absorb Ada's fragmented internal tooling, and that could one day ship as standalone APIs for evaluating any customer service transcript, human or AI.