Register or log in to access this video

New York • September 8 & 9, 2027
Loved LDX3 New York? Pre-sale tickets for 2027 are now available.
Have you ever inherited a technical decision that went sideways? Where teams were running their own experiments, trust was eroded, and every conversation felt like “my way vs. your way”?
Last year, I took over an AI code review tool evaluation that was heading in exactly that direction. An enthusiastic rollout had backfired, developers had lost trust, and teams were championing different tools with no shared criteria for success.
Six weeks later, we had a clear winner that exceeded our targets. More importantly, we had buy-in from developers who’d been skeptical that any AI tool could work, including those whose preferred tools didn’t make the cut.
The evaluation process was featured in The Pragmatic Engineer newsletter, but this talk goes deeper into the framework that made it work. I’ll share the methodology we used to turn a contentious tool evaluation into a process everyone could trust.
You’ll leave with a reusable framework for making technical decisions where stakeholder buy-in matters as much as the metrics.
Key takeaways:
- How to define success criteria before testing (and why this is the most critical step)
- What to measure: quantitative metrics (time savings) and qualitative signals (developer satisfaction, signal-to-noise ratio)
- How to structure fair comparisons when testing multiple tools in production
- When to drop tools early vs. giving them more time
- How transparency and maker-owner culture rebuild trust after failed rollouts