London

June 28–29, 2027

New York

September 15–16, 2026

Berlin

November 9–10, 2026

Evaluating AI developer tools without the drama

This talk explores how to turn a divided technical evaluation into a decision everyone can trust. Through a real-world AI code review tool rollout, you’ll learn a practical framework for setting shared criteria, rebuilding developer confidence, and making technical decisions with genuine stakeholder buy-in.

Speakers: Christina Chan

Register or log in to access this video

Create an account to access our free engineering leadership content, free online events and to receive our weekly email newsletter. We will also keep you up to date with LeadDev events.

Register with google

We have linked your account and just need a few more details to complete your registration:

Terms and conditions

 

 

Enter your email address to reset your password.

 

A link has been emailed to you - check your inbox.



Don't have an account? Click here to register
September 29, 2026
NYC 27 Pre sale ticket block image

Have you ever inherited a technical decision that went sideways? Where teams were running their own experiments, trust was eroded, and every conversation felt like “my way vs. your way”?


Last year, I took over an AI code review tool evaluation that was heading in exactly that direction. An enthusiastic rollout had backfired, developers had lost trust, and teams were championing different tools with no shared criteria for success.


Six weeks later, we had a clear winner that exceeded our targets. More importantly, we had buy-in from developers who’d been skeptical that any AI tool could work, including those whose preferred tools didn’t make the cut.


The evaluation process was featured in The Pragmatic Engineer newsletter, but this talk goes deeper into the framework that made it work. I’ll share the methodology we used to turn a contentious tool evaluation into a process everyone could trust.


You’ll leave with a reusable framework for making technical decisions where stakeholder buy-in matters as much as the metrics.

Key takeaways:

  • How to define success criteria before testing (and why this is the most critical step)
  • What to measure: quantitative metrics (time savings) and qualitative signals (developer satisfaction, signal-to-noise ratio)
  • How to structure fair comparisons when testing multiple tools in production
  • When to drop tools early vs. giving them more time
  • How transparency and maker-owner culture rebuild trust after failed rollouts