London

June 28–29, 2027

New York

September 15–16, 2026

Berlin

November 9–10, 2026

Building trustworthy agents at scale

This talk explores that journey: how we built testing pipelines for non-deterministic systems, defined “ethical success criteria,” and aligned engineers, product managers, and legal teams around shared principles for responsible AI.

Register or log in to access this video

Create an account to access our free engineering leadership content, free online events and to receive our weekly email newsletter. We will also keep you up to date with LeadDev events.

Register with google

We have linked your account and just need a few more details to complete your registration:

Terms and conditions

 

 

Enter your email address to reset your password.

 

A link has been emailed to you - check your inbox.



Don't have an account? Click here to register
September 29, 2026
NYC 27 Pre sale ticket block image

As large language models move from research to production, engineering leaders are facing a new kind of challenge: how do you test, deploy, and govern systems that can behave unpredictably? At GitHub, we found that reliability wasn’t enough — trust became the core metric. Building our internal coding agents required rethinking how we validated AI output, handled uncertainty, and collaborated across disciplines.

This talk explores that journey: how we built testing pipelines for non-deterministic systems, defined “ethical success criteria,” and aligned engineers, product managers, and legal teams around shared principles for responsible AI. I’ll share the technical patterns that worked — like controlled prompt experiments, sandboxed agent evaluation, and feedback loops — as well as the cultural lessons from introducing ethical review processes into fast-moving engineering work. Attendees will leave with concrete tools for building trustworthy AI systems, a playbook for leading cross-functional conversations about safety and ethics, and practical insights for balancing innovation with accountability.

In a world where AI capabilities evolve faster than our processes, this story is a reminder that trust is built, not assumed — and that engineering leaders have a critical role in making it real.

Key takeaways:

  • Learn how to test and evaluate non-deterministic LLM behavior in production systems.
  • See frameworks for defining and measuring “ethical success” beyond accuracy.
  • Understand how to align engineers, PMs, and legal on responsible AI principles.
  • Apply lessons for balancing speed, experimentation, and accountability in AI projects.