London

June 28–29, 2027

New York

September 15–16, 2026

Berlin

November 9–10, 2026

The data canary: How Netflix validates catalog metadata

This talk tells the story of how we built the Data Canary: an automated system that validates data transformations using real production traffic, detects regressions in 2.5–4 minutes, and blocks bad data from publishing, all within a 10-minute window.

Speakers: Celina Amados

Register or log in to access this video

Create an account to access our free engineering leadership content, free online events and to receive our weekly email newsletter. We will also keep you up to date with LeadDev events.

Register with google

We have linked your account and just need a few more details to complete your registration:

Terms and conditions

 

 

Enter your email address to reset your password.

 

A link has been emailed to you - check your inbox.



Don't have an account? Click here to register
September 29, 2026
NYC 27 Pre sale ticket block image

At Netflix, catalog metadata is mission-critical. It defines what titles exist, whether they can be played, and where they’re available, powering the experience for millions of members. When a production incident revealed that corrupted data could break streaming while every code canary showed green, we recognized a fundamental gap: we had rigorous deployment validation for code, but none for data.

This talk tells the story of how we built the Data Canary: an automated system that validates data transformations using real production traffic, detects regressions in 2.5–4 minutes, and blocks bad data from publishing, all within a 10-minute window.

We’ll cover three core innovations:

  • The Orchestrator Pattern: a dedicated cluster architecture that separates baseline from canary, avoids self-testing, and provides a generic integration point extensible to other data sources at Netflix.
  • Extending Our Chaos Platform: customizing experiment thresholds for tight time constraints, using sticky canaries to prevent cross-contamination, and discovering that behavioral metrics (Starts Per Second) detect catalog corruption far more reliably than latency or error rates.
  • Production-Hardened Engineering: handling in-flight experiments during redeployment, leader election across concurrent orchestrator instances, and running controlled failure injection experiments to validate the validator before going live. The patterns we landed on aren’t Netflix-specific. Any team running high-velocity data pipelines that directly impact customers should be asking: what’s your MTTD for data corruption? Can you validate with production traffic safely? How do you detect emergent issues in transformed data? This talk offers a practical, replicable blueprint for bringing code deployment rigor to data.

Key takeaways:

  • Data deployments carry the same production risk as code deployments and deserve equivalent validation rigor
  • Behavioral metrics like Starts Per Second can be more reliable signals for data corruption than latency or error rates
  • The orchestrator pattern, with permanent baseline and canary clusters, provides a reusable, extensible architecture for data pipeline validation
  • Sticky canaries and real production traffic are essential for detecting emergent failures that shadow traffic cannot surface
  • Chaos experiments can be tuned for tight time windows by prioritizing speed of detection over statistical confidence