Register or log in to access this video

New York • September 8 & 9, 2027
Loved LDX3 New York? Pre-sale tickets for 2027 are now available.
At Netflix, catalog metadata is mission-critical. It defines what titles exist, whether they can be played, and where they’re available, powering the experience for millions of members. When a production incident revealed that corrupted data could break streaming while every code canary showed green, we recognized a fundamental gap: we had rigorous deployment validation for code, but none for data.
This talk tells the story of how we built the Data Canary: an automated system that validates data transformations using real production traffic, detects regressions in 2.5–4 minutes, and blocks bad data from publishing, all within a 10-minute window.
We’ll cover three core innovations:
- The Orchestrator Pattern: a dedicated cluster architecture that separates baseline from canary, avoids self-testing, and provides a generic integration point extensible to other data sources at Netflix.
- Extending Our Chaos Platform: customizing experiment thresholds for tight time constraints, using sticky canaries to prevent cross-contamination, and discovering that behavioral metrics (Starts Per Second) detect catalog corruption far more reliably than latency or error rates.
- Production-Hardened Engineering: handling in-flight experiments during redeployment, leader election across concurrent orchestrator instances, and running controlled failure injection experiments to validate the validator before going live. The patterns we landed on aren’t Netflix-specific. Any team running high-velocity data pipelines that directly impact customers should be asking: what’s your MTTD for data corruption? Can you validate with production traffic safely? How do you detect emergent issues in transformed data? This talk offers a practical, replicable blueprint for bringing code deployment rigor to data.
Key takeaways:
- Data deployments carry the same production risk as code deployments and deserve equivalent validation rigor
- Behavioral metrics like Starts Per Second can be more reliable signals for data corruption than latency or error rates
- The orchestrator pattern, with permanent baseline and canary clusters, provides a reusable, extensible architecture for data pipeline validation
- Sticky canaries and real production traffic are essential for detecting emergent failures that shadow traffic cannot surface
- Chaos experiments can be tuned for tight time windows by prioritizing speed of detection over statistical confidence