Register or log in to access this video

New York • September 8 & 9, 2027
Loved LDX3 New York? Pre-sale tickets for 2027 are now available.
When Fetch aired its Super Bowl commercial, we had one chance to get it right. We ran a live $1.2M sweepstakes during the final two minutes of the game: anyone who opened the app entered automatically, and 120 winners were drawn and announced in-app in real time. Our normal signup load was ~1 RPS. Our target: 150,000. There was no rehearsal, and the sweepstakes carried legal obligations that meant certain operations had to be both fast and provably correct.
This talk focuses on the engineering leadership decisions that made this survivable. I’ll share how we decomposed the signup critical path into three buckets (must be synchronous, can be deferred, can be dropped) and how that framework shifted when a legally binding lottery meant some operations couldn’t be deferred or dropped. I’ll cover the architectural trade-offs: stripping the signup flow with feature flags, moving deferred work off the critical path with Kafka and SQS, and shifting verification to dedicated Redis clusters. Then I’ll explain how we decided which safety checks to temporarily relax, how we bounded that risk, and how we built confidence when our stress environment couldn’t replicate production.
Our stress testing caught some surprises early: Redis became CPU-bound at 150K RPS while DynamoDB held steady, the opposite of what we expected. But game day brought its own. Postgres connection limits quietly became a scaling ceiling, and an SMS rate limit we knew about still cost us valuable minutes of cross-team coordination during the live broadcast. Each surprise required real-time judgment calls, and our ability to respond came down to how we’d structured launch-day operations: defined roles, communication channels, and escalation paths designed with the same redundancy we’d built into the architecture.
If you’re preparing for a traffic spike with an immovable deadline, you’ll leave with a practical framework for decomposing a critical path, launch levers for controlling risk in real time, and hard-earned lessons about the gap between knowing a risk and being ready for it.
Key takeaways:
- How to decompose a critical path into synchronous, deferred, and droppable work using feature flags and queues as scaling levers.
- How stress testing at 150K RPS revealed counterintuitive bottlenecks (Redis, not DynamoDB, became the ceiling) and why preparing to be wrong mattered more than trying to be right.
- How in-memory static assets and gradual rollout patterns can shield backend services when millions of users transition from a live event back into the regular app.
- How to structure launch-day team operations (roles, comms, escalation) with the same redundancy you design into your architecture.