Building resilient systems with Sam Newman
Summary
Sam Newman, who coined the term microservices, now calls it an "architecture of last resort"—most outages stem from resource saturation, not architectural choices. His four dimensions of resilience (robustness, rebound, graceful extensibility, sustained adaptability) provide a framework for building systems that survive failure rather than prevent it.
Key Takeaways
- The three rules of distributed systems explain 90% of production failures: information can't travel instantaneously, nodes disappear unpredictably, and resource pools are finite. Most outages trace to resource saturation, not design flaws.
- Production is the source of truth, not specs or code. Specs describe intent but production reveals reality—monitoring and observability should drive architectural decisions, not theoretical models.
- Avoid idempotency keys when possible; use fingerprints instead. Idempotency keys create distributed state management problems that amplify failure modes in microservices architectures.
- The four dimensions of resilience—robustness, rebound, graceful extensibility, and sustained adaptability—replace the binary thinking of "up" vs "down." Build systems that degrade gracefully and adapt continuously rather than attempting perfection.
- Microservices adoption fails when teams ignore organizational readiness. Newman emphasizes microservices as a last resort after simpler architectures prove insufficient—most teams adopt them prematurely for perceived benefits that don't materialize.
Related topics
Transcript Excerpt
I describe microservices as being an architecture of last resort. >> Really? >> Yeah. Yeah. Absolutely. What a lot of people get wrong, I think, is they adopt microservices and they think other magic stuff's going to happen. >> You talk about the three rules of distributed systems. What are these three rules? The first is that you can't send information instantaneously from point A to point B. Rule number two is sometimes the thing you want to talk to isn't there. And the third rule is that resource pools are not infinite. The vast majority of the time when you have a system outage, just down to that, some resource has been totally saturated somewhere. What are the things that we know that we do not know about AI? >> I think the tech world in general is quite naive about AI and fundamental…