Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Even more than with tech stacks, deferring on-call process to "Google does it this way so we should do it this way" feels like a terrible idea. There's maybe ten companies in the world that have Google's scale and needs in this regard, and even though it would probably be good for their developers for Facebook to adopt Google's processes based on what I see in this thread, they also probably won't.

The rest of us have to muddle with questions like, how do we do it if we only have 20 people and they're only in two time zones and only half of those really know how to diagnose and recover a corrupted filesystem? A Google-like approach to error-budget-centric risk management just doesn't fit into that world.



If you only have 20 guys and a corrupted filesystem is one of your potential problems, you're doing it wrong. That's why people have switched to cloud services - You pay for the ability to flatten your systems, and if you've not architected with the ability to flatten your systems, you're gonna be SOL in those situations.

Ultimately, you're resilient for what you prepare for. There's a lot of tradeoffs in spend. I get that that's an example and probably not a real pressing concern, but the point is: You shouldn't have everyone trained in everything. You should have escalation paths for everything non-obvious. You should also train your people better.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: