logoalt Hacker News

riknos314today at 3:36 PM2 repliesview on HN

> Whenever corruption occurred, we had to stop the control plane process on the shard while we repaired or restored the database. This was painful for tailnets on that shard, because their entire control plane disappeared during that recovery window.

Gotta love single points of failure...


Replies

kccqzytoday at 3:44 PM

What are some solutions to avoid database corruption being single points of failure? I can’t think of any off the top of my head. I don’t think people typically consider database corruption to be a kind of failure common enough to design for, unless you have unusual requirements.

show 1 reply
spockztoday at 4:24 PM

The shard was already a way to make it not a single point of failure.