Scaling can take time but they've had at least a year to prepare. It's not like any demand increase they've seen was overnight...
At most normal companies you monitor your systems and address potential bottlenecks before they reach a tipping point. And generally you want enough headroom that a sudden 2-3x increase in demand wouldn't take out the service.
Either Github's technical leadership/talent is completely out of their depth (seems unlikely), or they're not hiring enough engineers to work on these capacity issues.
I got the impression from a previous thread that the continual issues were largely driven by the required move to Azure, and underlying Azure stability
More likely, as seen in a lot of companies, perhaps infra teams got whittled down (or frozen HC, less than BAU, etc) with resources reallocated to AI org units.
> Either Github's technical leadership/talent is completely out of their depth (seems unlikely), or they're not hiring enough engineers to work on these capacity issues.
Regarding their leadership, I'd argue the latter implies the former.
So yeah I'd say they're out of their depth.