As an elixir dev fly was a company I very much wanted to succeed but I twice had to leave them because they just never found the right balance between cool engineering and ops.
For a long time they would have global outages where the status page would list everything as green and the only reason you'd know something was wrong was the forum threads. Their response was that they were too busy fixing the issues to update the status board. Right.
Then once they finally started updating the board they would go several hours between updates with no indication of what was happing "Region x is down" then nothing, again "too busy fixing the issue".
When they came out with paid support I signed up immediately only to find that it gave me an email address to contact support with fast responses promised but most of the time no one was monitoring it. I'd report large outages and not get back responses until the next day and on a few occasions I'd get back a response several days later being like "Can you explain the problem you're seeing?" No mate it was an outage several days ago.
This would all just be annoying customer service if it was a rare occurrence but the reality of using them was that I was dealing with major outages almost monthly at points. If you scroll back through my post history here there were brief periods where stability would improve and I'd get optimistic about them only for it to crash back down to regular outages again days after I posted.
I eventually had to move everything off them to self hosted again which is a pain but far less of a pain since my uptime is dramatically better(!) and when something breaks I know what it is.
I still want fly to succeed but if you're going to host peoples stuff you need to accept responsibility and move budget to ops. Otherwise just kill the hosting entirely and try to become another HashiCorp.
> For a long time they would have global outages where the status page would list everything as green and the only reason you'd know something was wrong was the forum threads. Their response was that they were too busy fixing the issues to update the status board. Right.
Are you talking about AWS or Fly here? I seriously cannot tell.