Youtube is very bad, but I have honestly not yet seen anyone implement this well on any platform, which indicates to me that it's not a solved problem. Very open to being proven wrong here though - do you have examples of Whisper Large in use in a large consumable sample?
It's not fully solved, there are still errors but its improved a lot the past few years especially compared to Youtube. The best approximation is the Whisper models atm as far as I have seen. There is a write up here on hugging face which talks more about the models and the benchmarks against it.
https://huggingface.co/openai/whisper-large-v3