Just added this to my benchmark site: https://multilingualsttbench.com/
It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.
I'm confused, doesn't your leaderboard clearly show it is the most accurate model? It's number one in the leaderboard. Am I missing something?
Thanks for this, really helpful.
I would also like to see benchmark for translation. I'm looking for live translated subtitles so my Japanese wife can enjoy any show with out waiting months for official VOD streams to release them.