I’ve had a lot of success in the past with fine tuning STT using synthetic data.
I was doing it for Veterinary (ambient recording -> SOAP notes) which has tons of complex domain-specific language AND it is critically important to get right.
“CPR” transcribing as “see pee are” just doesn’t cut it in that industry.
Which open source STT models have you had success with for fine tuning?
Synthetic-data fine-tuning is the other credible answer to domain vocabulary. Curious whether you re-benchmark the fine-tune when new base models ship?