logoalt Hacker News

sipjcalast Sunday at 3:46 AM1 replyview on HN

Yep, but I am in the process of also porting NVIDIAs Sortformer for multi speaker diarization as well :)

I’m not sure how many specific models will be supported as the library is more focused on transcription specifically. But the models which support diarization natively must be supported I think. And parakeet multitalker was the primary driving force for this change


Replies

oezilast Sunday at 4:40 AM

How close do you aim for when it comes to drop-in vs whisper.cpp? Are timestamps per word and character something aimed for? How about multi-lingual transcription or hallucination suppression?

The github page doesn't seem to go into depth on these orthogonal topics. May have missed it.

show 1 reply