This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!
here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd
thanks for shearing!
Amazing quality for small size, but definitely not that enjoyable to listen to.
IMHO, its at about the same quality level of historic TTS tools.
amazing quality for such small size!
Couple highlights:
> Complete local text-to-waveform speech synthesis under 10M parameters.
In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.
> English only, with one fixed male voice. This is not zero-shot voice cloning.
(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)