logoalt Hacker News

Inflect-Micro-v2: complete voice in 9.36M parameters

99 pointsby nateb2022today at 12:36 AM7 commentsview on HN

Comments

yjftsjthsd-htoday at 3:21 AM

Couple highlights:

> Complete local text-to-waveform speech synthesis under 10M parameters.

In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

> English only, with one fixed male voice. This is not zero-shot voice cloning.

(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

modinfotoday at 5:48 AM

This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!

here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd

thanks for shearing!

itaketoday at 6:29 AM

Amazing quality for small size, but definitely not that enjoyable to listen to.

IMHO, its at about the same quality level of historic TTS tools.

show 1 reply
tmalytoday at 2:01 AM

This is impressive. I wish there were a voice clone option.

show 1 reply
jsomedontoday at 2:55 AM

amazing quality for such small size!