Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:
https://www.youtube.com/watch?v=WAeHgE94rVo
No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.
[1]: https://deepmind.google/models/gemma/gemma-4/
[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design
It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion
<grumble grumble people putting in links they expect you to follow to arbitrary goatse youtube videos for all I know>
The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.
Our imaginations and minds continue to rot under the weight of endless and effortless entertainment
Awesome! I had been meaning to build something like this for a while now, but never got around to it.
Is it possible to annotate your text with extra 'stage directions' that influence how the book is read out?