Thank you for the explaination. So it is a library and would require someone to adjust it to their use case. That's what got me confused: wide open use case, but nothing tangible. I mean, if an ESP device is doing STT+TTS, then what am I supposed to be talking to? But if it is an ESP device that can do ANY->TTS and/or STT->ANY with some code customization needed, then I think I get it.
I wrote the README from the perspective of being understandable by the lay person, and I think I left out some vital clues about this being a software building block for making your own ESP32 projects more intelligent. I pushed a README update to make it 0.1% more clear.
As for what to talk to.. or the ultimate purpose lol.. you could do simple device control scenarios ("lights off"), do hands-free sensor readings ("temperature at 95 degrees"), change wifi settings via voice ("switch access points"), etc.
I use it to control an agent-powered diverse device sensor network in my home and my truck. Much of the capability needs Internet access, but having on-device speech to text and text to speech means I can still do some stuff when I didn't bring the Starlink with me. On-device STT opens up a lot of low bandwidth (lora) opportunities too.