Show HN: Babytalk: Offline speech to text and text to speech on ESP32
github.com/tlack
I have been playing with a lot of $20 devices with no keyboards or screens. I wanted to use speech to interact with them, but without it having to run through a cloud service or use a fixed command vocabulary.
I worked with Claude to implement high performance 4bit int kernels for the ESP32S3 and P4 and quantize some larger STT and TTS models. Surprisingly, it's quite usable on these tiny boards.
Also includes a finetuned model that can better cope with my diesel drones, hacking coughs, and mumbling. Works with Micropython (RAM scarcity issues) and AtomVM (better).
[hidden]
- need a real time original voice to clone voice open source offline library
- need it for recording gaming videos while talking into mic with my voice but output is clone voice
atmanactive[5 comments hidden]
tlack[4 comments hidden]
Many ESP32S3 dev boards have a microphone and speaker header available conveniently on the board, so I have been using that to respond via Babytalk's text to speech capability.
I'll improve the README to make it a little more clear.
atmanactive[3 comments hidden]
tlack[2 comments hidden]
As for what to talk to.. or the ultimate purpose lol.. you could do simple device control scenarios ("lights off"), do hands-free sensor readings ("temperature at 95 degrees"), change wifi settings via voice ("switch access points"), etc.
I use it to control an agent-powered diverse device sensor network in my home and my truck. Much of the capability needs Internet access, but having on-device speech to text and text to speech means I can still do some stuff when I didn't bring the Starlink with me. On-device STT opens up a lot of low bandwidth (lora) opportunities too.
atmanactive[hidden]