Training data for human-like voice
Geko’s mission is to make voice models more human-like.
To do so, we focus on two things:
- Collecting training data with specialists, and co-authoring the research with them: linguists, phoneticians, voice actors
- Automating real-time audio RL environment creation, so a model improves on how it sounds in a live conversation
Scraped audio is thin outside a handful of languages and rarely carries what makes speech sound human: timing, breath, emphasis, switching language mid-sentence. Recording it with the people who study voice puts those back in.
An environment that listens in real time lets a model improve on the thing that matters, how it sounds in a conversation, instead of on a transcript after the fact. Both efforts compound: better data raises the ceiling, better environments get a model to it faster.
Email us at amirlan@geko.sh, or book a call