Training data for human-like voice

Geko’s mission is to make voice models more human-like.

To do so, we focus on two things:

  • Collecting training data with specialists, and co-authoring the research with them: linguists, phoneticians, voice actors
  • Automating real-time audio RL environment creation, so a model improves on how it sounds in a live conversation

Scraped audio is thin outside a handful of languages and rarely carries what makes speech sound human: timing, breath, emphasis, switching language mid-sentence. Recording it with the people who study voice puts those back in.

An environment that listens in real time lets a model improve on the thing that matters, how it sounds in a conversation, instead of on a transcript after the fact. Both efforts compound: better data raises the ceiling, better environments get a model to it faster.

Email us at amirlan@geko.sh, or book a call