Voice AI for the world’sunderserved languages
First-party models, trained on voice we record ourselves. Speech in and out, on one API key
Сәлеметсіз бе! Тапсырыс нөмірі 152, сомасы 5500 теңге. Рахмет!
Four pilots, live in Kazakhstan
Two national operators, a coffee chain and an AI sales agent
Every figure is the partner’s own, and sizes the business the pilot runs inside rather than Geko’s traffic. Kcell JSC, FY2024 results and VEON integrated annual report are published; pleep and Oneshott Coffee reported theirs to us directly.
Making voice in AI sound indistinguishable from human speech
We record the data, train the models, run the inference
Global providers ship one multilingual model and let the long tail take what falls out. We build one model per language instead.
There is no clean corpus to scrape for these languages, so we record our own. That is the part nobody can shortcut.
Built for the callspeople actually make
Three shapes cover almost every request today, and all three are presets in the playground.
Reception and support
Greetings, hold messages, answers read back. Six voices catalogued by role.
Тапсырыс нөмірі 152, сомасы 5500 ₸, жеңілдік 15%.
Order and status calls
Numbers, sums and dates spoken as a person says them. 5500 never comes out digit by digit.
Сәлеметсіз бе, можно узнать тапсырыс нөмірін?
Calls that switch language
Kazakh to Russian mid-sentence, inside a single word. No language parameter.
Two models, one key
Tokay-1.0 speaks, Seta-1.0 listens, on one key and one balance.
Whisper scores 44.95% on the same Kazakh audio
An hour of audio transcribed in about ninety seconds
Three female, three male, catalogued by role
Roughly 30 minutes of speech, or 2.8 hours of transcription
Tokay-1.0
Text in, WAV out. Six voices, and long passages stream sentence by sentence.
- Model
- tokay-kk-v1
- Output
- WAV, 24 kHz, 16-bit PCM mono
- Voices
- Six — three female, three male
- Streaming
- Sentence by sentence
Seta-1.0
Audio in, text out. Kazakh and Russian in one vocabulary, so switching costs nothing.
- Model
- seta-kk-ru-v2
- Accuracy
- 8.71% WER, against Whisper's 44.95%
- Speed
- 40× faster than real time
- Switching
- Mid-sentence, and inside a word
Both ends of the loop
Hear, decide, speak. Geko owns the first and the last — 8.71% WER in, 24 kHz streaming out.
- 1The caller speaksseta-kk-ru-v2
- 2Your logic answersyour code
- 3Tokay says it backtokay-kk-v1
The caller says: Сәлеметсіз бе, можно узнать тапсырыс нөмірін? Your own code looks up order 152 and finds it ready, total 5500. Tokay speaks the reply: Сәлеметсіз бе! Тапсырыс нөмірі 152, сомасы 5500 теңге. Рахмет!
Chirp
In developmentThe same loop, hosted. Not shipped yet — don’t plan around it.
Two lines, not a migration
OpenAI-shaped endpoints, both directions. Change the base URL and the model, keep every call site.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://geko--tokay-serve-web.modal.run/v1",
apiKey: process.env.GEKO_API_KEY,
});
const res = await client.audio.speech.create({
model: "tts-1",
voice: "Aigerim",
input: "Сәлеметсіз бе!",
response_format: "wav",
});| Endpoint | Method | Key |
|---|---|---|
| /v1/tts | POST | Required |
| /v1/tts/stream | POST | Required |
| /v1/transcribe | POST | Required |
| /v1/audio/speechOpenAI-compatible | POST | Required |
| /v1/audio/transcriptionsOpenAI-compatible | POST | Required |
| /v1/models | GET | Open |
| /v1/voices | GET | Open |
| /health | GET | Open |
- Cold starts
- The GPU backends scale to zero. A first call can take tens of seconds.
- Billing
- Per character out, per second in, on successful requests only.
Need a voice that doesn’t exist yet?
Tell us the language. We record the speakers and train the model


