Voice AI for the world’sunderserved languages

First-party models, trained on voice we record ourselves. Speech in and out, on one API key

tokay-kk-v1

Сәлеметсіз бе! Тапсырыс нөмірі 152, сомасы 5500 теңге. Рахмет!

Four pilots, live in Kazakhstan

Two national operators, a coffee chain and an AI sales agent

Every figure is the partner’s own, and sizes the business the pilot runs inside rather than Geko’s traffic. Kcell JSC, FY2024 results and VEON integrated annual report are published; pleep and Oneshott Coffee reported theirs to us directly.

Making voice in AI sound indistinguishable from human speech

We record the data, train the models, run the inference

Global providers ship one multilingual model and let the long tail take what falls out. We build one model per language instead.

There is no clean corpus to scrape for these languages, so we record our own. That is the part nobody can shortcut.

Built for the callspeople actually make

Three shapes cover almost every request today, and all three are presets in the playground.

Aigerimreception
Yerlansupport
Ainurpromos
+ three more

Reception and support

Greetings, hold messages, answers read back. Six voices catalogued by role.

Тапсырыс нөмірі 152, сомасы 5500 ₸, жеңілдік 15%.

→spoken as words, never digit by digit

Order and status calls

Numbers, sums and dates spoken as a person says them. 5500 never comes out digit by digit.

Сәлеметсіз бе, можно узнать тапсырыс нөмірін?

kkrukkone request

Calls that switch language

Kazakh to Russian mid-sentence, inside a single word. No language parameter.

Two models, one key

Tokay-1.0 speaks, Seta-1.0 listens, on one key and one balance.

8.71%
Word error rate

Whisper scores 44.95% on the same Kazakh audio

40×
Faster than real time

An hour of audio transcribed in about ninety seconds

6
Kazakh voices

Three female, three male, catalogued by role

$1.00
Free to start

Roughly 30 minutes of speech, or 2.8 hours of transcription

Text to speech

Tokay-1.0

Text in, WAV out. Six voices, and long passages stream sentence by sentence.

Model
tokay-kk-v1
Output
WAV, 24 kHz, 16-bit PCM mono
Voices
Six — three female, three male
Streaming
Sentence by sentence
The datasheet →
Speech to text

Seta-1.0

Audio in, text out. Kazakh and Russian in one vocabulary, so switching costs nothing.

Model
seta-kk-ru-v2
Accuracy
8.71% WER, against Whisper's 44.95%
Speed
40× faster than real time
Switching
Mid-sentence, and inside a word
The datasheet →

Both ends of the loop

Hear, decide, speak. Geko owns the first and the last — 8.71% WER in, 24 kHz streaming out.

One turn of a call
  1. 1The caller speaks
    seta-kk-ru-v2
  2. 2Your logic answers
    your code
  3. 3Tokay says it back
    tokay-kk-v1

The caller says: Сәлеметсіз бе, можно узнать тапсырыс нөмірін? Your own code looks up order 152 and finds it ready, total 5500. Tokay speaks the reply: Сәлеметсіз бе! Тапсырыс нөмірі 152, сомасы 5500 теңге. Рахмет!

Chirp

In development

The same loop, hosted. Not shipped yet — don’t plan around it.

Build it yourself today

Two lines, not a migration

OpenAI-shaped endpoints, both directions. Change the base URL and the model, keep every call site.

speech.ts
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://geko--tokay-serve-web.modal.run/v1",
  apiKey: process.env.GEKO_API_KEY,
});

const res = await client.audio.speech.create({
  model: "tts-1",
  voice: "Aigerim",
  input: "Сәлеметсіз бе!",
  response_format: "wav",
});
The whole APIThe reference →
EndpointMethodKey
/v1/ttsPOSTRequired
/v1/tts/streamPOSTRequired
/v1/transcribePOSTRequired
/v1/audio/speechOpenAI-compatiblePOSTRequired
/v1/audio/transcriptionsOpenAI-compatiblePOSTRequired
/v1/modelsGETOpen
/v1/voicesGETOpen
/healthGETOpen
Cold starts
The GPU backends scale to zero. A first call can take tens of seconds.
Billing
Per character out, per second in, on successful requests only.

Need a voice that doesn’t exist yet?

Tell us the language. We record the speakers and train the model