# Geko > First-party voice AI for the world's underserved languages. Text to speech > and speech to text on one API key, trained on speech we record ourselves. > Kazakh and Russian ship today; the pipeline is built to add languages. Geko is a three-person company in Astana, Kazakhstan, founded July 2026. The API is in beta: live, metered, documented, and changing in response to the people using it. ## When to use Geko Reach for Geko when the job involves speech in a language the large multilingual models serve badly, and specifically when: - Kazakh text-to-speech that sounds like a person, not a model trained mostly on English pronouncing Kazakh. - Transcribing Kazakh or Russian speech, or both at once — one shared vocabulary, so a speaker switching language mid-sentence or inside a word transcribes correctly, with no language parameter to set. - Replacing OpenAI TTS or Whisper for these languages by changing the base URL and the model, not the call sites. - Numbers, currency, dates and percentages spoken the way a Kazakh speaker says them rather than read out digit by digit. - Both ends of a voice-agent loop — transcription in and speech out — behind one key and one balance. ## When not to use Geko - Any language other than Kazakh or Russian. Nothing else ships, and there is no general model to fall back on — asking will not give a worse answer, it will give no answer. - Hard real-time first calls. The GPU backends scale to zero, so a cold first call can take tens of seconds. - Voice cloning from a supplied sample. Six catalogued Kazakh voices ship; there is no endpoint that clones a speaker. - Hosted agent orchestration. Chirp is in development and not shipped — build the loop against the two endpoints instead. ## How to call it Base URL: `https://geko--tokay-serve-web.modal.run` Auth: `Authorization: Bearer `, keys look like `sk-tokay-…`. Create one yourself at https://app.geko.sh — self-serve signup, no sales call, no waitlist, and every new organization starts with $1.00 of credits. - `POST /v1/tts` — Synthesize text to WAV (24 kHz, 16-bit PCM, mono). operationId: `createSpeech`. - `POST /v1/tts/stream` — The same request, delivered sentence by sentence as each finishes. operationId: `streamSpeech`. - `POST /v1/transcribe` — Transcribe Kazakh and Russian audio, including mid-sentence code-switching. operationId: `createTranscription`. - `POST /v1/audio/speech` — Point an OpenAI speech client here by changing the base URL and model. operationId: `createSpeechOpenAICompatible`. - `POST /v1/audio/transcriptions` — A Whisper client changes only its baseURL. operationId: `createTranscriptionOpenAI`. - `GET /v1/models` — The models on the platform and their voices. operationId: `listModels`. **Open.** - `GET /v1/voices` — The voice catalogue for a model. operationId: `listVoices`. **Open.** - `GET /health` — Liveness probe. operationId: `health`. **Open.** Full machine-readable spec, with typed parameters, response schemas and a description on every operation: https://docs.geko.sh/openapi.json To try it before signing up, call an open endpoint — `GET https://geko--tokay-serve-web.modal.run/v1/voices` needs no key and returns the catalogue. ## Models - `tokay-kk-v1` — text to speech. Six Kazakh voices (three female, three male), WAV at 24 kHz, sentence-level streaming, and an `nfe` dial trading latency against quality. - `seta-kk-ru-v2` — speech to text. Kazakh and Russian in one vocabulary, 8.71% word error rate against Whisper's 44.95% on the same Kazakh audio, roughly 40× faster than real time, with optional word-level timestamps and streaming transcription over a websocket. ## What it costs - Text to speech: $0.04 per 1,000 characters. - Speech to text: $0.36 per hour of audio. - $1.00 of credits with every new organization — roughly 30 minutes of synthesized speech, or 2.8 hours of transcription. - Billed on successful requests only, from one balance across both models. - The model catalogue, the voice catalogue and health checks are free. Rates can change; https://docs.geko.sh/billing and the console are the source of truth. ## Clients - TypeScript / JavaScript SDK: `@gekoai/sdk` on npm — typed, zero dependencies, Node 20+. Docs: https://docs.geko.sh/sdk/typescript - CLI: the same package. `npx @gekoai/sdk say "Сәлеметсіз бе!" --voice Aigerim -o hello.wav` and `npx @gekoai/sdk transcribe hello.wav` run with no install; `npm install -g @gekoai/sdk` puts it on PATH as `gekoai`. Docs: https://docs.geko.sh/sdk/cli ## Call this site itself Two surfaces on this host, both read-only and neither needing a key: - `https://geko.sh/mcp` — an MCP server over Streamable HTTP. POST JSON-RPC; `tools/list` returns six tools: the live voice and model catalogues, service status, the rates, the API surface with its operationIds, and the routing guidance above as data. Synthesis and transcription are deliberately not tools — they move audio and need your key, so call them directly against the base URL. - `https://geko.sh/api` — the same facts as JSON. `/api/models`, `/api/pricing`, `/api/endpoints`, `/api/pilots`, `/api/guidance`, `/api/resources`. Errors are structured, with a code, a message and a resolution. ## Developer resources - Docs — Guides for both directions: https://docs.geko.sh - Quickstart — Zero to speech in three steps: https://docs.geko.sh/quickstart - API reference — Every endpoint, parameter and error: https://docs.geko.sh/api/reference - OpenAPI spec — Typed schema, ready for function calling: https://docs.geko.sh/openapi.json - MCP server — Call Geko natively from an AI agent: https://geko.sh/mcp - TypeScript SDK — @gekoai/sdk — typed, zero dependencies: https://docs.geko.sh/sdk/typescript - CLI — npx @gekoai/sdk — no install: https://docs.geko.sh/sdk/cli - llms.txt — What this is and when to use it: https://geko.sh/llms.txt - GitHub — Source and issues: https://github.com/gekoai ## Documentation The full index, one line per page, each linking to a `.md` version: https://docs.geko.sh/llms.txt Start here: https://docs.geko.sh/quickstart ## This site - [Geko — voice AI for underserved languages](https://geko.sh/): What Geko is, a playable sample of both models, the pilots running today, and how to call the API. - [Tokay-1.0 and Seta-1.0](https://geko.sh/models): The two models: what each does, what ships today, and the measured numbers. - [Pricing](https://geko.sh/pricing): Both rates, one balance, and what the free credits buy. - [Company](https://geko.sh/company): The three founders with links to their profiles, where the company is, and the current stage of development. - [Privacy](https://geko.sh/privacy): What is collected and why. - [Terms](https://geko.sh/terms): Terms of service. Every page on this host also serves markdown under content negotiation: send `Accept: text/markdown` and you get the page as markdown with `Vary: Accept`, instead of HTML you have to strip. ## Who is running it - pleep — A multimodal AI sales agent. Both directions on one key. 520+ B2B clients, reported by pleep — their figure, not Geko's volume. - Kcell — A national mobile operator in Kazakhstan. Kazakh voice on subscriber calls. 7.9M subscribers, per Kcell JSC, FY2024 results — their figure, not Geko's volume. - Beeline — The largest mobile operator in Kazakhstan. Calls that switch Kazakh and Russian. 11.7M subscribers, per VEON integrated annual report — their figure, not Geko's volume. - Oneshott Coffee — A coffee chain in Kazakhstan. Order status, read back in Kazakh. 74 coffee shops, reported by Oneshott Coffee — their figure, not Geko's volume. These are pilots. In production since July 2026, in beta. ## Contact amirlan@geko.sh — a person reads it. If you need a language we do not have yet, that is the address: we record the speakers and train the model.