The full capability set.
Base transcription plus a suite of optional add-ons — diarization, Indic PII, custom vocabulary, and delivery mechanisms. Turn on only what a job needs.
24 Indic languages, one model.
One accuracy-tuned model, purpose-built for Indic languages and Indian call audio. Enterprise-comparable Hindi content capture at a fraction of the cost.
Tested tier: hi, en, ta, te, bn, gu, kn, ml, mr, or, pa, as, ur, ne, mai, bho, sa, sat, mni. Experimental: ks, sd, doi, brx, kok — routed to the same backend, flagged on the pricing page.
Pass language_codes: ["hi", "en"] and the model transcribes code-mixed calls in one pass. No splitting audio, no separate requests.
Boost proper nouns, brand names, and industry jargon so the model catches them on the first pass. Reduces manual correction on your transcripts.
Pick 'fast' for a budget tier (planned) or 'pro' for the accuracy-tuned production model. Same API surface, different price point.
Who spoke when, with confidence.
Speaker diarization runs in parallel with transcription — no extra wall-clock cost. Configure min/max speakers per call.
Identifies distinct speakers across a call and aligns them to speech spans. Runs concurrently with transcription so you don't pay the wait twice.
Set min_speakers and max_speakers per request when you know how many parties are on the call. Improves accuracy on short calls where diarizer inference can drift.
Every word carries a start_ms, end_ms, and confidence score. Build waveform players, aligned subtitles, or confidence-gated review queues without post-processing.
Indic PII redaction, built-in.
Detect and redact India-specific personally identifiable information — the entities your compliance team actually cares about, not just US SSNs.
PAN, Aadhaar, phone numbers (Indian format), UPI IDs, GSTIN. Add to the base entity set (person names, credit card numbers, emails, dates) that most ASR APIs ship.
Redact only the transcript text, or muffle the audio at the timestamp, or scrub all downstream fields. Pick the mode that fits your storage and compliance model.
Replace redacted spans with type-preserving placeholders (e.g. [PAN], [UPI_ID]) so downstream NLP pipelines don't break on empty strings.
Audio never leaves India-region infra once we're on the India-region deploy. Per-request audit log. RBAC-scoped API keys. SOC 2 planned for v2.
Get results in the format you already use.
Output formats and delivery mechanisms that match how transcripts actually get consumed downstream.
Pass output_format in your params — the same audio comes back shaped for a video player, a caption editor, an analytics pipeline, or a human reviewer.
Set a webhook_url and a webhook_auth_header. Job-completion POSTs are HMAC-signed and retry with exponential backoff. Also poll GET /v1/jobs/:id if you prefer pull-based.
Every request accepts an idempotency_key. Retries with the same key return the same job_id instead of creating duplicates. Safe to retry on network hiccups.
Batch SLA you can plan around.
Batch-only tier for MVP. Results within 30 minutes, P50 ~15-20 minutes. Real-time streaming API planned for a later phase.
Results within 30 min of upload, P50 15-20 min. Backed by a queue-drain worker that fires on volume (≥30 min audio queued) or age (oldest wait ≥20 min).
Cap each API key's monthly audio-hours or ₹ spend. Cap the whole organisation on top. Never wake up to a surprise bill because a client library retried in a loop.
Every admin action (key revoke, cap change, refund) is logged. Kill switch flips the batch worker off in seconds if anything goes wrong.
Start transcribing in Hindi + 18 more languages
Batch tier. Launching soon — get early access.