Feature set

The full capability set.

Base transcription plus a suite of optional add-ons — diarization, Indic PII, custom vocabulary, and delivery mechanisms. Turn on only what a job needs.

Transcription

24 Indic languages, one model.

One accuracy-tuned model, purpose-built for Indic languages and Indian call audio. Enterprise-comparable Hindi content capture at a fraction of the cost.

19 tested + 5 experimental languages

Tested tier: hi, en, ta, te, bn, gu, kn, ml, mr, or, pa, as, ur, ne, mai, bho, sa, sat, mni. Experimental: ks, sd, doi, brx, kok — routed to the same backend, flagged on the pricing page.

Hinglish and multi-language input

Pass language_codes: ["hi", "en"] and the model transcribes code-mixed calls in one pass. No splitting audio, no separate requests.

Custom vocabulary bias

Boost proper nouns, brand names, and industry jargon so the model catches them on the first pass. Reduces manual correction on your transcripts.

Model tier selection

Pick 'fast' for a budget tier (planned) or 'pro' for the accuracy-tuned production model. Same API surface, different price point.

Speaker intelligence

Who spoke when, with confidence.

Speaker diarization runs in parallel with transcription — no extra wall-clock cost. Configure min/max speakers per call.

Speaker diarization

Identifies distinct speakers across a call and aligns them to speech spans. Runs concurrently with transcription so you don't pay the wait twice.

Min/max speaker constraints

Set min_speakers and max_speakers per request when you know how many parties are on the call. Improves accuracy on short calls where diarizer inference can drift.

Word-level timestamps and confidence

Every word carries a start_ms, end_ms, and confidence score. Build waveform players, aligned subtitles, or confidence-gated review queues without post-processing.

Compliance

Indic PII redaction, built-in.

Detect and redact India-specific personally identifiable information — the entities your compliance team actually cares about, not just US SSNs.

Detect Indic identifiers

PAN, Aadhaar, phone numbers (Indian format), UPI IDs, GSTIN. Add to the base entity set (person names, credit card numbers, emails, dates) that most ASR APIs ship.

Text, audio, or all-fields redaction

Redact only the transcript text, or muffle the audio at the timestamp, or scrub all downstream fields. Pick the mode that fits your storage and compliance model.

Substitution mode

Replace redacted spans with type-preserving placeholders (e.g. [PAN], [UPI_ID]) so downstream NLP pipelines don't break on empty strings.

DPDP-aligned processing

Audio never leaves India-region infra once we're on the India-region deploy. Per-request audit log. RBAC-scoped API keys. SOC 2 planned for v2.

Delivery

Get results in the format you already use.

Output formats and delivery mechanisms that match how transcripts actually get consumed downstream.

SRT, VTT, JSON, or plain text

Pass output_format in your params — the same audio comes back shaped for a video player, a caption editor, an analytics pipeline, or a human reviewer.

Signed webhooks with custom auth

Set a webhook_url and a webhook_auth_header. Job-completion POSTs are HMAC-signed and retry with exponential backoff. Also poll GET /v1/jobs/:id if you prefer pull-based.

Idempotency keys

Every request accepts an idempotency_key. Retries with the same key return the same job_id instead of creating duplicates. Safe to retry on network hiccups.

Reliability

Batch SLA you can plan around.

Batch-only tier for MVP. Results within 30 minutes, P50 ~15-20 minutes. Real-time streaming API planned for a later phase.

Batch SLA: 30 minutes

Results within 30 min of upload, P50 15-20 min. Backed by a queue-drain worker that fires on volume (≥30 min audio queued) or age (oldest wait ≥20 min).

Per-key + org-wide usage caps

Cap each API key's monthly audio-hours or ₹ spend. Cap the whole organisation on top. Never wake up to a surprise bill because a client library retried in a loop.

Kill switches, audit log, RBAC

Every admin action (key revoke, cap change, refund) is logged. Kill switch flips the batch worker off in seconds if anything goes wrong.

Start transcribing in Hindi + 18 more languages

Batch tier. Launching soon — get early access.