How does speech to text work? Whisper, real-time transcription architecture, Indian languages, and building a live transcription feature

How automatic speech recognition turns audio into text, what changed with Whisper, how to build real-time transcription with streaming and voice activity detection, how to handle Hindi and mixed-language speech, the accuracy metric that matters, and how I built a live speech-to-text feature into a document workspace.

Taking questions during a session
Taking questions during a session

Key takeaways

  • Speech to text converts audio to text in stages: capture, voice activity detection, chunking, acoustic and language modelling, and post-processing. Modern models fuse the middle stages into one transformer.
  • Whisper made high-quality, multilingual, open-weight transcription free to run. It is the default starting point for most projects.
  • Real-time transcription is an engineering problem: stream short chunks, detect speech boundaries, and reconcile partial results. The model is the easy part.
  • Word error rate on your own audio is the only accuracy number that matters. Benchmarks are recorded in quiet rooms; your users are not.
  • Indian languages and code-mixing need multilingual models and testing on real speakers; do not trust English benchmarks.

What is speech to text and how does it work?

Speech to text, also called automatic speech recognition or ASR, converts spoken audio into written text. Audio is captured as a waveform, converted into a spectrogram of frequencies over time, split into short segments at speech boundaries, and fed to a model that maps acoustic patterns to words, using a language model to pick likely sequences. Post-processing adds punctuation, casing, numbers and speaker labels. Modern systems such as Whisper do the acoustic and language modelling in one transformer trained end to end.

Also asked as: what is speech to text · how does speech to text work · how speech recognition works · what is asr · automatic speech recognition explained · speech to text meaning · asr vs speech to text · how does voice to text work · speech recognition in ai

Two ideas made end-to-end recognition practical: connectionist temporal classification, which lets a network output text without aligning every frame to a character [6], and self-supervised pre-training on unlabelled audio, which wav2vec 2.0 showed could cut labelled data needs dramatically [5]. Whisper then trained a transformer on 680,000 hours of weakly labelled multilingual audio and made robust, open transcription available to anyone with a GPU or a patient CPU [1][2].

I built live speech-to-text into the OCR & Speech Workspace in 2026, alongside page-batched OCR and document chat, so that a spoken question could be transcribed, matched against a document and answered with page citations [13]. This page is the pipeline and the decisions underneath that feature.

Transcription accuracy on a benchmark is a fact about the benchmark. The only number I ship on is word error rate measured on my users' audio, on their phones, in their rooms. Pranjul Rathour

What is Whisper, and why did it matter?

Whisper is OpenAI's open-weight speech recognition model family, released in 2022, trained on 680,000 hours of multilingual and multitask audio [1]. It transcribes in dozens of languages, translates to English, detects language, and is robust to accents, background noise and technical vocabulary in ways earlier open models were not. It comes in sizes from tiny to large, so it runs on a laptop CPU at the small end and on a modest GPU at the large end. It made high-quality transcription free to self-host, which changed what students and small teams can build.

Also asked as: what is whisper ai · how to use whisper ai · whisper ai tutorial · whisper speech to text · openai whisper explained · how to install whisper ai on windows · how to install whisper ai on mac · whisper model sizes · is whisper free · whisper vs google speech to text

For most projects, faster-whisper on a server or whisper.cpp on a device is the practical choice; the official implementation is the reference [2][3][4].

How do I build real-time speech to text?

Capture audio in small frames, run voice activity detection to find where speech starts and stops, buffer frames into chunks of a few seconds, transcribe each chunk as it closes, and stream partial text to the user while a final pass reconciles boundaries and punctuation. Latency is decided by chunk length and model size, not by the network. Over WebSockets or WebRTC from the browser, with the model on a server or on the device, this is a solved architecture; the work is in the boundaries.

Also asked as: real time speech to text · live transcription · how to build real time transcription · streaming speech to text · speech to text websocket · low latency speech recognition · real time speech to text python · browser speech to text · live captions architecture

Silero VAD is the small, fast detector most pipelines use to find speech boundaries [7]. Speaker diarisation, who spoke when, is a separate model; pyannote.audio is the common open choice [8]. My architecture write-up covers the WebSocket flow and the trade-off between chunk length and latency in detail [14].

How accurate is speech to text, and how do I measure it?

Accuracy is measured as word error rate, WER: substitutions plus deletions plus insertions divided by the words in the reference transcript, so lower is better. Modern models reach single-digit WER on clean English benchmarks and much worse on noisy rooms, accents, domain vocabulary and phone audio. Measure on fifty clips of your own users' audio with hand-corrected transcripts; jiwer computes WER in a few lines [11]. Then fix the biggest error class, which is usually names and numbers.

Also asked as: speech to text accuracy · word error rate · how to measure asr accuracy · wer speech recognition · how accurate is whisper · speech to text accuracy comparison · why speech to text is inaccurate · improve speech recognition accuracy

How do I handle Hindi, Indian languages and code-mixing?

Use a multilingual model and test it on real speakers, because English benchmark numbers say nothing about Hindi, Tamil or Hinglish. Whisper supports many Indian languages with widely varying quality across them [1]. AI4Bharat publishes models and datasets tuned for Indic languages that often outperform general models on Indian speech [9], and Common Voice has crowd-sourced Indian-language audio you can evaluate on [10]. Code-mixed speech, Hindi and English in one sentence, remains the hardest case; measure it separately.

Also asked as: speech to text hindi · hindi speech recognition · speech to text indian languages · hinglish speech to text · whisper hindi accuracy · speech to text tamil · multilingual speech recognition · indic asr · speech to text for indian accent

KrishGyan, the assistant that won Changethon at IIT Roorkee, took voice notes from farmers in their own language and answered by voice; the recognition step was where the demo lived or died, and the fix was always testing on real recordings from real phones rather than trusting a model card.

Taking questions during a session
Taking questions during a session

Should I use an API or self-host a model?

Self-host with faster-whisper or whisper.cpp when audio is sensitive, volume is high, or you need to run offline or on device. Use a transcription API when you want the best quality on hard audio without owning GPUs, when volume is low, or when you need features such as diarisation and word timestamps out of the box. Many products do both: on-device for live captions, an API for the final high-accuracy transcript. Price per minute of audio and privacy terms decide it.

Also asked as: speech to text api · best speech to text api · whisper api vs self hosted · speech to text pricing · free speech to text api · open source speech to text · google speech to text vs whisper · azure speech vs whisper · deepgram vs whisper · assemblyai vs whisper

The OCR & Speech Workspace uses a provider's audio capability for transcription alongside its document models, because the product's users upload documents to the same provider anyway and the privacy boundary was already drawn there [12][13].

What is speaker diarisation, and do I need it?

Diarisation labels which speaker said each segment, turning a transcript into a dialogue. It is a separate model that clusters voice embeddings over time, and it adds latency and errors of its own, especially when speakers overlap or the count is unknown. You need it for meeting minutes, interviews and call analysis. You do not need it for dictation, voice commands, captions of a single speaker, or search over a recording. Add it only when the product shows speakers.

Also asked as: speaker diarization · who spoke when · speaker identification speech to text · diarization whisper · pyannote · meeting transcription speakers · speech to text with speaker labels

What are the common uses of speech to text?

Voice assistants and voice-first apps for users who do not type; meeting and lecture transcription with summaries; live captions for accessibility; call centre analytics; medical and legal dictation; searchable archives of recordings; voice input for forms in the field. In India, voice-first interfaces in regional languages are the highest-leverage use, because they reach the people every text interface excludes.

Also asked as: speech to text use cases · applications of speech recognition · speech to text for accessibility · voice first apps · speech to text for meetings · speech to text for lectures · voice input forms

Speech to text versus text to speech, and other terms

Speech to text, STT or ASR, turns audio into words. Text to speech, TTS, turns words into audio. Voice cloning is TTS in a specific voice. Natural language understanding interprets the transcribed text. A voice assistant chains all of them: recognise, understand, respond, speak. People search for these interchangeably; they are different models with different failure modes.

Also asked as: speech to text vs text to speech · what is text to speech · tts vs stt · voice recognition vs speech recognition · difference between speech recognition and voice recognition · what is tts · nlp vs speech recognition

What are common speech recognition interview questions?

Explain the pipeline from waveform to text. Define word error rate and its three components. Explain what voice activity detection does and why chunking matters for latency. Compare self-hosting Whisper with an API. Describe how you would evaluate a model for Hindi. Explain what diarisation is and when it is unnecessary. Each answer is a section above, and the best proof is a live transcription demo you built.

Also asked as: speech recognition interview questions · asr interview questions · speech to text interview · audio ml interview questions

Where should I start with a speech project?

Record ten voice notes on your phone, transcribe them with faster-whisper on your laptop, correct the transcripts by hand, and compute the word error rate. Then try a smaller and a larger model and a second language. That afternoon teaches the whole page. For a hands-on session on voice-first apps, speech recognition and document AI at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (Whisper, 2022)arxiv.org
  2. Whisper repository, OpenAIgithub.com
  3. faster-whisper: CTranslate2 reimplementation of Whispergithub.com
  4. whisper.cpp: Whisper inference in C/C++github.com
  5. Baevski et al., wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations (2020)arxiv.org
  6. Graves et al., Connectionist Temporal Classification (2006)cs.toronto.edu
  7. Silero VAD: voice activity detectorgithub.com
  8. Bredin et al., pyannote.audio: neural building blocks for speaker diarization (2020)arxiv.org
  9. AI4Bharat IndicWhisper and Indic speech resourcesai4bharat.iitm.ac.in
  10. Mozilla Common Voice datasetcommonvoice.mozilla.org
  11. jiwer: word error rate computation in Pythongithub.com
  12. Mistral AI audio and transcription documentationdocs.mistral.ai
  13. OCR & Speech Workspace source code, Pranjul Rathourgithub.com
  14. Real-time speech-to-text architecture, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Vision, speech & OCR

All Vision, speech & OCR guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur