

CrisperWhisper 2.0: Every word. Every pause. Precisely timed.
CrisperWhisper 2.0 is a speech-to-text model that transcribes speech as it actually happened - with fillers, repetitions, false starts, and vocal sounds included. It is the first released layer of our research program.

Hover a word for its timestamps; click it to jump there.

Nothing gets lost
Word-level timestamps and best-in-class verbatim transcription accuracy across languages.
- #1 verbatim transcription accuracy
- #1 word-timing accuracy
- 10+ languages
Word-level timing
The most accurate word-level timestamps on the market: 30 ms word boundaries via guided cross-attention.
Recover the verbatim layer
Turn existing spontaneous-speech corpora into faithful verbatim data with the Verbatimize task, while preserving the transcripts you already trust.

Open, fast, and production-ready
Open weights and code on Hugging Face, seamless longform, stronger hallucination resistance, and optimized inference make the model ready for real applications.
Performance you can verify
Open evaluation shows industry-leading transcription accuracy across languages and challenging real-world speech.
View the open benchmarkGerman
Disfluency F1 · higher is betterEnglish
Disfluency F1 · higher is betterOur research program
From spoken detail to speech understanding
Our program follows a deliberate order: preserve speech, measure its structure, separate the factors it carries, and build models that can learn from the whole signal.
Latest releases
CrisperWhisper 2.0
World's #1 verbatim speech recognition: every filler, pause and repair, exactly as spoken.
The nyra verbatim speech benchmark
Fifteen ASR systems scored on the fillers, repetitions and cut-offs that make speech real: 20+ typed metrics, not one WER.
Word-level timing from attention
Word boundaries to within 30 ms, supervised straight from the cross-attention heads that already lean toward alignment. No separate aligner.