Research
Models, benchmarks, datasets, and the publications behind them. Everything here serves one program: understanding speech holistically.
Featured
CrisperWhisper 2.0
The most accurate verbatim speech recognition you can run in production: controllable, multilingual, and timed to the word.
Measuring verbatimness
An open suite that scores not just the words, but the fillers, repetitions, cut-offs, and vocal sounds that make speech real.
Turning emergent cross-attention into a precise aligner
How CrisperWhisper 2.0 reads millisecond-accurate word timings out of an ASR decoder's cross-attention: by making the output policy explicit, then supervising the heads that already lean toward alignment.
Closing the verbatim data gap
How to upgrade existing spontaneous-speech corpora into faithful verbatim datasets while preserving the transcripts you already trust.
Longform transcription with conditional continuation
How CrisperWhisper 2.0 transcribes audio longer than 30 seconds: each window is prompted with the last words of the previous one and trained to continue them, with no timestamp tokens and predictable batching.
All research
How CrisperWhisper 2.0 transcribes audio longer than 30 seconds: each window is prompted with the last words of the previous one and trained to continue them, with no timestamp tokens and predictable batching.
The most accurate verbatim speech recognition you can run in production: controllable, multilingual, and timed to the word.
An open suite that scores not just the words, but the fillers, repetitions, cut-offs, and vocal sounds that make speech real.
How to upgrade existing spontaneous-speech corpora into faithful verbatim datasets while preserving the transcripts you already trust.
How CrisperWhisper 2.0 stops Whisper from inventing text on silence and from looping, and how we made it fast enough for production with a CTranslate2 port and speculative decoding.
How a handful of decoder-prefix tokens turn transcription style into an explicit switch, activating verbatim and intended transcription that Whisper already latently knew, and transferring it across languages zero-shot.
How CrisperWhisper 2.0 reads millisecond-accurate word timings out of an ASR decoder's cross-attention: by making the output policy explicit, then supervising the heads that already lean toward alignment.
A parallel-invariant speech tokenizer: when many voices say the same words under different conditions, content is the only shared factor. PINT keeps it and discards the rest.
Phoneme-level forced alignment for real speech: robust to the spontaneous and imperfect, and explicit about the fillers and vocal sounds that break other aligners.