Nothing is noise

Speech carries more than words. nyra labs conducts research and builds open models that preserve what was said and measure when, who, and how - so real human speech can be understood in full.

CrisperWhisper 2.0: Every word. Every pause. Precisely timed.

CrisperWhisper 2.0 is a speech-to-text model that transcribes speech as it actually happened - with fillers, repetitions, false starts, and vocal sounds included. It is the first released layer of our research program.

Verbatim

Intended

Hover a word for its timestamps; click it to jump there.

Nothing gets lost

Word-level timestamps and best-in-class verbatim transcription accuracy across languages.

  • #1 verbatim transcription accuracy
  • #1 word-timing accuracy
  • 10+ languages

Word-level timing

The most accurate word-level timestamps on the market: 30 ms word boundaries via guided cross-attention.

Recover the verbatim layer

Turn existing spontaneous-speech corpora into faithful verbatim data with the Verbatimize task, while preserving the transcripts you already trust.

Open, fast, and production-ready

Open weights and code on Hugging Face, seamless longform, stronger hallucination resistance, and optimized inference make the model ready for real applications.

Performance you can verify

Open evaluation shows industry-leading transcription accuracy across languages and challenging real-world speech.

View the open benchmark

German

Disfluency F1 · higher is better
OpenAI Whisper v30.6
NVIDIA Canary-1B v21.0
AssemblyAI Univ-326.2
Inworld STT46.9
CrisperWhisper 1.058.3
Microsoft MAI-Transcribe85.0
CrisperWhisper 2.089.9
CrisperWhisper 2.0 Pro96.0

English

Disfluency F1 · higher is better
OpenAI Whisper v39.7
NVIDIA Canary-1B v226.6
AssemblyAI Univ-367.9
CrisperWhisper 1.071.4
Microsoft MAI-Transcribe84.0
Inworld STT84.4
CrisperWhisper 2.090.7
CrisperWhisper 2.0 Pro93.2

Our research program

From spoken detail to speech understanding

Our program follows a deliberate order: preserve speech, measure its structure, separate the factors it carries, and build models that can learn from the whole signal.