The what
ShippedVerbatim recognition with CrisperWhisper 2.0: every word exactly as spoken - fillers, repetitions, false starts, and vocal sounds included.
Our research program
We work along two connected lines: measuring real speech faithfully and representing what it carries interpretably. Together, they form our path toward speech health foundation models.
Pinning down what actually happened in an utterance: every word, its timing, its speaker, and how it was said.
Verbatim recognition with CrisperWhisper 2.0: every word exactly as spoken - fillers, repetitions, false starts, and vocal sounds included.
Phoneme-level timing from the nyra forced aligner, robust to spontaneous, disfluent, real-world speech - not just read-aloud recordings.
Native diarization: every word attributed to the right voice, with overlapping speech kept, not merged away.
Prosody and the clinical dimensions of the voice, measured at the level where each phenomenon actually occurs.
With the what, the when, the who, and the how in place, labeling stops being the bottleneck: we can annotate real-world speech at scale, at a quality that today only expert annotators reach.
Compressing speech into representations where each attribute lives in its own place, at its own natural granularity, instead of everything entangled in one stream.
PINT, our parallel-invariant tokenizer, distills the one factor many voices share - the words - and discards the rest.
Codecs that give content, speaker, prosody, and acoustics each their own place, at their own natural granularity.
where both lines converge
Faithful measurement gives us annotation of real-world speech at expert quality and scale; factorized representations give us the substrate to learn from it. Together they are the data and the architecture we intend to build speech health foundation models on.