Our research program

From spoken detail to speech understanding

We work along two connected lines: measuring real speech faithfully and representing what it carries interpretably. Together, they form our path toward speech health foundation models.

Measuring speech faithfully

Pinning down what actually happened in an utterance: every word, its timing, its speaker, and how it was said.

The what

Shipped

Verbatim recognition with CrisperWhisper 2.0: every word exactly as spoken - fillers, repetitions, false starts, and vocal sounds included.

The when

In development

Phoneme-level timing from the nyra forced aligner, robust to spontaneous, disfluent, real-world speech - not just read-aloud recordings.

The who

In development

Native diarization: every word attributed to the right voice, with overlapping speech kept, not merged away.

The how

Next

Prosody and the clinical dimensions of the voice, measured at the level where each phenomenon actually occurs.

With the what, the when, the who, and the how in place, labeling stops being the bottleneck: we can annotate real-world speech at scale, at a quality that today only expert annotators reach.

Representing speech interpretably

Compressing speech into representations where each attribute lives in its own place, at its own natural granularity, instead of everything entangled in one stream.

Content

Published

PINT, our parallel-invariant tokenizer, distills the one factor many voices share - the words - and discards the rest.

Fully factorized codecs

In progress

Codecs that give content, speaker, prosody, and acoustics each their own place, at their own natural granularity.

where both lines converge

Speech health foundation models

Faithful measurement gives us annotation of real-world speech at expert quality and scale; factorized representations give us the substrate to learn from it. Together they are the data and the architecture we intend to build speech health foundation models on.