← All projects

Speech / Evaluated demonstration

Speech Studio

Turn text into controllable speech.

Real demo recording · English narration and captions · Fictional data

THE PROBLEM

Internal workflows can need synthetic narration with defined voices and output formats.

THE WORKFLOW

Prepare English text, adjust voice and speed, synthesize a local WAV, listen and download the audio.

  1. 01English text
  2. 02Voice settings
  3. 03Local synthesis
  4. 04Playback and download

RECORDED EVIDENCE

Results, with context.

Results on synthetic data
CheckCases
Valid PCM format8/8
24 kHz sample rate8/8
Audio signal present8/8
Dataset
synthetic-regression-v3-english
Completed cases
8/8
Median / p95 per case
0.206 s / 8.76 s
Evaluation record (UTC)
2026-10-08

Source: the project’s evaluation record, using local inference. Recorded hardware: NVIDIA GeForce RTX 5060 Ti, 595.84, 16311 MiB. Latency includes the full HTTP case and may include initial model loading. These measurements have not been rerun on the machine serving this website.

WHAT WE COULD BUILD WITH YOU

A private text-to-speech integration with local voices whose licences are reviewed for the intended use.

What we would define first

Voice and model licences, languages, pronunciation and audio format.

Demonstration limits

Eight tests check PCM output at 24 kHz, duration and signal. They do not measure perceptual quality or clone voices.

Discuss your use case ↗
Screenshot of the interface for Speech Studio
Demonstration review interface.

Explore other projects