Skip to main content
Run diarization to get who-spoke-when segments, the number of speakers, and confidence scores — without separating the audio. It returns the same diarization JSON that Multi-Speaker Separation includes alongside its stems, so use multi_voice instead when you also need the per-speaker audio.

Create a Task

Check Task status to monitor progress and download results, or use webhooks to be notified when the target completes.

Output

The task returns one diarization JSON output:
diarization.json
All fields are relative to results.diarization:
  • segments — one entry per speech segment, with speaker and start/end in seconds. The timeline is overlap-aware: segments from different speakers can overlap when people talk at once.
  • num_speakers — number of distinct speakers detected.
  • scores — assignment_confidence in [0, 1], per-frame at 20 ms resolution and for the whole file: how confident the model is that speech is attributed to the right speaker.

Use cases

  • Attribute transcripts to speakers without paying for audio separation
  • Count speakers and measure talk time in meetings, calls, and interviews
  • Pre-screen recordings to decide which need full Multi-Speaker Separation

Multi-Speaker Separation

Get per-speaker audio stems along with the same diarization output.

Speech Recovery

Denoise and de-reverb speech recordings.