multi_voice instead when you also need the per-speaker audio.
Create a Task
Output
The task returns onediarization JSON output:
diarization.json
results.diarization:
segments— one entry per speech segment, withspeakerandstart/endin seconds. The timeline is overlap-aware: segments from different speakers can overlap when people talk at once.num_speakers— number of distinct speakers detected.scores—assignment_confidencein [0, 1], per-frame at 20 ms resolution and for the whole file: how confident the model is that speech is attributed to the right speaker.
Use cases
- Attribute transcripts to speakers without paying for audio separation
- Count speakers and measure talk time in meetings, calls, and interviews
- Pre-screen recordings to decide which need full Multi-Speaker Separation
Multi-Speaker Separation
Get per-speaker audio stems along with the same diarization output.
Speech Recovery
Denoise and de-reverb speech recordings.