/tasks request to generate multiple outputs from the same file.
See Formats for supported input and output file types.
Instrument stem separation
Use these models to split songs into musical components for remixing, post-production, education, and interactive experiences.Speech
Models for multi-speaker separation and Speech Recovery (denoise and de-reverb).Maximum input length for
multi_voice is 1.5 hours.Post-production
Separation models for dubbing, dialogue cleanup, and audio editing workflows.Copyright compliance
Models for detecting, identifying, and removing music in content.Lyric transcription
Use these models to produce transcripts and time-synced text. These models are state of the art for lyric transcription.
For
alignment, provide one audio source (url or assetId) and optionally a transcript input (transcriptUrl or transcriptAssetId).
Maximum input length for
transcription and alignment is 45 minutes.