Skip to main content
AudioShake models define the type of processing applied to your source audio. You can combine multiple models in a single /tasks request to generate multiple outputs from the same file. See Formats for supported input and output file types.

Instrument stem separation

Use these models to split songs into musical components for remixing, post-production, education, and interactive experiences.

Speech

Models for multi-speaker separation and Speech Recovery (denoise and de-reverb).
Maximum input length for multi_voice is 1.5 hours.

Post-production

Separation models for dubbing, dialogue cleanup, and audio editing workflows. Models for detecting, identifying, and removing music in content.

Lyric transcription

Use these models to produce transcripts and time-synced text. These models are state of the art for lyric transcription. For alignment, provide one audio source (url or assetId) and optionally a transcript input (transcriptUrl or transcriptAssetId).
Maximum input length for transcription and alignment is 45 minutes.