Skip to main content

Category

Audio source separation in production: Demucs, UVR5, vocal isolation and ASR preprocessing

Source separation splits one audio track into its parts — voice, drums, bass, accompaniment. The applications are wide: karaoke tracks, dubbing a video into another language while keeping its music, lifting transcription accuracy in noisy recordings, remixing and transcription by ear. This cluster is built around Demucs v4, the strongest openly available model, and UVR5 (MDX-Net) for vocal work: how to pick a tool from your requirements, how to wire it into an ASR preprocessing pipeline, and how to run it in production.

12 articles in total

Foundational guide

Foundational guide (start here)

音源分離
Demucs
UVR5
技術選定
音声処理

How to choose a source-separation tool: selecting Demucs / UVR5(MDX-Net) / Spleeter / Open-Unmix by requirements

A cross-comparison of the major music-source-separation OSS — Demucs v4, UVR5(MDX-Net), Spleeter, Open-Unmix — by quality, speed, license, setup difficulty, and memory. It explains, with real code, a decision framework you can reverse-look-up from requirements ('which to choose for which project') and the license pitfalls you must always confirm for commercial use.

11 min read

Related practical articles