Detecting AI music in MP3 requires a smarter dual-engine approach

In addressing the challenges of detecting AI-generated music, I encountered issues similar to those documented by Deezer regarding CNN-based detection on compressed audio.

3 min readMachine Learning

This is the kind of practical problem-solving that actually moves a field forward. The researcher here didn't try to brute-force a CNN into working under real-world conditions. They looked at how music is actually distributed, compressed MP3 files, not pristine WAVs, and built a system that acknowledges that reality. That distinction matters because it separates a lab demo from a tool that could function in the wild.

The core insight is elegant. Human-recorded music carries physical artifacts: room acoustics, microphone bleed, the messy interactions of instruments recorded together. When you separate that track into stems and reconstruct it, those imperfections create measurable differences. AI-generated music, by contrast, is built from independent synthetic layers. The separation and reconstruction process returns something nearly identical to the original. The dual-engine approach takes advantage of this fundamental difference without relying on fragile spectral patterns that compression destroys. The CNN handles the obvious cases quickly, and the more expensive reconstruction engine only activates when the first model is uncertain. That is smart resource management, not just clever detection.

For anyone distributing or consuming music at scale, this matters. The false positive rate sits at roughly one percent for human recordings, while catching over eighty percent of AI-generated tracks across multiple compression formats. Those numbers are not perfect, but they are functional. The limitations are honest: performance varies across different AI generators, the source separation model is non-deterministic, and this has only been tested on music, not speech or sound effects. That transparency is refreshing in a field often prone to overclaiming.

The real question is whether this approach can be hardened for production use. The non-deterministic behavior of Demucs on borderline cases is a genuine weakness. A system that flips its verdict between runs on the same file is not deployable at scale. But the hybrid architecture itself is worth exploring further. If the reconstruction analysis can be made deterministic, or if a third engine can be added to resolve those edge cases, this could become a practical standard. The work is not finished, but the direction is sound.

From Machine Learning

I've been working on detecting AI-generated music and ran into the same wall that Deezer's team documented in their paper, CNN-based detection on mel-spectrograms breaks when audio is compressed to MP3.

The problem: A ResNet18 trained on mel-spectrograms works well on WAV files, but real-world music is distributed as MP3/AAC. Compression destroys the subtle spectral artifacts the CNN relies on.

Read the original at Machine Learning