Music doesn’t have to end at the final mix
- Written by
- Shaad Sufi
- Published
- Reading time
- 6 min read

From carrying music to shaping it
On October 23, 2001, Apple introduced the iPod with a promise you could understand without knowing anything about audio engineering: up to 1,000 songs in your pocket. A collection that once occupied shelves could come with you on the train.
Twenty years later, Apple announced Spatial Audio with Dolby Atmos for Apple Music, including playback on compatible AirPods. The experience had expanded again: instruments could seem to come from around you, rather than simply from the left or right.
Yet consider a familiar moment. You are learning a bass line, but the vocals keep pulling your attention away. Or you are editing a video and the music’s lead melody competes with the dialogue. Turning the whole song down only gets you so far.
At Wubble, our SNC research explores a question behind those moments: what if the music file kept its individual parts available, so the experience could change with the listener’s needs?
What we changed inside the music file
A song begins as parts: a voice, a drum pattern, a bass line, and the instruments around them. Producers work with those parts separately. When we receive a typical stereo download, we receive their combined mix.
Our research asks what happens if those parts travel with the finished song. We call the approach SNC, the Stem-Native Codec. Music stems are separate audio layers that a player can address individually. SNC packages them together in one file.
For our experiment, we separated a recording into four stems: vocals, drums, bass, and other instruments. We compressed each one, then added a fifth track containing the difference between the combined stems and the reference mix. This correction layer helps account for changes introduced by separation and compression.
The motivation is practical. Keeping the layers available could let someone adjust the music without having to separate it all over again. The file also carries information about playback, including spatial positions and artist-defined adjustment limits. The music’s structure becomes something an application can work with.
What the numbers show
We tested the approach on an electronic/rock track lasting 138.71 seconds—a little over two minutes. The resulting SNC file was 7.76 MB, including all four stems and the correction layer.
The FLAC version of the reference recording occupied 12.55 MB. SNC used 4.79 MB less space, a 38.2% reduction, while retaining separate access to the musical parts. The correction track accounted for 1.05 MB, or 13.5% of the SNC file.
The stereo Opus and MP3 versions were smaller, at 4.39 MB and 5.29 MB respectively. That comparison makes the choice tangible: how much space would you set aside to keep a song adjustable? These encodings make different choices about audio quality; the chart compares storage, not equivalent fidelity.
For this track, our experiment puts a number on the cost of carrying the layers together. Wider testing and listening studies will tell us how that balance changes across recordings. The figures below come from Table III of our paper.
What this could mean across industries
The same song can serve very different purposes. Once its parts remain accessible, the question for an application becomes: which parts does this moment need? Here are four ways that could matter.
Games: a soundtrack that follows the scene
Imagine walking through a quiet forest in a game. The melody is gentle. As danger approaches, percussion enters and the bass grows stronger, while the underlying theme stays familiar. Game audio teams already work creatively with musical layers. SNC explores a way to carry those layers and their playback information together, for a game engine to interpret. Its value here would be in how the music is delivered and controlled.
Film and advertising: room for the story
An editor finds the right track, but its vocal gets in the way of a line of dialogue. Access to stems could let them lower that vocal while keeping the rhythm and atmosphere. For a film team or a brand making several versions of a campaign, one music package could support different balances for a spoken introduction, a product reveal, and a closing scene.
Music education: hear it, then play it
A student could bring the bass forward to hear a phrase, then lower it and play the part themselves. A singer could rehearse with the accompaniment while reducing the recorded vocal. The recording becomes a practice partner, with the learner choosing which part to focus on.
Listening platforms: more personal playback
Someone listening on a busy commute might want the voice more prominent. At home, they might prefer the original balance. A compatible player could offer those choices using the same file. Spatial playback could also position individual layers around a listener. These applications would need support from the player or editing software; they are directions opened up by access to the stems.
Keeping the artist in the conversation
A mix is a creative decision. The way a voice sits behind a guitar, or a drum enters a chorus, can be part of what makes a song work. Giving listeners more control should leave room for those choices.
That is why our paper includes artist-defined adjustment limits alongside the audio. An artist could offer a practice mode or allow a small vocal adjustment while keeping other balances fixed. A player would need to interpret those instructions, and any reuse would still depend on the rights attached to the recording.
For the industry, this opens a useful conversation: what kinds of participation would artists want to offer? A lesson, an interactive soundtrack, and an immersive listening session each call for a different relationship with the music.

What comes after pressing play?
The iPod made a music collection portable. Spatial audio gave listeners another way to experience a recording. Our question at Wubble is about what happens when the parts of a song remain within reach.
SNC brings four stems, a correction track, and playback information into one package. Our experiment gives us an initial storage comparison. The next work is to test more recordings, listen carefully across devices, and explore how applications should handle changes to the mix.
For our work in music creation, the possibility is compelling: a track could keep more of its creative usefulness after export. The same recording might accompany a singer, follow a player through a game, or give an editor room for dialogue.
We spent years finding ways to take more music with us. It is worth asking how much more we could do with each song.
Read Wubble’s SNC research on arXiv for the methods and full results. This article draws on version 1, submitted on February 8, 2026.


