I produce music in Ableton, and I have spent a lot of time with Suno. It’s fun to toy with ideas and get inspiration, but I haven’t been able to actually produce any good music. I would argue that no producers are making good music on Suno. My problem is broader: most creative AI tools — and AI tools in general — don’t understand or can’t produce the layer that creatives actually want to work with. Suno is just the clearest place to see it.
Suno is best described as a consumer experience company. Its CEO is explicit that the future belongs to taste rather than skill, and that most people don’t even enjoy the process of making music. Judged on those terms, the product works. Around a hundred million people have used it, the generation loop is fun, and someone who has never opened a DAW can make a song that sounds finished in two minutes. Finished output is the product, and for most users that’s exactly right.
The problem appears when you try to actually work with what it makes. And to see why, you have to understand something about how professionals create anything.
A producer doesn’t work with a finished song. They work with the layer before it — MIDI notes, individual tracks, instruments, effects chains, an arrangement they can move around. The waveform you eventually hear is the last step, the render. The work happens upstream of it, in a representation where every element is separate and adjustable. Don’t like the bass tone? Swap the instrument and add effects. Melody comes in too early? Shift the notes. Drums too loud? Move one fader. None of that touches the final audio directly, because the final audio isn’t where you work. It’s what you export when you’re done.
This is the layer Suno skips. It generates audio first, then reconstructs the editable layer afterward. Stems get separated out of a finished mix. MIDI gets transcribed from those stems by analyzing the audio. I’ve run the exports and they always feel off. Chords come back incomplete, rhythms drift, and complex parts get simplified until they’re unusable. Edit one line and the model re-rolls the render, changing a feel you already locked in. There is no project file underneath, only guesses at one. You’re handed the output and asked to reverse-engineer the work from it.
The obvious counterargument is that this gets better with time. Stem separation improves every year. Transcription gets sharper. Give it a few releases and the reconstruction will be clean enough that the distinction stops mattering.
Even with improvement, I still have doubts. Every gain here is a post-hoc approximation of something you could have just had natively, and approximations hit an asymptote. Going from “can’t isolate instruments at all” to “rough stems” was a huge leap. Going from rough stems to perfect ones is exponentially harder, and even if you got there, it still wouldn’t give you tempo changes, arrangement edits, or instrument swaps. The gap doesn’t close. It moves. Each time the pipeline gets good at one thing, the next thing a producer needs is still locked inside the render.
More critically, you can’t compose flat outputs the way you can compose structured ones. Even with flawless separation, you can’t take the bassline from one generation, the drums from another, and the melody from a third and combine them cleanly. That requires the parts to exist as parts. Reusing a phrase across arrangements, dropping a generated element into an existing session, building a library you compound on over time — all of it depends on structure that a finished mix simply doesn’t carry. Composition, reuse, and precision are the entire point of professional work, and they’re exactly what audio-first generation can’t deliver.
The change I would make is to flip the architecture. Generate natively at the symbolic layer, with composition, arrangement, and sound parameters as the first-class output and audio as the render, not the other way around. Generate the MIDI, the instrument choices, the effects chain, the structure, and let the audio fall out of that the same way it does in a real session. That’s what producers actually work with, and it’s what turns Suno from a “slot machine” you visit into a workspace you stay in.
Suno building Studio, with its stem separation and MIDI export, is itself the tell. The most successful company in AI music felt compelled to bolt a structured layer onto a flat product, because professionals kept needing it. But bolting structure on after the fact is not the same as generating it from the start, and the information lost in waveform to stems to MIDI is lost for good. The company that generates the structure natively wins the part of the market that Suno can only approximate.
This is the strategic point underneath the technical one. Suno wins today on consumer delight, and that’s a good business, but the durable business in AI music belongs to whoever owns the layer where the work actually happens.