Home › How I test
Methodology
How I test AI music generators
Five criteria, forty identical briefs, three runs each, and the prices read off the vendor's own page.

Every tool on this site is scored on the same five criteria, weighted the same way, against the same set of briefs. The weights are below and they do not change between reviews. If a score moves, it is because a vendor changed something.
The rubric
The briefs
Forty jobs, identical for every tool, chosen to look like what people actually need rather than what demos show: three birthday songs with names in the lyrics, a wedding first dance, a leaving-party song with an in-joke, two podcast beds, a sixty-second product-video score with a mood change at thirty-eight seconds, four looping background tracks, a lullaby, and twenty-six more across genres from country to drum and bass.
I run each brief three times per tool. I keep the best of the three, which is generous to every tool equally, and I record how many of the three were usable at all. That second number is what the output score mostly reflects, because a tool that lands one in three is a different product from one that lands three in three even when their best takes are identical.
Prices and rights
Every price on this site is read off the vendor's own pricing page on the date shown at the top of each article, currently 21 September 2026. I never repeat a price from press coverage. Where a vendor advertises one price and puts commercial rights on a higher tier, I say so in the review and it costs them value points.
What I pay for, and what I take
I buy the subscriptions I test. No vendor on this site has been given a free account, an embargoed preview or any opportunity to see a review before it was published, and no vendor pays for placement or for a score.
There are no advertisements on this site, no sponsored posts and no affiliate links. I name tools in plain text and do not link out to them, so there is nothing on any page here that can earn money from a click. The rubric above is the only thing that determines an order. The current standing is Udio (83), ElevenLabs Music (82), Suno (79), Soundraw (68), Beatoven.ai (59), Boomy (55), Mubert (50).
Why the weights are what they are
Output quality carries thirty points because it is the reason anybody opens these tools, and only thirty because after five weeks I could not reliably tell the top three apart on a phone speaker. Rights and licensing carries twenty, which surprises people until they discover that the tier they bought is not the tier the commercial licence lives on.
Value carries fifteen and is scored against the real price of the thing you actually need, not the headline. Mubert loses value points because commercial use costs 32.49 USD and the pricing page leads with 14 USD. That is a deliberate choice in the rubric and it is the single most common way these tools mislead a buyer.
What I do not score
I do not score interface polish, community size, roadmap promises or anything a vendor says it is about to ship. I have been wrong before about what a company would ship next, and a score that moves on a press release is worth nothing.
I also do not score the thing everybody asks me about, which is whether a tool sounds like a particular artist. It is not a useful axis, it changes weekly, and chasing it is how a review turns into a highlight reel.
When scores change
A score moves when a vendor changes a price, a tier, a licence or a capability I tested. It does not move because a competitor released something. When a score changes I note the date on the review, and the price-checked date at the top of every page tells you how old the numbers you are reading are.
Corrections
If a price here is wrong or a vendor has changed its terms, tell me and it gets fixed with the date of the change noted on the page. Write to the address on the about page.
