Music & Sound with AI
An original soundtrack, a jingle or a full song — from text, in minutes. How to create AI music and use it correctly (and legally).
Disclosure: this page contains affiliate links. If you sign up through them we may earn a commission — at no extra cost to you. We only recommend tools we've tested. Details ›
What AI music is
AI music tools generate a finished track — melody, instrumentation, and in some cases vocals singing lyrics you supply — from a text description. You write "energetic pop with electric guitar, fast beat" and get audio back in minutes, with no musical training and no studio.
Music made for your video, at the length your video actually is, without a subscription per track. That last part is the genuine advantage — the legal picture is more complicated, and the next two sections are about that.
"No copyright strikes" is the claim to be careful with
The usual pitch for AI music is that generating your own removes any risk of a claim on your video. That is partly true and is worth stating precisely, because people make commercial decisions on it.
What it does remove: the risk that you used someone else's recording without a licence. That is the most common cause of a claim, and generating your own genuinely avoids it.
What it does not remove:
- Automated matching is imperfect. Rights-detection systems compare audio fingerprints, and generic material — a common chord progression at a common tempo — can trigger a match against something in the database. Claims like this are usually disputable, and disputing takes time you were not planning to spend.
- Someone else can register music you generated. This is the failure worth knowing about: people have registered AI-generated tracks into rights databases, after which anyone else who generated something similar can be claimed against. You are not protected by having made it first if you never registered it and cannot easily prove the date.
- Platform rules on synthetic audio are their own layer, separate from copyright, and they have been changing. Check the current policy where you publish.
The practical habit: keep your generation records — the prompt, the tool, the date, the original file. It costs nothing and it is what resolves a dispute quickly.
Do you actually own it?
This is unsettled, and anyone telling you otherwise with confidence is overstating the position.
Two separate questions get mixed together. The first is what the tool permits — that is a contract between you and the vendor, it is written in the terms, and it is the one that governs your day-to-day use. Read it: commercial rights usually depend on the plan, free tiers often reserve them or require credit, and terms change between versions.
The second is whether copyright subsists in the output at all, and in several jurisdictions the answer for purely machine-generated work is uncertain or restrictive — with human authorship generally being what the law protects. There is also ongoing litigation about the material these models were trained on, whose outcome nobody can predict.
What that means in practice, ranked by exposure:
- Background music in your own videos — low risk. The vendor's licence covers your use, and if it later turned out you could not stop someone else using the same track, you would not care.
- A jingle or sonic logo for a brand — medium. Here you may want exclusivity and the ability to stop others using it, which is exactly the part that may not exist.
- A track you intend to release and monetise as music, or client work where you promise ownership — treat carefully, and do not promise a client rights you have not verified you can transfer.
None of this is legal advice, and it differs by country. For anything where ownership matters commercially, that is a conversation with someone qualified rather than a paragraph on a guide.
Tools — Suno and others
- Suno — the standout for full songs with vocals and lyrics from text. Wide stylistic range.
- Udio — a direct competitor with an emphasis on audio quality and control.
- Background-music tools — instrumental beds for video and podcast, often with length controls built for editors.
Before committing to any of them, check three things that matter more than the demo: whether your plan grants commercial use, whether you can download stems (see below), and what happens to tracks you generated if you stop paying.
Writing a prompt that gets you something usable
The prompt decides the result more than the tool does. Describe:
- Genre and style — "lo-fi hip hop", "epic orchestral", "80s synthwave".
- Mood — energetic, calm, tense, warm.
- Instrumentation — piano, acoustic guitar, analogue synth, brushed drums.
- Tempo and feel — slow and spacious against fast and driving.
- Lyrics — write your own and have the tool sing them.
The parts people leave out
- Say what you do not want. "No vocals", "no drums for the first thirty seconds", "no big finish" — for music under narration, the exclusions matter more than the inclusions.
- Describe the production, not just the genre. "Recorded in a small room, slightly imperfect timing" produces something different from a polished master, and often something that sits better under a voice.
- Structure tags, where the tool supports them — intro, verse, chorus, break. Without direction these tools tend to reach a chorus fast, which is wrong for background use.
- Reference by description, not by artist. Naming a living artist is restricted by most tools' terms and is the kind of output likeliest to cause trouble later. "Sparse piano with tape hiss" gets you there without the problem.
- Generate several and extend the good one. Most tools can continue or vary an existing track — that is how you get a three-minute bed that stays coherent rather than three unrelated minutes.
What these tools are still bad at
Knowing the failure modes saves you from discovering them in an edit.
- Endings. Generated tracks often fade oddly or stop abruptly. Plan to fade it yourself in the edit.
- Long-form coherence. Beyond a couple of minutes, material drifts. Extending deliberately beats asking for length in one go.
- Vocal artefacts. Consonants smear, words get invented, and it is worst at speed. Listen on headphones before publishing — laptop speakers hide exactly these.
- Precise timing. Asking for a hit on a specific beat to match a cut is not something you can reliably prompt for. Edit the audio to the video instead.
- Silence and space. These models fill. If you want a sparse bed, ask for sparse explicitly and expect to try several times.
Stems are the feature that decides whether it is usable
A stem is one element of the mix exported on its own — drums, bass, vocals, the rest. Whether you can get them changes what you can do with the track, and it is buried in the plan comparison rather than on the marketing page.
With stems you can drop the drums under a voiceover and bring them back for the outro, remove a vocal that competes with narration, or build a version at a different length without it sounding cut. Without stems you have one stereo file and your only controls are volume and fades.
For anything going under speech — video, podcast, advertising — stems are worth more than a marginally better generation.
Putting music under a voice
The most common use, and the one most often done badly. Generated tracks arrive mastered loud, because that is what music is supposed to sound like on its own — which is exactly wrong underneath a narrator.
- Pull it well down. Background music should be clearly quieter than the voice — if you are aware of it while listening to the words, it is too loud.
- Duck it. Every editor has automatic ducking: the music drops when someone speaks and comes back in the gaps. This one setting does more than any amount of regenerating.
- Cut the midrange out of the music if you can. The voice lives there, and music competing in the same range is why a mix feels muddy even at low volume.
- Mind the platform's loudness normalisation. Platforms adjust playback level; a track mastered very loud does not end up louder, it ends up more squashed.
- Check on phone speakers. Most people will hear it there, and it is where a bass-heavy bed disappears and a bright one becomes harsh.
Writing lyrics the model can sing
If you are using the vocal side, the lyrics are where most of the quality lives — and lyrics that read well often sing badly, which surprises people the first time.
- Count syllables, not words. Lines of roughly equal length sing; wildly uneven ones get compressed into a mumble to fit the bar.
- Avoid consonant clusters. Words that are hard for a person to sing are harder for a model, and they are where the artefacts appear.
- Repeat the chorus verbatim. Slight variations between choruses cause the generation to drift, and repetition is what makes a chorus a chorus anyway.
- Keep proper nouns out where you can. Brand names and unusual names are mispronounced confidently, exactly as they are in text-to-speech.
- Read it aloud, in rhythm, before generating. Anything you stumble over will come back mangled, and generating costs credits while reading costs nothing.
If you are drafting lyrics with a model, give it the syllable constraint and the rhyme scheme explicitly. Asked for "a chorus about resilience" it produces something generic and unsingable; asked for "four lines, seven to nine syllables each, ABAB, plain words, no metaphors about fire or storms" it produces something you can use.
Use cases
- Video soundtracks — video and short-form, at the exact length you need.
- Podcast intro and outro — see podcasts with AI. A consistent theme across episodes is worth more than a better one-off.
- Jingles and sonic branding — with the ownership caveat above if it is to become a brand asset.
- Background music for a space or an event — check whether local public-performance rules apply to the venue regardless of the music's source.
- Demos and scratch tracks — a quick sketch to show a client or a composer what you mean. Genuinely useful and completely uncontroversial.
A workflow that ends with a track that fits
The usual path — generate, download, drop it in the timeline, wonder why it feels wrong — wastes most of the advantage. This one takes about twenty minutes and produces music that belongs to the video.
- Edit the video first. You cannot write to a length you have not decided, and cutting picture to music is the harder direction.
- Write down what the music has to do in one sentence: carry energy under a fast montage, sit invisibly under narration, mark a transition. Different jobs, different prompts.
- Generate three or four candidates from variations of the same prompt, not one prompt refined four times. Breadth first.
- Audition them against the actual video, at the volume they will be heard, not in the generator's player. Most candidates that sound best alone lose here, because they are doing too much.
- Extend the winner to the length you need, rather than looping a short clip — loops are audible and the seam is where attention goes.
- Download stems if your plan has them, then duck, filter and fade in the edit.
- Save the prompt and the date with the project file. This is your record if a claim ever arrives, and your starting point when the next video needs the same feel.
That last step is the one that compounds. After a few projects you have a short list of prompts that reliably produce music that works under your voice, in your style — and that list is worth considerably more than any individual track.
When a licensed library is the better answer
This page carries an affiliate link, so it owes you the cases where you should not use these tools.
- When you need certainty. A subscription music library gives you a clear licence, indemnity in some cases, and a paper trail. For client work or advertising, that certainty is often worth more than the saving.
- When you need genuinely good music. Library catalogues are full of work by real composers and the ceiling is higher, particularly for anything exposed — where the music is the point rather than a bed.
- When you need it to be exclusive. Neither route gives you much of this, but a library at least tells you the terms.
- When you need one track, once. A single licence may cost less than a month of a subscription you will forget to cancel.
The strongest case for generating is volume and fit: a lot of short pieces, each cut to a specific video, where the music's job is to not be noticed.
Saying it is AI
Platform requirements around labelling synthetic media have been expanding, and music sits inside that in some contexts — particularly anything with a generated vocal, and absolutely anything imitating a real performer's voice, which is separately prohibited by most tools' terms and by law in a growing number of places.
For a music bed under a video, disclosure is usually not required and not expected. For a released track, or one with vocals that a listener might take for a person singing, labelling it is both increasingly required and simply the honest thing — and it removes the story of having been caught not saying so.
The part the tool cannot give you
Generation removes the barrier to producing music. It does not give you the ear to tell whether a piece is right for the thing it is sitting under, and that judgement is now the whole job.
The failure is specific and common: creators pick the track they enjoyed most in the generator, which is usually the busiest and most interesting one, and it fights the video for the rest of its life. Background music is supposed to be slightly boring. A bed that you notice is a bed that is taking attention from the words you spent all day writing.
Two habits build the judgement faster than listening to more output:
- Pay attention to music in work you admire. Watch something well made and notice when the music enters, when it drops out entirely, and how much quieter it is than you would have guessed. The silences are usually the surprise.
- Try the version with no music at all. Often it is better, and comparing against it stops you adding a bed out of habit. Music should earn its place by doing something — carrying a transition, setting a tone, covering a cut — rather than being present because tracks are free now.
Common mistakes
- Commercial use without reading the plan. The most expensive error, and it is a two-minute check.
- Assuming generated means claim-proof. It removes the main risk and not all of it. Keep your generation records.
- A vague prompt. Genre, mood, instruments, tempo — and what to leave out.
- Taking the first result. Generation is cheap; generate several and extend the one that works.
- Leaving it mastered-loud under a voice. Duck it, drop it, and check on a phone.
- Promising a client ownership you have not verified you can transfer.
Next step
Add the soundtrack to your videos and podcast, or explore more tools.