Skip to main content
Level: Advanced Updated: September 2026

AI Avatars — A Talking Head

Turn text into a video with a talking presenter — no camera, studio or crew. The technology, the tools and the use cases.

What an AI avatar is

An AI avatar is a video of a human figure — real or synthetic — that "speaks" text you wrote, with lip movement, expression and a generated voice. You write a script, choose a figure and a voice, and the system renders a finished video. No camera, lighting, location or crew.

The real advantage

Not that the first video is cheap — that the fiftieth revision is. Change a line of text, regenerate, done. That is a different economics from filming.

What avatars are actually for

The tempting summary is "video without a shoot", and it leads people to use avatars for the wrong things and conclude the technology is not ready. It is ready — for a specific job.

The job is delivering information that changes. A compliance module that gets rewritten every quarter, a product walkthrough that goes stale with each release, onboarding that mentions a tool you no longer use. Filmed video is expensive to change, so in practice it never gets changed, and organisations quietly serve outdated training for years. An avatar video is edited by editing text. That is the whole value proposition and it is a real one.

The job it does badly is anything that has to be believed, felt or trusted. A founder's message about layoffs. A sales pitch that depends on rapport. A thank-you to customers. Here the format works against you: a synthetic person expressing sentiment reads as someone who could not be bothered to say it themselves, and that is exactly the message received regardless of what the words said.

So the test before you generate anything: is this video transferring information, or asking for trust? Information is an excellent fit. Trust is not, and no improvement in rendering quality changes that — the problem is not that it looks fake, it is that it is not you.

Where the illusion breaks

Current avatars are convincing in short bursts and get less convincing the longer they run. Knowing the specific failure points lets you design around them.

The design response is structural: keep segments short, cut to screen recordings and graphics often, and let the avatar introduce and hand off rather than present continuously. A talking head on screen for the entire runtime is the hardest possible version and the one people try first.

Writing for an avatar is not writing for a page

Most disappointing avatar videos are script problems wearing a technology costume. A few mechanical things change the output more than switching tools would.

Voice is most of the impression

People say "the avatar looks fake" when what they actually noticed was the voice. Audio carries more of the credibility than the face does, and it is the cheaper half to get right.

Match the voice to the figure — an accent or an apparent age that contradicts the face creates a mismatch viewers feel without identifying. Listen to the generated audio on its own, without the video, before you approve it; problems that hide behind a moving face are obvious with your eyes shut. And if you clone your own voice with ElevenLabs or similar, record the source samples properly: a quiet room and a decent microphone for a few minutes beats an hour of noisy audio, because the clone reproduces the room as faithfully as the voice.

Cloning yourself, and the part nobody plans for

Most platforms will build an avatar from a few minutes of footage of a real person. For a company this is attractive: your actual trainer, your actual founder, in videos you can update forever.

Forever is the problem.

A cloned likeness is an asset that outlives the relationship. The employee who recorded it leaves, and their face is still delivering your onboarding. They change roles, or fall out with the company, or simply stop wanting their face used — and the clone is in fifty modules. This is not hypothetical; it is the predictable end state of the useful version of this feature.

Handle it before you record, in writing, covering four things: what the likeness may be used for (internal training only, or public marketing too), for how long, what happens when the person leaves, and how they withdraw consent and what you undertake to do then. Give the person a copy. If your organisation has an HR or legal function, this is a conversation for them and not a checkbox in a tool.

Add an operational habit: keep a register of every video that uses a cloned person, so that honouring a withdrawal is a list to work through rather than an archaeology project.

Translation and lip sync

Translating a video into several languages with matched lip movement is the most impressive thing these tools do, and it is genuinely useful for training material that has to reach several markets.

Three practical cautions. Length changes. The same sentence takes noticeably longer in some languages than others, which pushes your timing out of step with anything on screen — so leave slack, and avoid tight synchronisation between narration and visuals. Translation is not localisation. Examples, currencies, regulations, names and humour need adapting, not converting, and a machine translation will carry all of them across unchanged. Have a native speaker watch it before it ships. Not read the transcript — watch it, because the errors that matter are tone and register, and those do not show up in text.

Worth being clear-eyed about the last one: a translated avatar video that no fluent speaker has reviewed is a public statement in a language you cannot check. The cost of that review is small and the alternative is discovering the problem from the audience.

Leading tools

They differentiate less on rendering quality — which converges quickly — than on the workflow around it: template and brand controls, how review and approval work, how many seats, and what the licence says about commercial use and about cloned likenesses. If this is for a company rather than a personal channel, read the likeness and data terms before comparing avatar galleries.

Workflow

  1. Script: write short, clear text (with AI content writing). Short and specific works best.
  2. Figure: pick a stock avatar, or clone your own — with consent and a written agreement.
  3. Voice: choose or clone a voice, set language and tone, load the pronunciation dictionary.
  4. Generation & editing: add captions, background, logo, b-roll and music — then export.

Budget properly for the fourth step. The generated clip is raw material; the assembly is where the video becomes watchable, and it is the same editing work it would have been with filmed footage. What you saved was the shoot, not the edit.

The library problem, which arrives in year two

The reason to use avatars is that videos are cheap to update. The consequence, which nobody plans for, is that you end up with a lot of videos.

A team that films twice a year has twelve videos and knows where they are. A team generating a video per release has hundreds within eighteen months, each with a script, a language variant or three, and a figure whose consent has a scope. At that size the bottleneck stops being production and becomes knowing what you have.

Three habits make the difference, and they cost almost nothing if you start with them:

The failure mode without these is not dramatic. It is a growing pile of videos nobody is sure are current, which is precisely the problem avatars were bought to solve — arrived at by a different route.

What it costs

Pricing is usually per minute of generated video, or a monthly allowance of minutes, with cloning often a separate charge or a higher tier. That shape has two consequences worth planning around.

Revisions consume the budget. You will regenerate for a mispronunciation, a script fix, a pacing problem. Estimate on finished minutes multiplied by three, not on finished minutes.

Long-form is where it gets expensive, and long-form is also where the illusion breaks. Both point the same way: shorter videos, more of them.

Where it genuinely wins

Your first one, in under an hour

The fastest way to find out whether this fits your work is to make one properly rather than to read more comparisons. Pick something real and low-stakes — an internal process explainer, a walkthrough of one feature.

  1. Write ninety seconds of script. Roughly 200 words. Short sentences, one idea per paragraph, and a specific thing the viewer should be able to do afterwards.
  2. Read it aloud and cut whatever you stumble on. This usually removes about a fifth of it, and the fifth it removes is the part that would have sounded synthetic.
  3. Load your proper nouns into the pronunciation dictionary before the first generation, not after you have heard them mangled.
  4. Generate, then listen with your eyes closed. Judge the audio on its own; it carries most of the impression.
  5. Cut away at least twice — to a screen recording, a diagram, anything. Do not leave the figure on screen for ninety unbroken seconds.
  6. Show it to one person who was not involved and ask what they noticed, not whether they liked it. The answer tells you whether the format suits your audience far better than your own judgement will, because you have been staring at it.

If the verdict is "it was clear and I learned the thing", you have found the right use. If it is "it was a bit odd", check whether you asked it to do a trust job rather than an information job — that is usually what the oddness is.

Where to use a real camera instead

Ethics & disclosure

Avatars are a powerful tool and a responsibility. The rules:

Captions are not optional

An avatar speaking is still audio-only information, and it is easy to assume the visual figure makes the video accessible. It does not. Add captions — most platforms generate them from your script, which means they are accurate for once — and make sure any on-screen text is readable at phone size and has enough contrast. If the video teaches something, a text version alongside it serves both accessibility and the substantial number of people who would simply rather read.

Common mistakes

Next step

Combine an avatar with quality voice and video, or explore more creation tools.