AI Avatars — A Talking Head
Turn text into a video with a talking presenter — no camera, studio or crew. The technology, the tools and the use cases.
What an AI avatar is
An AI avatar is a video of a human figure — real or synthetic — that "speaks" text you wrote, with lip movement, expression and a generated voice. You write a script, choose a figure and a voice, and the system renders a finished video. No camera, lighting, location or crew.
Not that the first video is cheap — that the fiftieth revision is. Change a line of text, regenerate, done. That is a different economics from filming.
What avatars are actually for
The tempting summary is "video without a shoot", and it leads people to use avatars for the wrong things and conclude the technology is not ready. It is ready — for a specific job.
The job is delivering information that changes. A compliance module that gets rewritten every quarter, a product walkthrough that goes stale with each release, onboarding that mentions a tool you no longer use. Filmed video is expensive to change, so in practice it never gets changed, and organisations quietly serve outdated training for years. An avatar video is edited by editing text. That is the whole value proposition and it is a real one.
The job it does badly is anything that has to be believed, felt or trusted. A founder's message about layoffs. A sales pitch that depends on rapport. A thank-you to customers. Here the format works against you: a synthetic person expressing sentiment reads as someone who could not be bothered to say it themselves, and that is exactly the message received regardless of what the words said.
So the test before you generate anything: is this video transferring information, or asking for trust? Information is an excellent fit. Trust is not, and no improvement in rendering quality changes that — the problem is not that it looks fake, it is that it is not you.
Where the illusion breaks
Current avatars are convincing in short bursts and get less convincing the longer they run. Knowing the specific failure points lets you design around them.
- Duration. Fifteen seconds is comfortable, two minutes is noticeably synthetic, ten minutes is uncomfortable. The small stillnesses accumulate.
- Emotional range. Neutral, informative delivery is handled well. Genuine enthusiasm, humour and gravity are where it flattens.
- Fast or emphatic speech. Lip sync holds up at measured pace and drifts when the delivery gets excited.
- The body below the neck. Faces have had the most attention; hands, shoulders and posture are where a loop or a stiffness shows.
- Names and unusual words. The voice mispronounces your product, your company and anything foreign, confidently.
The design response is structural: keep segments short, cut to screen recordings and graphics often, and let the avatar introduce and hand off rather than present continuously. A talking head on screen for the entire runtime is the hardest possible version and the one people try first.
Writing for an avatar is not writing for a page
Most disappointing avatar videos are script problems wearing a technology costume. A few mechanical things change the output more than switching tools would.
- Short sentences. Long subordinate clauses expose the lack of natural breath and phrasing. If you would need a comma to survive saying it aloud, split it.
- Punctuation is your timing control. A full stop gives a pause; a comma gives a shorter one. Paragraph breaks often produce a beat. This is most of the direction you get, so use it deliberately.
- Spell out what the voice will get wrong. Numbers, dates, currencies and acronyms are read by rule and the rule is often not yours. Write "twenty twenty-six" or "A-P-I" if that is what you want to hear. Most platforms also accept a pronunciation dictionary — worth filling in once with your product and company names.
- No stage directions. "(laughs)" or "*emphasis*" gets read aloud or ignored; it does not direct the performance.
- Read it out loud before you generate. Anything you stumble over, the avatar will render badly — and generating costs credits while reading costs nothing.
Voice is most of the impression
People say "the avatar looks fake" when what they actually noticed was the voice. Audio carries more of the credibility than the face does, and it is the cheaper half to get right.
Match the voice to the figure — an accent or an apparent age that contradicts the face creates a mismatch viewers feel without identifying. Listen to the generated audio on its own, without the video, before you approve it; problems that hide behind a moving face are obvious with your eyes shut. And if you clone your own voice with ElevenLabs or similar, record the source samples properly: a quiet room and a decent microphone for a few minutes beats an hour of noisy audio, because the clone reproduces the room as faithfully as the voice.
Cloning yourself, and the part nobody plans for
Most platforms will build an avatar from a few minutes of footage of a real person. For a company this is attractive: your actual trainer, your actual founder, in videos you can update forever.
Forever is the problem.
A cloned likeness is an asset that outlives the relationship. The employee who recorded it leaves, and their face is still delivering your onboarding. They change roles, or fall out with the company, or simply stop wanting their face used — and the clone is in fifty modules. This is not hypothetical; it is the predictable end state of the useful version of this feature.
Handle it before you record, in writing, covering four things: what the likeness may be used for (internal training only, or public marketing too), for how long, what happens when the person leaves, and how they withdraw consent and what you undertake to do then. Give the person a copy. If your organisation has an HR or legal function, this is a conversation for them and not a checkbox in a tool.
Add an operational habit: keep a register of every video that uses a cloned person, so that honouring a withdrawal is a list to work through rather than an archaeology project.
Translation and lip sync
Translating a video into several languages with matched lip movement is the most impressive thing these tools do, and it is genuinely useful for training material that has to reach several markets.
Three practical cautions. Length changes. The same sentence takes noticeably longer in some languages than others, which pushes your timing out of step with anything on screen — so leave slack, and avoid tight synchronisation between narration and visuals. Translation is not localisation. Examples, currencies, regulations, names and humour need adapting, not converting, and a machine translation will carry all of them across unchanged. Have a native speaker watch it before it ships. Not read the transcript — watch it, because the errors that matter are tone and register, and those do not show up in text.
Worth being clear-eyed about the last one: a translated avatar video that no fluent speaker has reviewed is a public statement in a language you cannot check. The cost of that review is small and the alternative is discovering the problem from the audience.
Leading tools
- HeyGen — realistic avatars, cloning your own figure, and translating videos into many languages with lip sync.
- Synthesia — geared to enterprise training, with a library of avatars and templates and multilingual support.
- D-ID — animating a static photo into a talking figure, and real-time avatar streaming.
They differentiate less on rendering quality — which converges quickly — than on the workflow around it: template and brand controls, how review and approval work, how many seats, and what the licence says about commercial use and about cloned likenesses. If this is for a company rather than a personal channel, read the likeness and data terms before comparing avatar galleries.
Workflow
- Script: write short, clear text (with AI content writing). Short and specific works best.
- Figure: pick a stock avatar, or clone your own — with consent and a written agreement.
- Voice: choose or clone a voice, set language and tone, load the pronunciation dictionary.
- Generation & editing: add captions, background, logo, b-roll and music — then export.
Budget properly for the fourth step. The generated clip is raw material; the assembly is where the video becomes watchable, and it is the same editing work it would have been with filmed footage. What you saved was the shoot, not the edit.
The library problem, which arrives in year two
The reason to use avatars is that videos are cheap to update. The consequence, which nobody plans for, is that you end up with a lot of videos.
A team that films twice a year has twelve videos and knows where they are. A team generating a video per release has hundreds within eighteen months, each with a script, a language variant or three, and a figure whose consent has a scope. At that size the bottleneck stops being production and becomes knowing what you have.
Three habits make the difference, and they cost almost nothing if you start with them:
- Keep the scripts outside the tool, in whatever your team already uses for documents. Scripts are the source; the video is the build output. Editing the script and regenerating should be the normal path, and that only works if the script is findable by someone who is not you.
- Put a version and a date in the video itself, on screen or in the description. When someone reports that a module is wrong, the first question is which version they watched.
- Keep a register of what each video covers, which avatar and voice it uses, and what it is embedded in. This is the same list that makes a consent withdrawal manageable, so it earns its keep twice.
The failure mode without these is not dramatic. It is a growing pile of videos nobody is sure are current, which is precisely the problem avatars were bought to solve — arrived at by a different route.
What it costs
Pricing is usually per minute of generated video, or a monthly allowance of minutes, with cloning often a separate charge or a higher tier. That shape has two consequences worth planning around.
Revisions consume the budget. You will regenerate for a mispronunciation, a script fix, a pacing problem. Estimate on finished minutes multiplied by three, not on finished minutes.
Long-form is where it gets expensive, and long-form is also where the illusion breaks. Both point the same way: shorter videos, more of them.
Where it genuinely wins
- Training and onboarding — the core case, because the content changes and filmed video never gets re-filmed.
- Product and release notes — a short video per release, which nobody would ever book a studio for.
- Internal communication at scale — policy updates, process changes, the same message in several languages.
- Personalisation — the same structure with a changing name or account detail, at a volume no person could record.
- Appearing without being on camera — for consistency, for privacy, or because you simply do not want to be.
Your first one, in under an hour
The fastest way to find out whether this fits your work is to make one properly rather than to read more comparisons. Pick something real and low-stakes — an internal process explainer, a walkthrough of one feature.
- Write ninety seconds of script. Roughly 200 words. Short sentences, one idea per paragraph, and a specific thing the viewer should be able to do afterwards.
- Read it aloud and cut whatever you stumble on. This usually removes about a fifth of it, and the fifth it removes is the part that would have sounded synthetic.
- Load your proper nouns into the pronunciation dictionary before the first generation, not after you have heard them mangled.
- Generate, then listen with your eyes closed. Judge the audio on its own; it carries most of the impression.
- Cut away at least twice — to a screen recording, a diagram, anything. Do not leave the figure on screen for ninety unbroken seconds.
- Show it to one person who was not involved and ask what they noticed, not whether they liked it. The answer tells you whether the format suits your audience far better than your own judgement will, because you have been staring at it.
If the verdict is "it was clear and I learned the thing", you have found the right use. If it is "it was a bit odd", check whether you asked it to do a trust job rather than an information job — that is usually what the oddness is.
Where to use a real camera instead
- Anything emotional or sensitive. Bad news, apologies, thanks, anything about people's jobs.
- Building a personal brand. The audience is there for a person; a synthetic stand-in is the opposite of the product.
- High-stakes sales. If the deal turns on trust, spend the twenty minutes filming on a phone. An imperfect real video outperforms a polished synthetic one here.
- When a phone would genuinely be faster. For one ninety-second video that will never be updated, recording it is quicker than scripting, generating and editing.
Ethics & disclosure
Avatars are a powerful tool and a responsibility. The rules:
- Clone only with informed, written consent, covering scope, duration and withdrawal. A verbal "sure, go ahead" is not enough for something this durable.
- Never clone a public figure or anyone who has not agreed, for any purpose, including parody you consider obvious.
- Disclose where a viewer could reasonably think it is a real person speaking to them. The major platforms have been adding requirements around labelling realistic synthetic people; check the current policy where you publish.
- No deception. Do not put words in a figure's mouth that create a false impression of who said them or of what happened.
- Be careful what you feed it. Scripts for internal training often contain unreleased or confidential material; check whether the vendor retains or trains on your inputs before pasting it in.
Captions are not optional
An avatar speaking is still audio-only information, and it is easy to assume the visual figure makes the video accessible. It does not. Add captions — most platforms generate them from your script, which means they are accurate for once — and make sure any on-screen text is readable at phone size and has enough contrast. If the video teaches something, a text version alongside it serves both accessibility and the substantial number of people who would simply rather read.
Common mistakes
- Long single takes. The uncanny effect is a function of duration. Keep segments short and cut away often.
- Neglecting the script. The rendering is good; if the content is weak, the video is weak, and now it is weak in a format nobody can skim.
- Skipping the pronunciation pass. Your product name said wrong, fifty times, across a training library.
- Cloning without paperwork. The awkward conversation happens later, when the person has left and the videos have not.
- Using it where trust is the point. The efficiency is real; spending your credibility to get it is not a trade worth making.
Next step
Combine an avatar with quality voice and video, or explore more creation tools.