Choose an Avatar
Text mode sends the script and voice to Avatar V. Cost depends on the estimated output-video duration.
Choose a Voice
Fill frame fills the whole video and may crop edges. Show full avatar keeps the full avatar visible and may add empty space.
WebM output is transparent. For vertical 9:16 videos, use MP4.
Works only with matting-enabled Avatar V avatars.
Confirm you have the rights and consent to use this identity, voice/audio, and avatar for a digital-twin/spokesperson video.
Sign in to spend tokens
Choose an Avatar is required
Your generated video will appear here
Avatar V / HeyGen Digital Twin creates a talking avatar video from text or audio using HeyGen's studio-grade digital-twin avatars — with output up to 4K, transparent WebM, and a library of 1,200+ avatars. On Prizmad it's built in: pick an avatar, add a script or recording, generate.
Avatar V is HeyGen's digital-twin avatar model: instead of animating a flat photo, it renders a pre-built, studio-quality avatar identity speaking your script with natural lip-sync. On Prizmad you get a searchable library of over 1,200 Avatar V-compatible avatars and 100+ voices, so you can cast a spokesperson without filming anyone.
It works in two modes. Generate from Text sends your script and chosen voice straight to Avatar V — the cost depends on the estimated length of the output video. Upload Audio takes a clean voice recording of up to 60 seconds and lip-syncs the avatar to it, which is the way to go when you already have a voiceover, a cloned voice, or a specific read you want to keep. Output options are unusually deep for an avatar tool: 720p, 1080p, or 4K resolution, 16:9 or 9:16 aspect ratio, MP4 or transparent WebM, plus a framing control and background removal on matting-enabled avatars.
On Prizmad, Avatar V uses the same token balance as AI image generation, video, voiceover, captions, and music. There is no separate HeyGen account or per-seat billing — you generate from the same dashboard as the rest of your creatives, and every video is yours to use commercially.
A polished presenter delivering your offer in 16:9 or 9:16 — studio quality without booking talent or renting a set.
Transparent WebM output lets you composite a talking presenter over product demos, screen recordings, or branded backgrounds.
Keep the same avatar and swap the script and voice to spin one winning ad into multiple languages and markets.
Walk viewers through your product, offer, or unboxing story with a consistent digital presenter across every touchpoint.
4K output covers YouTube, CTV-style placements, and website heroes where a 720p avatar clip would look soft.
A searchable library of over 1,200 Avatar V-compatible avatars with preview thumbnails — cast a presenter that matches your brand and audience.
Type a script and pick from 100+ voices, or upload a clean voice recording up to 60 seconds for exact lip-sync to your own audio.
Render at 720p, 1080p, or 4K — enough headroom for large-format placements, not just mobile feeds.
Export as MP4 or as WebM with a transparent background, so you can layer the presenter over your own footage or designs. For vertical 9:16 videos, use MP4.
Strip the background entirely on matting-enabled Avatar V avatars — drop the presenter straight onto product shots or branded scenes.
Fill frame crops to fill the whole video; Show full avatar keeps the entire avatar visible — pick per placement instead of cropping in an editor.
Open the avatar picker and browse or search the 1,200+ avatar library. Preview thumbnails help you cast the right presenter for your niche.
Choose your audio source: Generate from Text or Upload Audio.
In text mode, write the script and pick a voice — the voice picker covers 100+ voices, most with playable previews. In audio mode, upload a clean voice recording up to 60 seconds.
Set framing (Fill frame or Show full avatar), output format (MP4, or WebM for a transparent background), resolution (720p, 1080p, or 4K), and aspect ratio (16:9 or 9:16).
Optionally toggle Remove background if your chosen avatar supports matting.
Confirm rights and consent for the identity, voice, and avatar you're using, then click Generate. The video renders in roughly 3–8 minutes and appears in your asset library.
Avatar V generations start at 8 tokens. In text mode the cost depends on the estimated duration of the output video, so shorter scripts cost less. Tokens come from your token balance.
Avatar V is unlocked on the same token balance you use for image generation, video, voiceover, and music — no separate HeyGen account, no API key, no per-seat billing.
If your tokens run out mid-campaign, buy a one-off top-up directly from the top-up page. Every video you generate stays yours with full commercial rights.
It runs on your Prizmad token balance. There is no separate charge for the tool — generations start at 8 tokens, and in text mode the cost scales with the estimated output-video duration.
From 8 tokens per generation. Text mode pricing depends on the estimated duration of the output video, so a short hook costs less than a full 60-second read. You see the cost in the tool before you generate.
The picker covers over 1,200 Avatar V-compatible avatars with searchable previews, and 100+ voices — most with playable audio previews so you can hear the read before committing tokens.
Yes. Switch the audio source to Upload Audio and attach a clean voice recording up to 60 seconds — MP3, WAV, AAC, OGG, and several other formats are accepted. The avatar lip-syncs to your recording exactly.
Two ways. WebM output format renders with transparency, so you can layer the avatar in an editor. Separately, the Remove background toggle strips the background at generation time — it works only with matting-enabled Avatar V avatars. Note that for vertical 9:16 videos, MP4 is the recommended format.
Yes — with one requirement: the tool asks you to confirm you have the rights and consent to use the chosen identity, voice or audio, and avatar for a digital-twin or spokesperson video. Once confirmed, the generated videos are yours to use in paid ads, landing pages, and organic content.
Avatar V uses HeyGen's pre-built digital-twin avatar library (1,200+ identities) and outputs up to 4K with transparent WebM and background removal — it starts at 8 tokens. HeyGen Avatar 4 animates any avatar image you upload, tops out at 1080p, offers more aspect ratios, and bills by duration from 1 token per 5 seconds. Use Avatar V for maximum polish and format flexibility; use Avatar 4 for custom faces and cheaper short clips.
Roughly 3–8 minutes depending on length and resolution. Jobs run asynchronously — queue the video and collect it from your asset library when it's ready.
Published 2026-07-16 · Last updated 2026-07-16