Produces AI avatar videos in HeyGen, from talking heads and explainers to short cinematic clips and multi-scene productions.
Creates avatar-led videos in HeyGen across four modes: a talking head of 60 seconds or more for explainers, training and VSLs; short cinematic clips of 5 to 15 seconds with body movement and camera work; a full production mixing dialogue scenes with cinematic hooks and closers; and B-roll only with no avatar. It delivers either a finished video and download link generated through the API, or a fully written prompt plus exact dashboard settings you paste in yourself. Every mode starts from an optimised prompt, because output quality tracks prompt quality almost entirely.
This skill has no button. You start it by saying what you want. Any of these will do it:
| If you actually want | Use this instead |
|---|---|
| writing the spoken script when the ask is a sales video | vsl-architect |
| writing the persuasive copy itself | legendary-copywriters |
| a multi-shot shot list with effects timelines and density maps | video-prompt-builder |
| still image prompts or overall art direction | creative-director |
| editing raw camera or screen footage you already have | auto-editor |
| slicing an existing long video into clips and posts | video-repurposer |
| What you need | Why | |
|---|---|---|
| A HeyGen account at app.heygen.com | needed for every mode, including the one where you paste the prompt into the dashboard yourself | Required |
A HeyGen API key set as the HEYGEN_API_KEY environment variable | required for the two API modes, generate-for-me and precise scene control. Without it you can still use the dashboard-prompt mode | Optional |
| HeyGen credits on the account | generation spends real credits. Cinematic clips run at 4 premium credits per second, so a 15 second clip is 60 | Required |
A check of protocol/current-tool-stack.md | if the workspace names a different avatar-video tool as the tool of record, the request should be routed there instead. The tool-stack file wins that decision | Required |
| Your topic or script, plus who the video is for | the skill needs the video type, duration and audience before it can write an optimised prompt | Required |
| Workspace brand files: owner profile, voice profile, ICA and verified claims | they set the styling, the script tone, the audience angle, and the grounding for anything claimed on screen or out loud | Optional |
| HeyGen MCP tools configured | preferred over raw HTTP calls when present. Without them the skill falls back to direct API calls with your key | Optional |
No toolchain to install. This skill needs one thing: your own HeyGen API key, for the modes that generate through the API.
From your HeyGen account settings. Without it the skill still produces the script and the prompt for you to run manually, so it degrades rather than blocks.
Should print your key. If it prints nothing, open a new terminal.
brand/[workspace]/assets.md recording the generated video| The mistake | Do this instead |
|---|---|
| Passing a talking photo ID as the avatar ID in an API call | The API does not accept them. Switch to the dashboard-prompt mode and select the talking photo there. |
| Sending a rough one-line prompt and judging the tool by the result | Never send an unoptimised prompt. Quality depends almost entirely on prompt quality, which is why the skill optimises before every generation in every mode. |
| Writing cinematic clip prompts with negations, such as 'no shaky camera' | Positive prompting only. Describe what you want, and close with smooth motion and stable framing. |
| Stacking camera movements to make a clip feel dynamic | One camera movement per shot, never stacked. Dynamism comes from the subject's action and the shot choice. |
| Trying to script one long production as a single generation | Split it. Dialogue scenes on the talking-head path, hooks and closers on the cinematic path, then edit together. |
When you need a multi-shot video rather than a single scene, run video-prompt-builder first. It produces the structured shot list with effects breakdown and paste-ready blocks that feed directly into this skill's cinematic clip mode, so you are not inventing camera language shot by shot inside the generation step.