Back to all skills
What it does
Takes a creative brief, anything from one line to a full storyboard, and returns a complete five-section video prompt. You get a shot-by-shot effects timeline, a master inventory of every effect used, a density map showing where the video is busy and where it breathes, a three-act energy arc, and a paste-ready prompt block with one self-contained prompt per clip plus a STYLE LOCK footer that keeps every clip visually consistent. The prompts are model-agnostic architecture: shot structure, camera language, lighting and constraints transfer to whichever video tool your workspace has on record.
Say this to start
This skill has no button. You start it by saying what you want. Any of these will do it:
/build-video-prompt
> build me a video prompt
> shot list for this video
> b-roll planning
> ad concept shot by shot
> brand film shot list
When to reach for it
- planning a cinematic brand film shot by shot
- a social ad under 10 seconds that needs a signature moment
- a product or SaaS demo sequence
- B-roll planning before any generation
- an ad concept that needs an effects breakdown
When NOT to use it
| If you actually want | Use this instead |
| a still image prompt or the overall visual strategy | creative-director |
| the spoken script of a sales video | vsl-architect |
| editing, cutting or grading footage you already have | auto-editor |
| slicing an existing long video into short clips | video-repurposer |
Before you start
| What you need | Why | |
| A creative brief, even a rough one | if the brief is too vague to build from, it will ask one focused clarifying question before proceeding, which costs you a round trip | Required |
The workspace tool-stack file at protocol/current-tool-stack.md | it names the video tool the prompts are targeted at, and the paste-ready section is formatted for it | Required |
voice-profile.md in the active workspace | supplies brand personality, colour preferences, energy level and aesthetic direction; without it the visual tone is generic | Optional |
positioning.md in the active workspace | market tier, luxury through to accessible, shapes visual treatment, pacing and effects density | Optional |
offer-output.md in the active workspace | gives product context and audience, which shape the shot list's narrative and CTA framing | Optional |
| An account, credits and API access for the video tool on record | only needed if you want it to chain straight into generation. The default path just hands you the prompts | Optional |
| A duration target | without one it defaults to 10 to 15 seconds, which may not be what your placement needs | Optional |
How it runs
- You give it the briefAnything from 'a runner in a stadium, 10 seconds' to a detailed storyboard. Useful additions are subject, setting, mood, brand context, specific camera moves, duration target, colour palette and any reference films or ads.
- It loads the rules and the tool of recordIt reads the prompting reference that holds the model constraints, checks the workspace tool-stack file for which video tool to target, and pulls the breakdown reference closest to your format: athletic brand film, product demo, short social ad or longer brand story.
- It picks up workspace contextBrand personality and colour cues from your voice profile, market tier from positioning, product context from your offer, and any recent feedback logged against this skill or creative-director. It mentions the adaptations it made.
- It calibrates shot count to durationDuration drives everything: 4 to 6 seconds gets 3 to 5 shots and one density peak, 20 to 30 seconds gets 14 to 20 shots and a full three-act arc. If you did not name a duration it uses 10 to 15 seconds.
- It writes the four analysis sectionsThe shot-by-shot timeline with effect, visual description, camera, timing, lighting, constraints and transition per shot; the master effects inventory with usage counts; the density map rating each 3 to 6 second segment high, medium or low; and the energy arc across acts. The most impactful shot is explicitly flagged as the signature visual effect.
- It writes the paste-ready blockEach shot collapses into a standalone 50 to 100 word prompt using the six-element formula: subject, action, environment, camera, style, constraints. A STYLE LOCK footer holds the style, lighting and constraint keywords you paste into every single generation so the clips match.
- You pick a pathThe default is prompt only: take the paste-ready block to whichever tool you use, no API keys needed. If you want end-to-end generation and have the platform's tools and key configured, it hands the prompts to the execution skill for the tool on record.
- Feedback loopIt asks whether the prompt is ready as-is, needs minor tweaks, or needs rework, and logs the answer with the video type and duration so the next prompt adapts.
What you get
- Section 1: a shot-by-shot effects timeline, each shot with effect, visual, camera, speed, lighting, constraints and transition
- Section 2: a master effects inventory listing every distinct effect, how often it appears and where
- Section 3: an effects density map rating each segment high, medium or low
- Section 4: an energy arc describing how the video opens, develops and resolves
- Section 5: paste-ready prompt blocks, one per clip, plus a reusable STYLE LOCK footer
- An entry appended to
brand/[workspace]/assets.md with shot count, duration, effects count, video type and signature effect
Honest limits
Read this before you rely on it
- It writes prompts. It does not generate video. On the default path you take the output somewhere else and generate it yourself.
- Current models generate roughly 4 to 15 seconds per clip, so anything longer is planned as separate generations with marked break points and assembled afterwards in an editor.
- Chaining straight into generation needs the platform's tools configured and its API key set. Without both, you stay on the prompt-only path.
- The prompt architecture transfers between models but the paste-ready wording may need light adaptation if you target a different platform to the one your tool-stack file names.
- Any specific claim in an on-screen line has to trace to your workspace's verified claims. Without a registry entry it stays generic and gets marked as a proof gap rather than invented.
Where people go wrong
| The mistake | Do this instead |
| Writing one sentence that moves both the camera and the subject | Split them. 'The dancer spins slowly in the centre of the stage. Camera holds fixed framing.' This is the most common failure mode across every current video model. |
| Stacking simultaneous camera movements to get a richer shot | One primary camera instruction per shot. Sequential compounds are fine, 'push-in then subtle rise'; simultaneous ones cause jitter. |
| Writing longer prompts to get more control | Stay at 50 to 100 words per shot. Longer prompts accumulate conflicting instructions and degrade quality. Be specific, not verbose. |
| Asking for fast motion throughout | Use fast sparingly, one fast element per shot at most. Unqualified 'fast' degrades output. Slow and medium have their own vocabulary that models read reliably. |
| Generating each clip from its own prompt and wondering why they do not match | Paste the STYLE LOCK block into every single generation. It exists to hold style, lighting and constraints identical across clips. |
Worth knowingIf you can only add one detail to a shot, add lighting. It is the single highest-impact element in a video prompt, and every shot in the timeline gets a lighting line for that reason.