Turns a brief or a reference image into art-directed image and video prompts that pass a scroll-stop quality gate.
Produces the creative direction behind a visual: a structured JSON prompt built from a reference image you upload or from a written brief, a plain-text director's brief for quick exploration, or a video scene brief for a single generated clip. Every prompt runs through the Visual Elevation Protocol, a five-principle gate whose job is to stop you shipping the stock-photo version of your idea, and a kill-switch question that fails any visual fully describable as a person doing an activity at a location. It also carries hard rules on fabricated data, text rendering and platform safe zones, so what comes back is usable rather than just pretty.
This skill has no button. You start it by saying what you want. Any of these will do it:
| If you actually want | Use this instead |
|---|---|
| shot-by-shot video sequences with effects timelines and B-roll planning | video-prompt-builder |
| AI-avatar talking-head video production | heygen |
| the headline or body copy that sits on an ad visual | meta-ads-copywriter |
| the written copy around a visual outside of ads | legendary-copywriters |
| designing an actual HTML page or funnel section | frontend-design |
| cutting, captioning or clipping footage that already exists | video-repurposer |
| What you need | Why | |
|---|---|---|
| A brief, or a reference image if you want it reverse-engineered | the decision tree routes on which of the two you have, and the output format differs | Required |
The workspace tool-stack file at protocol/current-tool-stack.md | it names the image and video tools of record, which decides whether you get a structured JSON prompt or a single flat prompt string | Required |
voice-profile.md and positioning.md in the active workspace | colour preference, energy and aesthetic come from voice; market tier from positioning changes the visual treatment substantially | Required |
epistemology.md in the active workspace | it holds what the brand rejects, which is what keeps the visuals off the category's default look. Without it you get a one-line warning and visuals built from aesthetic cues alone, which drift towards cliché | Optional |
policy.md in the active workspace | any never-touch topic is excluded from prompts and references | Optional |
| A media generation tool connected to your workspace | with one connected you get actual generated images and video. Without one you still get complete briefs and prompts to paste into whatever you use | Optional |
| The exact real data for anything that must appear on the visual: phone, email, URL, address, price | it will never invent these. Anything you have not supplied comes back as a placeholder plus a production note telling you to add it yourself | Optional |
The client's own brand files when you are working in a client-- workspace | the client's identity applies there, not your agency's | Optional |
brand/[workspace]/assets.md, plus a logged learning on how the visuals landedepistemology.md in the workspace the visuals default to the category's consensus look. You get a warning, not a fix.| The mistake | Do this instead |
|---|---|
| Accepting the first prompt because it reads professionally | Run the kill switch on it yourself. If you can describe it fully as a person doing an activity at a location, it will generate a stock photo no matter how well written it is. |
| Expecting it to fill in the business phone number, price or URL | Supply the exact data in the session, or accept the placeholder and add the text yourself in post. This is a zero-tolerance rule, not a limitation to argue with. |
| Loading long rendered stats or numbers into the image itself | Keep rendered text to 2 to 5 words and simple round numbers. Put anything that must be letter-perfect on as a post-production overlay. |
| Skipping safe zones on vertical or feed output | Load the safe zones for the target platform before writing the prompt. UI overlays will otherwise sit on top of your subject or your text. |
| Asking for both the JSON prompt and the five-part brief | Pick one. The JSON schema replaces the director's brief format; producing both means two versions of the same direction drifting apart. |
Faces are worth keeping in frame, they raise click-through, but a face alone is exactly what fails the kill switch. Pair the face with one elevated element, an unexpected scale, a juxtaposition, or something that should not physically be possible, and you keep the recognition benefit without the stock-photo penalty.