Agency OS Skill Library

Creative Director

Turns a brief or a reference image into art-directed image and video prompts that pass a scroll-stop quality gate.

Video First result: 30 minutes creative-director
Back to all skills

What it does

Produces the creative direction behind a visual: a structured JSON prompt built from a reference image you upload or from a written brief, a plain-text director's brief for quick exploration, or a video scene brief for a single generated clip. Every prompt runs through the Visual Elevation Protocol, a five-principle gate whose job is to stop you shipping the stock-photo version of your idea, and a kill-switch question that fails any visual fully describable as a person doing an activity at a location. It also carries hard rules on fabricated data, text rendering and platform safe zones, so what comes back is usable rather than just pretty.

Say this to start

This skill has no button. You start it by saying what you want. Any of these will do it:

/create-visuals
> build a JSON prompt from this
> recreate this image
> prompt from reference
> give me a JSON prompt
> ad variants as JSON

When to reach for it

When NOT to use it

If you actually wantUse this instead
shot-by-shot video sequences with effects timelines and B-roll planningvideo-prompt-builder
AI-avatar talking-head video productionheygen
the headline or body copy that sits on an ad visualmeta-ads-copywriter
the written copy around a visual outside of adslegendary-copywriters
designing an actual HTML page or funnel sectionfrontend-design
cutting, captioning or clipping footage that already existsvideo-repurposer

Before you start

What you needWhy
A brief, or a reference image if you want it reverse-engineeredthe decision tree routes on which of the two you have, and the output format differsRequired
The workspace tool-stack file at protocol/current-tool-stack.mdit names the image and video tools of record, which decides whether you get a structured JSON prompt or a single flat prompt stringRequired
voice-profile.md and positioning.md in the active workspacecolour preference, energy and aesthetic come from voice; market tier from positioning changes the visual treatment substantiallyRequired
epistemology.md in the active workspaceit holds what the brand rejects, which is what keeps the visuals off the category's default look. Without it you get a one-line warning and visuals built from aesthetic cues alone, which drift towards clichéOptional
policy.md in the active workspaceany never-touch topic is excluded from prompts and referencesOptional
A media generation tool connected to your workspacewith one connected you get actual generated images and video. Without one you still get complete briefs and prompts to paste into whatever you useOptional
The exact real data for anything that must appear on the visual: phone, email, URL, address, priceit will never invent these. Anything you have not supplied comes back as a placeholder plus a production note telling you to add it yourselfOptional
The client's own brand files when you are working in a client-- workspacethe client's identity applies there, not your agency'sOptional

How it runs

  1. It loads the workspace and the tools of recordVoice profile for personality and colour, positioning for market tier, epistemology for what the brand rejects, policy for excluded topics, past feedback for your visual preferences. It reads the tool-stack file to know which image and video tools it is directing for, and mentions any adaptation it made from your logged preferences.
  2. It answers three questions before writing anythingWhat emotion should the viewer feel, what action should they take, and what makes this memorable. These drive the colour temperature, lighting and composition choices that follow.
  3. It routes to a formatA reference image you want recreated, or a brief for ad variants and structured single-image prompts, goes to the JSON schema. A quick in-conversation exploration goes to the plain-text director's brief. A video request goes to the video scene brief. It picks one and does not output both.
  4. Visual Elevation ProtocolAt least one of five principles is applied to every prompt, two or more for thumbnails, hero images and ad creatives: metaphor over literal, visual tension through juxtaposition, scale disruption, emotional extremes, and one element that breaks the rules. Text-on-solid designs, tutorial screens, literal product demos and simple data graphics are exempt.
  5. The kill switchBefore finalising, it asks whether the visual can be fully described as a person doing an activity at a location with nothing memorable left over. If yes, the prompt is reworked. A face in frame is fine; a face and nothing else is not.
  6. Safe zones and text rulesAny social-platform or vertical output gets safe zones applied so platform UI does not cover the subject or the text. Rendered text is kept to 2 to 5 words, multi-digit numbers are avoided because generators garble them, and anything that must be letter-perfect is marked as a post-production overlay.
  7. Anti-patterns passIt checks the output against the six tells of synthetic imagery: dead-centre symmetry, plastic overlit skin, stock-photo energy, floating or illegible text, uniformly oversaturated colour, and melted background detail.
  8. Output and feedbackJSON mode returns analysis, the JSON prompt, and tweaks. Director's brief mode returns the prompt, the creative rationale, technical specs, two to three variations to test, and production notes. Generated assets are logged to your workspace assets file and it asks how the visuals landed.

What you get

Honest limits

Read this before you rely on it

Where people go wrong

The mistakeDo this instead
Accepting the first prompt because it reads professionallyRun the kill switch on it yourself. If you can describe it fully as a person doing an activity at a location, it will generate a stock photo no matter how well written it is.
Expecting it to fill in the business phone number, price or URLSupply the exact data in the session, or accept the placeholder and add the text yourself in post. This is a zero-tolerance rule, not a limitation to argue with.
Loading long rendered stats or numbers into the image itselfKeep rendered text to 2 to 5 words and simple round numbers. Put anything that must be letter-perfect on as a post-production overlay.
Skipping safe zones on vertical or feed outputLoad the safe zones for the target platform before writing the prompt. UI overlays will otherwise sit on top of your subject or your text.
Asking for both the JSON prompt and the five-part briefPick one. The JSON schema replaces the director's brief format; producing both means two versions of the same direction drifting apart.
Worth knowing

Faces are worth keeping in frame, they raise click-through, but a face alone is exactly what fails the kill switch. Pair the face with one elevated element, an unexpected scale, a juxtaposition, or something that should not physically be possible, and you keep the recognition benefit without the stock-photo penalty.