Turns one video transcript into six platform-native, SEO-structured captions in your voice, all delivered at once.
Takes a video transcript and returns six copy-paste ready outputs: TikTok, Instagram, YouTube Shorts, YouTube Long-Form, LinkedIn and Facebook. Each one is written natively for its platform, with the primary keyword front-loaded into the zone that platform actually indexes, hashtag counts that follow that platform's rules, and character counts inside its limits. YouTube gets separate title, description, hashtag and tag fields. Instagram gets alt text and a format recommendation. It writes in your voice by taking the voice signal from the transcript itself first, ahead of any brand file.
This skill has no button. You start it by saying what you want. Any of these will do it:
| If you actually want | Use this instead |
|---|---|
| Writing a social post from scratch with no transcript to work from | legendary-copywriters |
| Paid ad copy for Meta, Facebook or Instagram | meta-ads-copywriter |
| A VSL or sales video script | vsl-architect |
| Turning one input into a full multi-asset set of carousels, threads and posts | content-repurposer |
| Cutting a long video into vertical clips | video-repurposer |
| A long-form SEO article | blog-writer |
| The 30-day multi-tier posting plan the captions slot into | content-system-architect |
| What you need | Why | |
|---|---|---|
| A transcript, as text or a file path | This is the primary input and the primary voice source. Without it there is nothing to caption and the skill has no fresh voice signal to work from | Required |
brand/[workspace]/verified-claims.md with entries marked consent: ok-as-stated | Numbers, results and stories spoken in the video still have to trace to the registry before they appear in a caption. A hook is not exempt | Required |
positioning.md with your actual offer names and CTAs | Offer alignment pulls real offer names from it. Without it the skill has no offer stack to match the video topic against | Optional |
voice-profile.md and any voice guide files in knowledge/ | They back up the transcript voice. The transcript still takes precedence, so these matter most when the transcript is short | Optional |
keyword-plan.md | It supplies the keywords the captions layer in. Without it the skill derives keywords from the transcript alone | Optional |
| The video file and ffmpeg installed, if you want visual context | It extracts a frame every five seconds to see screen recordings, whiteboards and text overlays. Without it the skill works from words only and cannot reference what is shown | Optional |
| Web search available | It researches current algorithm behaviour and hashtag patterns per platform. If unavailable it notes that research was skipped and falls back to its built-in rules | Optional |
positioning.md and picks the CTA style that fits. This runs alongside the research step. It will not force an offer where none fits naturally, and the alignment note is for you, not for publishing.learnings.md, and any pattern seen three times becomes a rule.captions-v1 schema artifact other skills can consume, and a log entry in brand/assets.mdverified-claims.md with consent, the caption comes back generic with a proof gap flag rather than repeating it. The skill states plainly that a hook is not exempt, so expect your best spoken line to be the one that gets flagged.| The mistake | Do this instead |
|---|---|
| Feeding it a cleaned-up summary instead of the raw transcript | Give it the full unedited transcript. The transcript is the primary voice source and outranks your brand voice file, so a polished summary strips out exactly the phrasing and rhythm it needs. |
| Asking it to add a CTA to a pure value video | Leave it value-only. Shoehorned offers are explicitly against the skill's rules, and the quality gate checks that a value-only video carries no hidden pitch. |
| Overriding the CTA the video already contains | If you said 'comment AGENCY' on camera, that exact CTA belongs in the caption. The skill preserves it verbatim for a reason: it has to match what the viewer just heard. |
| Asking for one platform at a time | Let it deliver all six at once. The sub-agents run in parallel, so asking for them one by one costs you time and produces less consistent keyword coverage. |
| Putting the link in the Facebook post body | Put it in the first comment. 98 percent of top-performing posts carry no outbound link in the body, and the quality gate checks this. |
| Treating the offer alignment note as caption copy | That block is for you, not for publishing. It explains why the skill matched the video to a given offer. |
The transcript outranks every brand file for voice, so the rawer the transcript the better the captions. If you were fired up on camera, the captions come back fired up. Send the verbatim transcript including the false starts and the asides, and send the video file too if you have ffmpeg, because the frames let the captions reference what was actually on screen rather than only what was said.