Agency OS Skill Library

Captions From A Transcript

Turns one video transcript into six platform-native, SEO-structured captions in your voice, all delivered at once.

Copy First result: about 20 minutes per video social-caption-writer
Back to all skills

What it does

Takes a video transcript and returns six copy-paste ready outputs: TikTok, Instagram, YouTube Shorts, YouTube Long-Form, LinkedIn and Facebook. Each one is written natively for its platform, with the primary keyword front-loaded into the zone that platform actually indexes, hashtag counts that follow that platform's rules, and character counts inside its limits. YouTube gets separate title, description, hashtag and tag fields. Instagram gets alt text and a format recommendation. It writes in your voice by taking the voice signal from the transcript itself first, ahead of any brand file.

Say this to start

This skill has no button. You start it by saying what you want. Any of these will do it:

> write captions
> caption this
> post this video
> social captions
> write a caption for

When to reach for it

When NOT to use it

If you actually wantUse this instead
Writing a social post from scratch with no transcript to work fromlegendary-copywriters
Paid ad copy for Meta, Facebook or Instagrammeta-ads-copywriter
A VSL or sales video scriptvsl-architect
Turning one input into a full multi-asset set of carousels, threads and postscontent-repurposer
Cutting a long video into vertical clipsvideo-repurposer
A long-form SEO articleblog-writer
The 30-day multi-tier posting plan the captions slot intocontent-system-architect

Before you start

What you needWhy
A transcript, as text or a file pathThis is the primary input and the primary voice source. Without it there is nothing to caption and the skill has no fresh voice signal to work fromRequired
brand/[workspace]/verified-claims.md with entries marked consent: ok-as-statedNumbers, results and stories spoken in the video still have to trace to the registry before they appear in a caption. A hook is not exemptRequired
positioning.md with your actual offer names and CTAsOffer alignment pulls real offer names from it. Without it the skill has no offer stack to match the video topic againstOptional
voice-profile.md and any voice guide files in knowledge/They back up the transcript voice. The transcript still takes precedence, so these matter most when the transcript is shortOptional
keyword-plan.mdIt supplies the keywords the captions layer in. Without it the skill derives keywords from the transcript aloneOptional
The video file and ffmpeg installed, if you want visual contextIt extracts a frame every five seconds to see screen recordings, whiteboards and text overlays. Without it the skill works from words only and cannot reference what is shownOptional
Web search availableIt researches current algorithm behaviour and hashtag patterns per platform. If unavailable it notes that research was skipped and falls back to its built-in rulesOptional

How it runs

  1. Read the transcript in fullBefore anything else it reads the whole transcript, and reads extracted frames if you gave it the video. It pulls the core topic, the strongest hook, three to five key moments, stories, data points, the emotional tone and the keywords you repeat naturally.
  2. Detect the CTAIt looks for a real CTA in the transcript: 'comment this word', 'DM me', 'link in bio'. If it finds one it preserves it exactly. If there is none it marks the video value-only and writes no pitch into any caption.
  3. Align to an offer, or do notIt matches the video topic against your offer stack from positioning.md and picks the CTA style that fits. This runs alongside the research step. It will not force an offer where none fits naturally, and the alignment note is for you, not for publishing.
  4. Research the platforms and the keywordsIt searches for current algorithm and hashtag behaviour per platform, then sets one primary keyword and three to five secondary keywords per platform, because the same video is searched differently on TikTok than on YouTube.
  5. Write six captions in parallelIt launches one sub-agent per output. Each gets the content analysis, offer alignment, platform trends, keywords, voice profile and that platform's playbook, and writes independently as native content rather than as a reformatted copy of the others.
  6. Run the quality gateThirteen checks before delivery, including: does every caption sound like you rather than a social media manager, is the primary keyword in the first line on every platform, are hashtag counts within each platform's rules, does LinkedIn land between 1,300 and 1,900 characters, is the transcript CTA preserved exactly, and does a value-only video contain no hidden pitch.
  7. Run the claim checksIt lints the captions against your verified-claims registry and spawns a claim grader over the final text. A failure blocks the caption until the flagged claim is grounded or made generic.
  8. Deliver everything at once and learnAll six outputs arrive in a single response, never one platform at a time. It then asks whether you shipped as-is, made minor tweaks, rewrote significantly or scrapped it, and which platform performed best. That feedback goes to learnings.md, and any pattern seen three times becomes a rule.

What you get

Honest limits

Read this before you rely on it

Where people go wrong

The mistakeDo this instead
Feeding it a cleaned-up summary instead of the raw transcriptGive it the full unedited transcript. The transcript is the primary voice source and outranks your brand voice file, so a polished summary strips out exactly the phrasing and rhythm it needs.
Asking it to add a CTA to a pure value videoLeave it value-only. Shoehorned offers are explicitly against the skill's rules, and the quality gate checks that a value-only video carries no hidden pitch.
Overriding the CTA the video already containsIf you said 'comment AGENCY' on camera, that exact CTA belongs in the caption. The skill preserves it verbatim for a reason: it has to match what the viewer just heard.
Asking for one platform at a timeLet it deliver all six at once. The sub-agents run in parallel, so asking for them one by one costs you time and produces less consistent keyword coverage.
Putting the link in the Facebook post bodyPut it in the first comment. 98 percent of top-performing posts carry no outbound link in the body, and the quality gate checks this.
Treating the offer alignment note as caption copyThat block is for you, not for publishing. It explains why the skill matched the video to a given offer.
Worth knowing

The transcript outranks every brand file for voice, so the rawer the transcript the better the captions. If you were fired up on camera, the captions come back fired up. Send the verbatim transcript including the false starts and the asides, and send the video file too if you have ffmpeg, because the frames let the captions reference what was actually on screen rather than only what was said.