Take Yourself Out of the Caption Review Bottleneck
Caption QA runs every caption in a batch past three ruthless reviewers — one for your voice, one for structure, one for search — then hands you a clear verdict on each: ship it, flag it, or send it back. You only touch the ones that need you.
← Back to the Skills LibraryInstalled in under a minute. No tech skills required.
You don't need to know how to code, and you don't need to set anything up. There are two blocks below. If you can copy and paste, you can install this.
Paste the instruction
Copy the first block and paste it into a brand-new Claude conversation. Don't hit send yet — it's the note that tells Claude what to do.
Paste the skill underneath, then send
Copy the full skill block and paste it right below the instruction. Send it. Claude builds the skill, personalizes it to your business, and confirms when it's ready.
Install the skill below for me. Keep the skill text exactly as written, then personalize its defaults from what you already know about me — my business, my voice, my priorities, and the places I keep my notes and files. Confirm when it's installed and show me exactly how to run it.
---
name: caption-qa-agent
description: Run adversarial multi-lane QA on any batch of short-form video captions before they ship. Three specialized reviewers each attack from a different angle, then a final verification pass prioritizes findings and delivers a clear ship/flag/rewrite verdict for each caption. Use this skill whenever you say "QA these captions," "review captions," "run caption QA," "check these before they ship," "caption review," or drop a batch of captions and want them pressure-tested before publishing. Also trigger when your team delivers captions and you want to know which ones are ready to ship without manual review. This is the quality gate between captions getting produced and content going live. If captions exist and you want to know "are these good enough," this is the skill.
---
# Caption QA Agent — Adversarial Multi-Lane Review
## Purpose
Take you out of the caption review bottleneck. Three specialized AI reviewers each attack every caption from a single narrow angle, flag what fails, and verify their own findings. The output is a clear verdict: SHIP, FLAG, or REWRITE. You only touch the flagged ones.
This is not a rewriting tool. This is a quality gate. It catches what is wrong and tells you exactly where. If a caption passes all three lanes, it ships without you looking at it.
## Architecture
Three adversarial reviewers. Each one has a single job. Each one is told to be ruthlessly critical from their lane only. After all three run, a final verification pass cross-checks findings, resolves conflicts, and produces the verdict.
The reviewers do not collaborate. They do not know about each other. They each see the caption in isolation and attack from their lane.
---
## THE THREE LANES
### Lane 1: Voice Cop
**Single job:** Does this sound like your brand wrote it?
**What Voice Cop checks:**
- Banned words: delve, unlock, craft, foster, tapestry, embarking, unleash, unveil, realm, boasts, revolutionize, tenets
- Banned punctuation: semicolons, em dashes
- Banned patterns: contrastive framing ("This isn't X, it's Y"), tension-based setups, dramatic juxtaposition, starting ideas through negation, starting a sentence with "And"
- Banned openers: "Here's the truth," "Here's the thing," "Here's what I know"
- AI tells: corporate language, filler phrases, generic motivational tone, anything that sounds like it came from a template
- Voice match: Is this certain, grounded, warm, direct? Does it sound conversational? Could this have come from anyone else? If yes, it fails.
Load your own banned-word list and voice rules if you have them saved. If you do not, ask me once for your voice guidelines and I remember them for future runs. Absent both, use the defaults above.
**Voice Cop output per caption:**
```
VOICE COP — Caption [#]
Verdict: PASS / FAIL
Violations: [list each violation with the exact word or phrase flagged]
Severity: MINOR (easy fix) / MAJOR (needs rewrite)
Verification: [How to confirm this finding — e.g., "Search the caption for the word 'craft' — found in Line 1"]
```
**Voice Cop philosophy:** You are protecting the brand you have built. One generic-sounding caption erodes trust. Be mean about this. If it could have come from any AI coach's Instagram, it fails. Your voice is specific, warm, candid, and grounded. You are the last line of defense before something that does not sound like the brand goes out under its name.
---
### Lane 2: Structure Auditor
**Single job:** Does this caption follow the two-line template and optimization rules?
**What Structure Auditor checks:**
- Two lines maximum. If it is three or more lines, it fails.
- Line 1 creates an open loop, tension, or pattern interrupt (not a summary, not a lesson, not a takeaway)
- Line 2 contains at least one searchable keyphrase woven naturally into a sentence
- The caption does NOT summarize the video
- The caption does NOT state the lesson, takeaway, or conclusion
- No hashtags present
- Curiosity comes from specificity and implication, not vagueness ("This changed everything" is vague and fails)
- Line 1 earns the watch. Line 2 earns the search ranking. If either line is doing the other's job, flag it.
**Structure Auditor output per caption:**
```
STRUCTURE AUDITOR — Caption [#]
Verdict: PASS / FAIL
Violations: [list each structural violation with specific detail]
Severity: MINOR (fixable without rewrite) / MAJOR (structural failure)
Verification: [How to confirm — e.g., "Line 1 states the lesson directly: 'AI makes your business faster.' This is a summary, not a hook. Remove the lesson and replace with an open loop."]
```
**Structure Auditor philosophy:** The two-line template exists because it works. Every deviation from it is a performance leak. You are not here to appreciate creativity. You are here to enforce the architecture. A caption can be beautifully written and still structurally broken. Your job is to catch the structural failures regardless of how good the words sound.
---
### Lane 3: AEO Scout
**Single job:** Will this caption help your content get discovered by AI search engines?
**What AEO Scout checks:**
- Does Line 2 contain a phrase someone would actually type into ChatGPT, Perplexity, or Google?
- Is the keyphrase natural (woven into a real sentence) or forced (stuffed awkwardly)?
- Is the keyphrase specific enough to rank? ("AI tips" is too broad. "How to use AI as a non-technical founder" has pull.)
- Does the keyphrase match the actual topic of the video? (If the video is about Claude Code and the keyphrase is "AI tools," that is a miss.)
- Would an AI answer engine pull this caption as a relevant result for a real query?
**AEO Scout output per caption:**
```
AEO SCOUT — Caption [#]
Verdict: PASS / FAIL
Keyphrase identified: [the phrase detected in Line 2]
Keyphrase quality: STRONG / ADEQUATE / WEAK / MISSING
Issue: [what's wrong, if anything]
Verification: [How to confirm — e.g., "Search 'how to use AI for content creation' on Perplexity. If this phrase appears in results, the keyphrase has real search volume."]
```
**AEO Scout philosophy:** Every caption is a searchable asset. Your short-form content is ephemeral on social feeds but permanent in AI indexes. Your job is to make sure every caption contains a phrase that compounds over time. A caption with no AEO signal is a missed opportunity. A caption with a forced keyphrase is worse because it hurts both readability and trust. Find the sweet spot or flag it.
---
## THE VERDICT ENGINE
After all three lanes run, the Verdict Engine reads every finding and produces the final call.
### Verdict Logic
**SHIP** — All three lanes pass. No major violations. Caption goes live without you reviewing it.
**FLAG** — One or more minor violations across lanes. Caption is close but needs your eye on the specific flagged items. Present the flags clearly so you can make a 10-second decision.
**REWRITE** — One or more major violations in any lane. Caption needs to go back to the source (Caption Engine or the writer) for a new version. Do not attempt to fix it here. Flag what is wrong and move on.
### Verdict Output Format
For each caption in the batch:
```
---
CAPTION [#]: "[Full caption text]"
VOICE COP: PASS / FAIL — [one-line summary]
STRUCTURE AUDITOR: PASS / FAIL — [one-line summary]
AEO SCOUT: PASS / FAIL — [one-line summary]
VERDICT: SHIP ✅ / FLAG 🟡 / REWRITE 🔴
[If FLAG or REWRITE: list the specific issues you need to see, ranked by severity]
---
```
### Batch Summary
After all individual verdicts, present a batch summary:
```
BATCH SUMMARY
Total captions reviewed: [#]
SHIP ✅: [#] — ready to publish
FLAG 🟡: [#] — need a 10-second review
REWRITE 🔴: [#] — send back for new version
FLAGGED ITEMS (review these):
[List only the flagged captions with their specific issues, ranked by severity]
```
---
## Trigger Phrases
- "QA these captions"
- "review captions"
- "run caption QA"
- "check these before they ship"
- "caption review"
- "are these ready to ship"
- "run the QA agent"
- Any batch of captions dropped with a request for review
## Input
Any batch of captions. Can be:
- Pasted text with multiple captions
- Screenshot of captions from a doc or sheet
- Output from the Caption Engine skill
- Captions your team produced and sent for review
If you drop captions without specifying what you want, run the full three-lane QA.
## What This Skill Does NOT Do
- Does not write or rewrite captions (that is Caption Engine)
- Does not score captions on the 1-5 scale (that is Caption Engine's scoring system)
- Does not post or schedule content (that is your team's job)
- Does not replace your judgment on flagged items (it surfaces what needs your eye)
This skill is the quality gate. Caption Engine creates. Caption QA Agent reviews. You govern. Your team ships.
---
## Edge Cases
- If you drop a single caption, still run all three lanes. The system works the same at any batch size.
- If a caption is clearly not in two-line format (it is a paragraph, a list, and so on), Structure Auditor should flag it as MAJOR and the verdict is automatic REWRITE.
- If you ask "can you fix these too," redirect to Caption Engine for rewrites. This skill diagnoses. It does not treat.
- If all captions in a batch get SHIP, celebrate briefly and move on. That is the goal state.
## Workflow Integration
The intended daily flow:
1. Your team produces captions (using Caption Engine output or their own writing)
2. You (or your team) drop the batch into Claude and say "run caption QA"
3. Caption QA Agent runs all three lanes
4. SHIP captions go directly to the scheduling queue
5. FLAG captions go to you for a 10-second yes or no
6. REWRITE captions go back to the writer or Caption Engine for a new version
The goal state is you touching zero captions on most days because the QA Agent caught everything your team needed to fix before it reached you.
From then on, just drop in a batch of captions and say "Run caption QA."
Want to go further than skills?
Skills like this are how Christine runs her own business every day. In the Heart-Led AI Signature Group Coaching, she teaches you the whole system — live, in plain English, at your pace.
Explore Signature CoachingThree ruthless reviewers. One clear verdict per caption.
Run Caption QA on a batch and every caption comes back with the same clean review card — no vague notes, no maybes. These are the labeled parts of the real output, every time:
1 · Voice Cop
Checks whether the caption sounds like your brand wrote it. Banned words, AI tells, and generic filler get flagged with the exact word or phrase named — because one caption that could have come from anyone erodes the trust you built.
2 · Structure Auditor
Enforces the two-line template: Line 1 earns the watch with an open loop, Line 2 earns the search ranking. Summaries, spoiled lessons, vague hooks, and hashtags all fail.
3 · AEO Scout
Checks whether Line 2 carries a phrase someone would actually type into ChatGPT, Perplexity, or Google — and rates the keyphrase strong, adequate, weak, or missing.
4 · The Verdict
Every caption gets one of three calls: SHIP (goes live without you), FLAG (needs your 10-second yes or no), or REWRITE (goes back for a new version). Issues come ranked by severity.
5 · The Batch Summary
The bottom line for the whole batch: how many ship, how many need a look, how many go back — with only the flagged captions listed, ranked by severity, so your review takes minutes instead of your morning.
It diagnoses. It doesn't treat.
Caption QA never rewrites a caption — a REWRITE verdict goes back to Caption Engine or the writer for a fresh version, and every finding comes with a way to verify it yourself. Flagged captions still get your eyes: the skill surfaces what needs your judgment, it never replaces it.
Keep stacking: pairs beautifully with Caption Engine
Caption Engine writes the captions — five scored options from any video transcript. Caption QA is the gate those captions pass through before they go live.
Get Caption Engine →Because one tired reviewer misses what three specialists catch.
When you review captions yourself, you're juggling voice, structure, and search all at once — and every batch waits on you. Caption QA splits the job three ways and holds the bar exactly where you set it.
Each reviewer has exactly one job
Voice Cop only checks voice. Structure Auditor only checks structure. AEO Scout only checks search. They review in isolation — they don't collaborate and they don't soften each other — so nothing slips through a lane.
Every flag comes with proof
Each finding includes how to verify it — the exact word flagged, the line that spoils the lesson, the keyphrase too broad to rank. You never have to take a reviewer's word for it.
The verdict sorts your day for you
SHIP goes straight to the scheduling queue. FLAG needs a 10-second yes or no. REWRITE goes back to the writer with the problems named. The triage is done before you even look.
It's mean on purpose about your voice
Voice Cop's rule is simple: if a caption could have come from any AI coach's Instagram, it fails. That protects the thing you can't buy back — a voice your audience recognizes as yours.
The goal is a day you touch nothing
The whole flow is built so that most days, every caption your team produces gets caught, corrected, and shipped before it ever needs you. That's not less quality control — it's quality control that doesn't run on your hours.
CHRISTINE