For media buyers, creative strategists and founders who want AI to break down video ads
AI ad creative analysis: what it can read
AI ad creative analysis means a model watches an ad and describes how it is built: the hook, the order of the beats, the claims, the offer and the call to action. A model can read only what is in the file, which is the frames, the speech and the text on screen, so it can say what an ad does and when, and it can never say what the ad earned. The answer is only as good as what the model was given: a video, a set of frames, or only the thumbnail and the copy. Ask for the second of every claim, because a second can be checked and an adjective cannot.
Published 10 min read
The short version
- A video ad file holds 3 things a model can read: the frames, the speech and the text on screen. It holds 0 numbers about spend, ROAS or CPA.
- Of 3 large model APIs, 1 takes a video file directly. Gemini samples it at 1 frame per second by default. Claude and OpenAI take images, so a video is cut into frames first.
- At 1 frame per second, a shot shorter than 1 second can fall between 2 samples. A 30 second ad with a cut every 0.8 seconds has about 37 shots and 30 samples.
- Ask for the second of every claim and every beat. 'Price on screen at 0:08 for 3 seconds' can be checked in 10 seconds, and 'the offer is strong' cannot.
- A model that sees only the thumbnail and the copy reads 1 frame of the ad. The hook, the pacing and the offer happen in time, and none of them is in 1 frame.
What is on this page and on no other
Every page about these products carries the same prices and the same feature lists. These are the facts this one produced itself, and how each one was made.
1 of 3 large model APIs accepts a video file. Gemini samples video at 1 frame per second by default. Claude accepts JPEG, PNG, GIF and WebP and uses only the first frame of an animation, and OpenAI accepts PNG, JPEG, WebP and non-animated GIF.
Gemini's own documentation says its default 1 frame per second may miss details in videos with rapid motion or quick scene changes, and puts a video at about 100 tokens a second at low media resolution.
A 30 second ad at 30 frames per second holds 900 frames. A thumbnail shows 1 of them, and sampling at 1 frame per second shows 30.
What an AI can read from an ad, and what it cannot
Everything in the file. Nothing about the auction.
| In the file | A model can report | Example |
|---|---|---|
| The frames | What is on screen and when: faces, the product, the setting, the cuts | Product in hand at 0:02, first cut at 0:01.2 |
| The speech | What is said and when, word for word | 'I tried everything for my back' at 0:00.5 |
| The text on screen | Captions, claims, prices and the call to action, with their seconds | '50% off today' from 0:21 to 0:25 |
| The music and sound | Whether there is voice-over, music only, or silence | Music only, no voice |
| Not in the file | Where it lives |
|---|---|
| Spend, impressions, CPM | The ad account, or the Meta Ad Library for some ads in the EU |
| ROAS, CPA, conversion rate | The ad account and your attribution tool |
| Hook rate and hold rate | The ad account, from view data |
| Who the ad was shown to | The ad account's targeting, never the video |
The answer depends on what the model was given
A video, a set of frames, or a thumbnail and the copy are three different inputs.
| What you send | What the model sees | What it misses |
|---|---|---|
| The thumbnail and the ad copy | 1 frame and the primary text | The hook, the pacing, the offer, the call to action, anything spoken |
| A few screenshots you took | The frames you chose | Everything between them, and the speech |
| The video to Gemini | 1 frame per second by default, and the audio | Shots shorter than 1 second can fall between samples |
| The video to Claude or OpenAI | Nothing until the video is cut into images. Claude uses the first frame of an animated GIF only | The speech, unless you send a transcript beside the frames |
| A tool that reads the ad for you | What the tool samples, and what it writes down | Depends on the tool. Ask for the second of every claim and check 3 of them |
Sampling decides more than the model does. A 30 second ad exported at 30 frames per second holds 900 frames. At 1 frame per second a model sees 30 of them. That is enough for an ad with a shot every 3 seconds, and it is not enough for an ad that cuts more often: a 30 second ad with a cut every 0.8 seconds has about 37 shots, so some shots are never sampled at all.
Text on screen has the same problem. A price that is up for half a second is up for less than 1 sample. So when a reading says 'no price shown', check whether it could have seen one.
The questions worth asking a model about an ad
Each one has an answer in the file, and each answer has a second.
What happens in the first 3 seconds?
Ask for what is on screen, what is said and what text appears, each with its second. The hook is a claim, a question, a result or a face, and the answer should say which.
What are the beats, with the start and length of each?
Hook, problem, demonstration, proof, offer, ask. The order and the length of each part is the structure you borrow in a brief.
What does the ad claim, word for word?
Every claim on screen or in the voice-over, quoted, with the second it is legible. A claim you cannot check against the frame is a claim you should not brief from.
When does the product first appear, and the first face?
Two seconds that decide how the ad opens. An ad that shows the product at 0:01 and one that shows it at 0:12 are different ads with the same script.
Does it work with the sound off?
Ask whether the story is carried by the text on screen and the picture alone. On a feed where most people scroll muted, that answer changes the brief.
What could it not read?
A good reading names what it could not resolve, such as text too small to read or a claim spoken over music. A reading that has an answer for everything has not told you where it guessed.

A checklist before you trust an AI ad analysis
Five checks, and the first one takes 10 seconds.
| Check | How | A fail looks like |
|---|---|---|
| Seconds are given | Every claim and beat has a time | 'The ad opens with a strong hook' |
| 3 seconds hold up | Scrub to 3 of the times it names | The text it quotes is not on screen at that second |
| No performance claim | It says nothing about spend or ROAS from the video alone | 'This ad likely converts well' |
| It saw the whole ad | The last beat ends near the ad's length | The beats stop at 0:15 on a 0:40 ad |
| It names what it missed | A list of what it could not read | An answer for every field, and no doubt anywhere |
Playhead is built to pass all five. It reads the ad from its link or from a file, so nothing is cut or sampled on your side. It returns the hook, every beat with its start and length, every claim quoted with its second, the cuts counted off the file, the second of the first face and the first product, and whether the ad works muted. It names what it could not resolve. And it never states spend or ROAS, because it connects to no ad account and a video holds none.
Put up to 10 ads in one set to see what the winners share, and ask a new question of any saved ad later for less than the first reading. It works in the browser and inside Claude or ChatGPT through its MCP server, on every plan.
The verdict
AI ad creative analysis is worth doing when the model sees the whole video and answers with seconds. A thumbnail shows 1 frame of the 900 in a 30 second ad, and 1 frame holds no hook, no pacing and no offer. Send the video, ask for the second of every beat and claim, check 3 of them, and keep spend and ROAS out of the question. Playhead does that from a link, free for 50,000 credits a month, and says what it could not read.
Common questions
Can AI analyse a video ad?
Yes, when the model receives the video or its frames. Gemini accepts a video file and samples it at 1 frame per second by default. Claude and OpenAI accept images, so the video is cut into frames first. A tool built for ads does the sampling for you and returns the beats and claims with their seconds.
Can AI tell me if an ad will perform well?
Not from the video alone. A video file holds the frames, the speech and the text on screen, and no spend, ROAS or CPA. A model can say what an ad does and when. What it earned comes from the ad account, and a prediction without that data is a guess.
What is the best way to analyse Meta ads with AI?
Give the model the ad video, not the thumbnail, and ask for the second of every beat and claim. Check 3 of the seconds against the ad. Then read up to 10 ads as a set: what the ads that keep running have in common is a better brief than any one reading.
Is ChatGPT or Claude enough to analyse an ad?
Both can read frames you send them, and neither reads a video file directly through its API, according to each provider's documentation read on 26 September 2026. That means you choose which frames the model sees. For an ad with a cut every second or less, the frames you do not send are the part it cannot report.
How much does AI ad analysis cost?
Playhead is free for 50,000 credits a month, and the paid plans are $19, $39 and $109 a month. A general model through its API bills by the tokens of the frames and audio you send: Gemini's documentation puts a video at about 100 tokens a second at low media resolution.