For media buyers, creative strategists and founders who want AI to break down video ads

AI ad creative analysis: what it can read

AI ad creative analysis means a model watches an ad and describes how it is built: the hook, the order of the beats, the claims, the offer and the call to action. A model can read only what is in the file, which is the frames, the speech and the text on screen, so it can say what an ad does and when, and it can never say what the ad earned. The answer is only as good as what the model was given: a video, a set of frames, or only the thumbnail and the copy. Ask for the second of every claim, because a second can be checked and an adjective cannot.

By The Playhead team

We build the video layer these numbers are measured on. Every figure about Playhead comes from the source file it lives in.

Published 10 min read

The short version

  • A video ad file holds 3 things a model can read: the frames, the speech and the text on screen. It holds 0 numbers about spend, ROAS or CPA.
  • Of 3 large model APIs, 1 takes a video file directly. Gemini samples it at 1 frame per second by default. Claude and OpenAI take images, so a video is cut into frames first.
  • At 1 frame per second, a shot shorter than 1 second can fall between 2 samples. A 30 second ad with a cut every 0.8 seconds has about 37 shots and 30 samples.
  • Ask for the second of every claim and every beat. 'Price on screen at 0:08 for 3 seconds' can be checked in 10 seconds, and 'the offer is strong' cannot.
  • A model that sees only the thumbnail and the copy reads 1 frame of the ad. The hook, the pacing and the offer happen in time, and none of them is in 1 frame.

What is on this page and on no other

Every page about these products carries the same prices and the same feature lists. These are the facts this one produced itself, and how each one was made.

  • 1 of 3 large model APIs accepts a video file. Gemini samples video at 1 frame per second by default. Claude accepts JPEG, PNG, GIF and WebP and uses only the first frame of an animation, and OpenAI accepts PNG, JPEG, WebP and non-animated GIF.

    Counted The video understanding page of the Gemini API, Anthropic's vision page and OpenAI's images and vision page, each read on 26 September 2026.

  • Gemini's own documentation says its default 1 frame per second may miss details in videos with rapid motion or quick scene changes, and puts a video at about 100 tokens a second at low media resolution.

    Counted The Gemini API video understanding page, read on 26 September 2026, quoted from its sampling and token sections.

  • A 30 second ad at 30 frames per second holds 900 frames. A thumbnail shows 1 of them, and sampling at 1 frame per second shows 30.

    Counted Arithmetic from the frame rate of a standard 30 fps export and the 1 fps default on the Gemini API video understanding page, read on 26 September 2026.

What an AI can read from an ad, and what it cannot

Everything in the file. Nothing about the auction.

What a video ad file holds, and so what a model can report
In the fileA model can reportExample
The framesWhat is on screen and when: faces, the product, the setting, the cutsProduct in hand at 0:02, first cut at 0:01.2
The speechWhat is said and when, word for word'I tried everything for my back' at 0:00.5
The text on screenCaptions, claims, prices and the call to action, with their seconds'50% off today' from 0:21 to 0:25
The music and soundWhether there is voice-over, music only, or silenceMusic only, no voice
What a video ad file does not hold, and so no model can report from it
Not in the fileWhere it lives
Spend, impressions, CPMThe ad account, or the Meta Ad Library for some ads in the EU
ROAS, CPA, conversion rateThe ad account and your attribution tool
Hook rate and hold rateThe ad account, from view data
Who the ad was shown toThe ad account's targeting, never the video

The answer depends on what the model was given

A video, a set of frames, or a thumbnail and the copy are three different inputs.

Three inputs, and what each one lets a model see. Figures from each provider's documentation, read on 26 September 2026
What you sendWhat the model seesWhat it misses
The thumbnail and the ad copy1 frame and the primary textThe hook, the pacing, the offer, the call to action, anything spoken
A few screenshots you tookThe frames you choseEverything between them, and the speech
The video to Gemini1 frame per second by default, and the audioShots shorter than 1 second can fall between samples
The video to Claude or OpenAINothing until the video is cut into images. Claude uses the first frame of an animated GIF onlyThe speech, unless you send a transcript beside the frames
A tool that reads the ad for youWhat the tool samples, and what it writes downDepends on the tool. Ask for the second of every claim and check 3 of them

Sampling decides more than the model does. A 30 second ad exported at 30 frames per second holds 900 frames. At 1 frame per second a model sees 30 of them. That is enough for an ad with a shot every 3 seconds, and it is not enough for an ad that cuts more often: a 30 second ad with a cut every 0.8 seconds has about 37 shots, so some shots are never sampled at all.

Text on screen has the same problem. A price that is up for half a second is up for less than 1 sample. So when a reading says 'no price shown', check whether it could have seen one.

The questions worth asking a model about an ad

Each one has an answer in the file, and each answer has a second.

  1. What happens in the first 3 seconds?

    Ask for what is on screen, what is said and what text appears, each with its second. The hook is a claim, a question, a result or a face, and the answer should say which.

  2. What are the beats, with the start and length of each?

    Hook, problem, demonstration, proof, offer, ask. The order and the length of each part is the structure you borrow in a brief.

  3. What does the ad claim, word for word?

    Every claim on screen or in the voice-over, quoted, with the second it is legible. A claim you cannot check against the frame is a claim you should not brief from.

  4. When does the product first appear, and the first face?

    Two seconds that decide how the ad opens. An ad that shows the product at 0:01 and one that shows it at 0:12 are different ads with the same script.

  5. Does it work with the sound off?

    Ask whether the story is carried by the text on screen and the picture alone. On a feed where most people scroll muted, that answer changes the brief.

  6. What could it not read?

    A good reading names what it could not resolve, such as text too small to read or a claim spoken over music. A reading that has an answer for everything has not told you where it guessed.

A list of the claims in an ad, each quoted word for word with the second it is legible and a frame beside it.
What an answer to 'what does the ad claim, word for word' looks like in Playhead: each claim quoted, with its second and the frame it appears on.

A checklist before you trust an AI ad analysis

Five checks, and the first one takes 10 seconds.

Five checks for any AI reading of an ad
CheckHowA fail looks like
Seconds are givenEvery claim and beat has a time'The ad opens with a strong hook'
3 seconds hold upScrub to 3 of the times it namesThe text it quotes is not on screen at that second
No performance claimIt says nothing about spend or ROAS from the video alone'This ad likely converts well'
It saw the whole adThe last beat ends near the ad's lengthThe beats stop at 0:15 on a 0:40 ad
It names what it missedA list of what it could not readAn answer for every field, and no doubt anywhere

Playhead is built to pass all five. It reads the ad from its link or from a file, so nothing is cut or sampled on your side. It returns the hook, every beat with its start and length, every claim quoted with its second, the cuts counted off the file, the second of the first face and the first product, and whether the ad works muted. It names what it could not resolve. And it never states spend or ROAS, because it connects to no ad account and a video holds none.

Put up to 10 ads in one set to see what the winners share, and ask a new question of any saved ad later for less than the first reading. It works in the browser and inside Claude or ChatGPT through its MCP server, on every plan.

The verdict

AI ad creative analysis is worth doing when the model sees the whole video and answers with seconds. A thumbnail shows 1 frame of the 900 in a 30 second ad, and 1 frame holds no hook, no pacing and no offer. Send the video, ask for the second of every beat and claim, check 3 of them, and keep spend and ROAS out of the question. Playhead does that from a link, free for 50,000 credits a month, and says what it could not read.

Common questions

Can AI analyse a video ad?

Yes, when the model receives the video or its frames. Gemini accepts a video file and samples it at 1 frame per second by default. Claude and OpenAI accept images, so the video is cut into frames first. A tool built for ads does the sampling for you and returns the beats and claims with their seconds.

Can AI tell me if an ad will perform well?

Not from the video alone. A video file holds the frames, the speech and the text on screen, and no spend, ROAS or CPA. A model can say what an ad does and when. What it earned comes from the ad account, and a prediction without that data is a guess.

What is the best way to analyse Meta ads with AI?

Give the model the ad video, not the thumbnail, and ask for the second of every beat and claim. Check 3 of the seconds against the ad. Then read up to 10 ads as a set: what the ads that keep running have in common is a better brief than any one reading.

Is ChatGPT or Claude enough to analyse an ad?

Both can read frames you send them, and neither reads a video file directly through its API, according to each provider's documentation read on 26 September 2026. That means you choose which frames the model sees. For an ad with a cut every second or less, the frames you do not send are the part it cannot report.

How much does AI ad analysis cost?

Playhead is free for 50,000 credits a month, and the paid plans are $19, $39 and $109 a month. A general model through its API bills by the tokens of the frames and audio you send: Gemini's documentation puts a video at about 100 tokens a second at low media resolution.

Written by

The Playhead team

We build the video layer these numbers are measured on. Every figure about Playhead comes from the source file it lives in.

How we write these articles, and how to send a correction