Why Gemini should be the eyes and ears for AI short dramas: 31 models tested on video and audio
An agent making AI short dramas can't watch or hear its own clips. We sent one mp4 to 31 models: GPT, Claude and Grok never saw it; Gemini and Qwen did.

You ask Claude Code or Codex to make a short drama. It renders dozens of takes, and it can’t watch a whole clip or hear one. Was the line spoken right, do the lips move, where should the cut go: you end up opening every clip yourself. Our first episode took 44 takes to keep 29.
Our fix is to give the agent eyes and ears. Send the whole mp4 to Gemini, let it watch and listen, and get back a report with timestamps. The agent then decides from the report whether to redo the take or cut it.

We picked Gemini by testing: this 5-second take went to 31 models with no hints, and we asked only what is said and what happens.
| Model | Result |
|---|---|
| 8 Gemini models | Described the picture and heard the line; 4 of them misheard one character |
| Alibaba Qwen3.8-Omni-Flash | Described the picture and heard the line |
| 8 GPT models | “I did not receive a video” |
| 12 Grok models | “I did not receive any video” |
| Claude Opus 5.5, Sonnet 5.5 | Returned an error |
The GPT and Grok APIs take no video, so the mp4 never reached them. We told them to say so if nothing arrived, and all 20 did.
What separates Gemini from GPT, Claude and Grok
| Gemini | GPT, Claude, Grok | |
|---|---|---|
| Picture | Takes the whole video, one frame a second by default | No video; you extract frames and send images |
| Sound | Goes into the same model as the picture | No way in; you add speech recognition |
| Time | A timestamp on every second | You label it in the prompt |
| Tokens | About 90 for a second of video, at low resolution | 450 to 570 for one frame |
Since its first version, Gemini has been trained with video frames and audio as inputs alongside text, and Google trains it on YouTube videos. The model docs for GPT, Claude and Grok list text and image as inputs.
So it can tell you which line runs from which second to which, and it is cheap: our 2 min 21 s episode costs it 13,000 tokens to read.
The four jobs we give it
| Job | What it returns | What the agent does with it |
|---|---|---|
| Check the lines | A word-for-word transcript, with the start and end of each line | A line that differs from the script means the take is redone |
| Find the in and out points | The second before the first action and the second after the last line | Hands them to ffmpeg |
| Spot defects | A prop that changes shape, eyes that change colour, a person appearing from nowhere, each with its second | Decides whether to keep the take |
| Be the first viewer | Having never read the script, it watches the episode and retells the story | Whatever it can’t retell was not shown clearly |
Which one to use
Nine models could watch and listen. We took three and ran each twice on the same takes and the same episode:
| Gemini 3.8 Flash (thinking high) | Gemini 3.1 Pro (thinking low) | Qwen3.8-Omni-Flash | |
|---|---|---|---|
| Lines in single takes, names aside | All correct, but one piece of footage got an empty answer every time | All correct | All correct |
| Start of a line | Off by 0.05 s on average | Off by 0.1 s on average | Right on half the takes, 0.7 to 1.4 s late on the rest |
| 28 cuts in the episode, within half a second | 24, then 11 | 28, then 28 | 22, then 21 |
| Wait per take | 9 s | 27 s | 42 s |
| With the audio track removed | Invented lines both times | Invented no lines | Invented no lines; once reported water sounds |
- Use Gemini 3.8 Flash for single takes. It is the fastest, and its lines and timing are right. For the take it answers with nothing, use one of the other two. We never found out why it does that.
- Use Gemini 3.1 Pro for a whole episode. It found every cut and all 10 lines in both runs.
- Have Qwen listen again when a line is in doubt. It is a second opinion from another vendor. Where a subtitle had dropped two characters, it wrote what was spoken. Keep thinking on; with it off, lines go missing.
Four ways it fools you
- No sound, and it still “hears”. We removed the audio track and sent the take again. 3.8 Flash returned lines both times, different words each time, timed to the mouth. Run ffprobe first to confirm the clip has an audio track. Keep the script line out of the prompt, let it transcribe blind, and compare with the script afterwards.
- Burned-in subtitles get read back. On the subtitled cut, three Gemini models wrote the rare name 苍鱼 character for character. Without subtitles they wrote 苍宇, 苍俞 and 苍余. One subtitle had dropped two characters of its line. Two of the three models dropped them too, and all three heard them once the subtitles were gone. To check the voice track, send the version without subtitles.
- Past one minute, 3.8 Flash writes seconds as minutes and seconds. In a 141-second episode it put 2:04.3 down as 204.3. It did this in both runs, and in the second it wrote the cuts that way too. Of the 5 line times it got wrong, 3 are still shorter than the episode, so checking against the length misses them. 3.1 Pro and Qwen did not do this in either run.
- Ask for too much at once and defects get missed. One take has three real defects: the staff changes shape, the eyes change colour mid-shot, and the man drops from a wave crest to flat sea. With ten items in one prompt, four Gemini models, two runs each, had 24 chances and found 7. Asked about defects alone, with the kinds to look for listed, they found 21.
Claude Opus 5.5, given frames and the same big prompt, found all three, along with some false ones. That was one run, at 5 times Gemini’s tokens. It is worth adding when picture defects are all you want.
Gemini looks at one frame a second by default, and 3.7 Flash twice reported a fast push-in as a hard cut. For frame-exact cuts use ffmpeg’s shot-change detection. It misses a cut between two shots of the same person at the same framing. Our episode has 4 of those, and 3.1 Pro found them in both runs.
Whether the voices sound good and whether a take stays is still a person’s call.
What to paste to your AI
Write a watch script for this project: send a whole mp4 to Gemini 3.8 Flash (thinking high), have it watch and listen, and return JSON only:
a timeline, camera moves and hard cuts, the lines (word for word, start and end, whether the mouth moves), suggested in and out points.
Do not give it the script line; compare its transcript with the script in code. First check with ffprobe that the clip has an audio track.
If the answer comes back empty, retry with Gemini 3.1 Pro. Ask about defects in a separate call that lists the kinds to look for.
For a whole episode use Gemini 3.1 Pro on the cut without subtitles, and check the cuts against ffmpeg's shot-change detection.
Run it on every new take: redo the take when the line differs from the script, otherwise cut at the suggested points.
How we tested
- The material is one episode of a fantasy short drama we made, with picture and Chinese voices both generated by AI: 3 single takes, 1 clip of three takes joined, and the 2 min 21 s episode.
- The lines are scored against the script, checked with Whisper. The rare name counts as right when only the tone differs. The cuts are scored against the edit list.
- Every test ran twice. The subtitled episode and Claude on frames ran once.
- Temperature was 0. Google recommends the default 1.0 for Gemini 3, so we reran the empty-answer take twice at the default and twice at 1.0. It stayed empty.
- GPT, Grok, Claude and Gemini were called through our own model gateway, Qwen through Alibaba’s Bailian. Qwen caps base64 uploads at 10 MB, so it got a more compressed episode.
- The 22 models that saw no video saw none in the second run either. Sent the other way, as a file attachment, the mp4 got this from the GPT API:
unsupported MIME type 'video/mp4'.
Sources: the Gemini video understanding docs, the Gemini 1.0 report, CNBC on YouTube training data, and the model docs of OpenAI, Claude and Grok.
How the AI makes the video in the first place is in One sentence to Opus 5.5, a finished video.
If you know how to make GPT, Grok or Claude take an mp4 directly, or you spot a flaw in the test, tell us and we will rerun it.
How does your AI check the videos it makes?