Video Generator, part 7: testing whether a character stays the same
Two people leave a bakery. One hands the other a croissant, and they both laugh. I watch it yet again, checking whether the glasses have disappeared.
This is one of my working experiments in Vini Studio. From the outside, it might look like I’m just generating cute cartoons. I like the cartoons too. But right now I’m interested in something quite specific: will a character still look like themselves once the picture starts moving?
In part five, I wrote about references for characters and objects. In part six, I showed how those pieces were becoming a studio. Behind those changes is some less visible work: designing a test, getting a result, and working out what it actually lets me claim. Sometimes that means crossing out my own explanation.
Mira, Bo, and one croissant
For this experiment, I picked two fictional adult characters. Mira has a red braid, round glasses, and a yellow coat. Bo is shorter, with a beard, a red hat, and a green sweater. They’re hard to mix up even in a small preview.

Mira. One sheet brings together the views and distinguishing features I use to recognize her in later frames.

Bo. His orange sneakers are also handy for testing: they stand out in a wide shot.
The story is deliberately simple. I need to see the characters walking together and passing an object. If I send them off to save the galaxy straight away, it’ll be harder to tell where things went wrong.
Even this scene has several requirements. These exact two characters need to stay in the shot. The croissant needs to pass from one hand to the other. Their appearance has to hold up in motion, beyond that first beautiful frame.
One of the results from Gemini Omni Flash. Watch the characters as they move and pass the croissant. Keep an eye on the bakery signs too: we’ll come back to those.
First, I decide what I’m testing
“The model makes good videos” is too broad a claim to base a development decision on. I need a question I can answer by looking at the result.
So I split up the tasks. In one experiment, the characters need to appear in a new scene based on their reference images. In another, I need to replace a character in an existing video while preserving what happens around them. I also tested the creation of the first frame separately: an error at that stage can carry over into the video.
These are different tests. A good opening frame doesn’t tell me what will happen to a face when the head turns. And successfully replacing a person in an existing video says nothing on its own about how the model will invent motion from scratch.

The source clip for the replacement test: a mannequin walks down a street. The frames run chronologically from left to right.

The result from Kling O3: Mira walks in place of the mannequin. Now I can compare her appearance separately from how well the street and movement were preserved. The replacement worked in this run: the street is recognizable, and Mira walks toward the camera, following the action in the source clip.
There’s another wrinkle in this comparison: the same task doesn’t always mean the same request. Different models accept images in different ways. I tested both the studio’s current setup and alternatives based on the model providers’ documentation. When several conditions change at once, the experiment helps test the approach as a whole, but it doesn’t let me attribute the result to one clever bit of wording.
What the different models produced
Here are a few observations from these runs. A single scene makes it easy to see how recognizable characters and a correctly executed task can be two different things.
Grok and Gemini: recognizable characters
Grok Imagine 1.5 preserved Mira and Bo: both are recognizable from the first frame, and the sheets showing their different views didn’t become part of the scenery. That’s already a useful result for me. I can recognize Mira’s braid, her glasses, and Bo’s clothes while watching an ordinary scene outside a bakery.

Grok Imagine 1.5. The croissant passes to Mira, and the characters remain recognizable throughout the scene. The model also drew lettering on the sign, so this result didn’t meet every requirement.
Gemini Omni Flash also managed the croissant handoff in the clip above. I like that I can check the whole action here: Bo had the object, Mira received it, and both reacted. But this version included lettering on the bakery too. Both examples did well at preserving the characters, while the requirement about text on screen still needed a separate mark against it.
Wan and HappyHorse: who’s in the frame
In one Wan 3.0 variant, Mira and Bo remained recognizable, but by the end the camera had moved so close that they no longer fit entirely in the frame. Look at their faces, and everything is familiar. Look at the staging, and it has changed. Look closely, and Bo also appears taller than Mira in these frames, even though he is shorter in the original references.
Wan 3.0, one variant of the experiment. Watch the final seconds: Mira and Bo are recognizable, but more and more of their bodies fall outside the frame.
HappyHorse 1.1 took a different small liberty: in one clip, Bo was partly hidden by the door at the start; in another, Mira lingered behind it. In an ordinary scene, coming out one after the other would be perfectly natural. But my task required both to be visible from the beginning. Pleasant animation can easily distract from an error like this.
Seedance: the test was over before the clip could begin
Seedance 2.0 Fast and 2.5 rejected Mira’s sheet: the filter reported that the image might contain a real person. Mira is fictional, yet Bo’s sheet passed the check. Frustrating: I wanted to see how she moved and didn’t even get that far.
With other variants of the input images, both models produced clips with recognizable characters. So I record the rejection separately from generation quality. It shows a problem with accepting a particular input; there’s no motion to judge in a clip that doesn’t exist.
Recognizing the character isn’t enough
There was plenty to like in the first results: recognizable characters, similar clothing, a lively scene. It would have been easy to tick the box and move on.
But here’s a separate result from the first-frame test with Nano Banana 2:

Nano Banana 2. The characters are there. So is the lettering, even though the instructions said not to draw any. Character recognizability and compliance with constraints produced different results here.
I like this picture myself. That’s exactly why it helps to write down what I intend to check before generating anything. Otherwise, it’s very easy to decide afterward that the signs actually make the scene nicer and quietly adjust the criterion to fit a result I like.
Here’s the first frame from another run, with GPT Image 2.5:

GPT Image 2.5. The characters are arranged differently: Bo is on the left, Mira on the right. This version of the task didn’t specify sides, so I don’t count that change as an error.
I also like this pair as a check on my own judgment. Lettering that was explicitly forbidden is a violation. A position I pictured in my head but never specified isn’t something the model has to guess. Mix those things up, and the comparison quickly becomes a contest to match the picture in my imagination.
Researchers make a similar distinction in VBench: consistency of a character’s appearance, motion smoothness, and the degree of movement are evaluated separately. My examples haven’t been evaluated with that benchmark; what I find useful here is the way it frames the problem.
There’s a trap in the other direction too. You can focus so much on preserving appearance that you stop noticing weak motion. In ConsistI2V, the authors note that their method sometimes produces motion with a limited range. That’s a limitation of that particular method, but it’s a good reason to watch the action when checking whether a character stays themselves. In my scene, keeping Mira’s glasses isn’t enough. The croissant still has to change hands.
The most useful part is when the hypothesis doesn’t hold up
Plenty of things worked in the tidy test scene with Mira and Bo. But a working project had previously produced very different results: reference images turned into inserts within the video, and extra characters appeared. It was tempting to write this off as an unlucky generation.
I tried to reproduce the failure. With the test characters, individual changes to the conditions didn’t bring it back. But when I reran the project’s request with its original materials, the characteristic errors returned. The conditions weren’t an exact match: one of the tests used a different resolution. Still, I now had an example I could use to test explanations.
The first explanation sounded convincing: the source images were the problem. Then I changed how those images were supplied and got an unwelcome answer. Some defects disappeared; others remained or gave way to new ones. I’d decided too soon that I understood it.
The next experiments showed that I needed to account for both the scene description and the visual inputs. Their combination mattered more than my first explanation had suggested. What matters to me here is the moment when the next test forced me to revise my conclusion.
That’s how I understand a scientific approach to my work. First, there’s an observation and a suspected cause. Then I need to design an experiment in which that explanation might fail, and actually run it. An inconvenient result still stays in the notes.
I try to change one condition at a time wherever possible. When several conditions change together, I record that separately. And even a careful comparison of individual generations is a lead for the next experiment: the randomness hasn’t gone away.
What this test doesn’t prove yet
The main matrix had one run per combination of conditions. That was enough to spot working options and find problems worth investigating. It doesn’t measure how often errors occur.
If a character stayed consistent once, I know that result was possible. I don’t yet know how many attempts the next project will take. Even a recurring failure doesn’t justify declaring a model unsuitable for every scene.
The main test also used stylized adult characters and short scenes. Some of the planned variants weren’t run. Conclusions about a long episode or a different visual style will need their own examples and repeated runs.
For the next stage, I need batches of runs across several scenes. I want to understand how often I have to redo a result and which errors return after changes. The images in this article show what happened in individual experiments; assessing reliability is still ahead of me.
What makes its way into the studio
This testing is already useful for development. It gives me a better idea of where to look for a problem and which example to use when checking a future fix. Sometimes the task needs a more precise description. Sometimes I need to investigate how the studio passed the materials to the model. Switching models doesn’t resolve that question on its own.
In the first part of this series, I took a similar approach to the scriptwriter. This time, I’m looking at characters with a croissant instead of text, but the habit is the same: save the input and the result so I can return to my own conclusion later and check it.
I like that this work also leaves me with pictures like these. I can enjoy looking at Bo simply because he’s charming. Then I can play the clip again and spot an error I missed in that first impression.