A Singapore HDB kitchen at night, one fluorescent tube. A woman in her 40s faces camera. A man in a dull white singlet sits in the foreground with his back to us.
He says, “Have you eaten?”
She says nothing. Six seconds.
Watch firstNo labels until you’ve seen it
Sound on · 6 seconds
I watched this clip without knowing how it had been prompted. My verdict was contempt. My note: “good and subtle but clear.”
When we revealed the prompt, it was the control. We had specified the shot and his line, with no backstory or emotional direction.
One of my favourite performances came from asking for almost nothing.
The other control take read as nothing much. So this wasn’t a reliable shortcut. It was the question that made the experiment interesting: how much of a performance can we deliberately prompt, and how much do we find in it ourselves?
After 42 takes, our working approach is to start with a few visible actions and judge three things separately: whether the emotion reads, whether the performance feels believable, and whether the shot follows the instructions.
What we tested
We used ByteDance Seedance 2.5: six-second clips at 480p with native audio, all starting from the same image. We didn’t lock the seed, so every generation was a fresh draw. The 42 takes included the main comparison, controls, and follow-ups.
We chose rage, sadness, and contempt. Behind rage was his third affair; behind sadness, the anniversary of their son’s death. For contempt, he had disappeared when the loan sharks came. She paid off his gambling debts with two years of overtime. Now he was back, eating from her rice cooker.
We tried four approaches, keeping the shot and dialogue the same:
- Backstory as script: explain what happened before the shot.
- Feeling word: name the emotion, such as “She is heartbroken.”
- Visible actions: describe what her face and body do, in time order.
- Actor direction: give her circumstances and something to try to do, such as hiding what his words have done to her.
Two takes per approach, for each emotion, gave us 24 in the main comparison.
I was the only reviewer. The clips were shuffled and randomly named to hide their intended emotion and prompting method. I recorded the emotion I read, its clarity, and whether the performance felt real. With one face, one setting, and two takes per prompt version, this is a record of what we observed and what we’ll try next, not a reliable success-rate estimate.
Recognising a feeling isn’t the same as believing it
All eight rage takes read as anger. Seven of the eight sadness takes read as sadness. None of the eight contempt takes read as contempt.
| Prompting approach | Matching emotion readings |
|---|---|
| Backstory as script | 3 of 6 |
| Feeling word | 4 of 6 |
| Visible actions | 4 of 6 |
| Actor direction | 4 of 6 |
Three approaches tied on recognition. The performances still felt different.
The visible-actions prompt for rage included a palm slammed on the counter, a lunge, and bared teeth. It produced the clearest rage in the test. It also pulled the camera along with her movement. Later prompts explicitly asked her to stay seated and the camera to remain still. An obvious emotion wasn’t enough if the shot stopped following instructions.
RageFour ways of prompting it, one take each
Sound on · 24 seconds · Each tile plays in turnSound on · 24 seconds · Clips play in sequence
Tile labels use earlier names: “Script” is Backstory as script; “Describing the face” is Visible actions.
For sadness, “Heartbroken” produced a take I described as “very subtle, not too clear.” Actor direction produced something “quite subtle.” The visible-actions prompt was:
After he speaks, her chin trembles once. Her eyes fill with water but no tears fall. She swallows, looks down for a moment, then looks back at him.
Both takes read as sadness. On the second, I wrote “feels almost real.” That was the distinction I cared about: the visible-actions prompt tied the feeling word on recognition, but produced a performance I believed more.
SadnessFour ways of prompting it, one take each
Sound on · 24 seconds · Each tile plays in turnSound on · 24 seconds · Clips play in sequence
Tile labels use earlier names: “Script” is Backstory as script; “Describing the face” is Visible actions.
More precise instructions changed the result
Next, we revised the visible-actions prompts using facial-coding research: more precise eyebrow movements for sadness, and a mouth movement on only one side for contempt. The original contempt prompt had produced a smile on both sides of her mouth.
Five of the six revised takes read as intended, compared with four of the six original visible-actions takes. Both revised contempt takes read as contempt, although the performances were weak. One looked like a “fake smile, mixed a little with anger.”
The more detailed sadness felt less convincing. I scored it as sad at the time, with the note “an obvious sad face, but it doesn’t make me feel sad.” Watching it again, it reads to me as uncomfortable rather than sad. I preferred the earlier trembling chin, filling eyes, and swallow.
SadnessOriginal prompt, then the detailed rewrite
Sound on · 12 seconds · Left plays, then rightSound on · 12 seconds · Clips play in sequence
The extra detail helped us get contempt recognised. It also pushed sadness towards discomfort. More precise instructions didn’t automatically make the acting more convincing.
Getting contempt to work
We kept exploring contempt. Here is how the attempts compared:
| What we tried | Contempt readings |
|---|---|
| Research-informed visible actions | 2 of 2 |
| An eye roll | 1 of 2 |
| Silently mouthing his words back | 1 of 2 |
| Backstory plus visible actions, at three strengths | 3 of 6 |
The behavioural prompts brought their own problems. An eye roll looked like something in her eye. Asking her to mouth his words silently made her speak them aloud, which felt robotic.
For the combined approach, the gambling-debts backstory came before the shot and reaction. Low, medium, and high strength each produced one contempt reading and one miss.
The medium and high versions each gave me a contempt performance I particularly liked. This was the medium reaction:
She looks at him. Her chin lifts slightly and her lids lower. The left corner of her mouth tightens and pulls up, only that side; no smile reaches her eyes. A small breath out through her nose. She looks past his shoulder. She stays seated; the camera does not move.
ContemptFollow-up attempts
Sound on · 18 seconds · Three takes in sequence
The combined prompts produced two of my favourite contempt performances. We’ll test whether that quality holds up across more takes; recognition alone gave us no reason to prefer the combination.
Watch the four original contempt attempts
ContemptThe four original ways of prompting it, one take each
Sound on · 24 seconds · Each tile plays in turnSound on · 24 seconds · Clips play in sequence
Tile labels use earlier names: “Script” is Backstory as script; “Describing the face” is Visible actions.
What we’ll put in the next prompt
We’re starting with short sequences of visible actions, adding detail where a take misses something specific. For this kitchen scene, our working structure is:
BACKSTORY, WHEN RELEVANT: This happened before the shot begins. Do not show it: [The relevant relationship and event.] THE SHOT: Static over-the-shoulder shot. In the first second, he says, “Have you eaten?” The woman does not speak. We stay on her reaction. HER REACTION: [A few visible actions, in time order.] She stays seated; the camera does not move.
Feeling words stay in the comparison because they matched visible actions on recognition. Neutral controls stay too: the opening clip earned its place. We still need to compare backstory plus visible actions against visible actions alone for rage and sadness.
Next, we’ll repeat promising prompts with more reviewers, another face and setup, and different editing contexts.
Back to the kitchen
None of the 42 takes felt fully real to me. The closest earned “almost real.” And one of my favourite contempt performances still came from the control: a woman, a kitchen, a line, and whatever I read into her reaction.
Getting the emotion right was only the first test. Simple emotion words matched visible-actions prompts on recognition, but visible actions produced some of the performances I believed most. That is why we’re building on them.