🧬 MiniMax H3 Prompt Guide

MiniMax's H3 video model reads a structured, field-labeled prompt rather than a loose paragraph β€” it separates what's on screen (integrated_multimodal_description / detailed_description) from ambient sound (overall_soundscape) and score (non_diegetic_music), and in full-reference mode it tracks every image, video and audio input through explicit <Subject N> / <Picture N> / <Video N> / <Audio N> labels. This page covers the five input modes, the exact field structure for each, two real prompts pulled from a working H3 creator on X, and the paired h3-prompt-writing Claude skill that writes these for you.

1. Overview & the five modes

H3 takes one of five inputs. The first four share one prompt structure (the base modes); the fifth β€” full-reference β€” uses a longer six-section rewrite format because it has to track multiple separate source assets at once.

ModeInputWhat the prompt does
T2VAText onlyBuilds the complete audiovisual timeline from nothing but the text.
I2VAFirst frameAnchors <Picture 1> as the actual 0.00s frame, then develops forward from it.
FL2VAFirst + last frameDescribes the continuous path connecting the two given frames β€” usually one single shot.
L2VALast frame onlyInfers a plausible opening, then converges onto the given last frame by the final shot.
Ref2VAAny mix of images / videos / audioFull-reference mode β€” every asset gets a tracked label and a six-section rewrite (see Β§6).

The four base modes always use the same three core fields, just with a different opening instruction line β€” see Β§2.

2. Base prompt structure (T2VA / I2VA / FL2VA / L2VA)

Every base-mode prompt has two parts: an instruction line (skipped entirely for T2VA), then a blank line, then the three core fields in this fixed order.

Part one β€” the instruction line

N is the actual shot index; S.SS is the target duration to exactly two decimal places.

T2VA β€” no instruction line at all, prompt starts directly at "integrated_multimodal_description:"

I2VA:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

FL2VA:
How the reference pictures align with the target video β€” Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.

L2VA:
How the reference pictures align with the target video β€” <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.

Part two β€” the three core fields

integrated_multimodal_description: [Shot 1] ...

overall_soundscape: ...

non_diegetic_music: ...
  • integrated_multimodal_description β€” visuals, actions, shots, speakers, dialogue, singing, and diegetic (in-scene) audio, laid out along the timeline.
  • overall_soundscape β€” ambience, physical action sound, and non-verbal human sound across the whole video. Never repeats dialogue or music.
  • non_diegetic_music β€” score the characters can't hear, only the audience. Instrumentation/tempo/dynamics only β€” no mood words.

How each keyframe mode develops the description

ModeRecommended arc
I2VAfirst-frame anchor β†’ action onset β†’ continuous development β†’ result or reaction
FL2VAfirst-frame state β†’ observable intermediate changes β†’ progressively narrowing differences β†’ last-frame state
L2VAplausible preceding state β†’ explicit action/transition β†’ gradual convergence in the final shot β†’ last-frame landing

FL2VA generally favors a single shot so the model can interpolate continuously between the two given frames β€” only split into multiple shots when explicitly asked for. In L2VA, the reference picture belongs to the last shot, not the first β€” don't describe Shot 1 as if it already matches the image.

Shots and cuts

No timestamp on [Shot 1]. Every later shot starts with a strictly increasing cut time inside the video's duration:

[Shot 2] At 00:03.500, the camera cuts to...

Use the camera cuts to / the shot cuts to / transitions to / changes to / switches to for an ordinary cut. Cross-dissolve, fade or wipe only when the user explicitly asks for one. A cut should introduce genuinely new information (subject, space, state, viewpoint, time) β€” if only distance or a slight angle needs to change, use camera motion instead of a cut.

On-screen text

Any banner, sign, label, subtitle, or neon text actually visible in the shot goes in English double quotes, original wording and punctuation preserved verbatim, never translated:

A red neon sign reading "θ₯业中" glows above the doorway.

3. Camera language

A complete camera move has up to three parts β€” motion type (required), amplitude, and speed. Only add amplitude/speed when they matter; medium amplitude and normal speed are usually left unstated. Write the move as a natural sentence inside the shot, not as stacked labels at the end.

DimensionExpressionMeaning
MotionZoom In / Zoom OutFocal length changes, camera body stays put
MotionPush In / Pull OutCamera physically moves forward / backward
MotionPan Left / Pan RightCamera stays put, lens pivots horizontally
MotionTruck Left / Truck RightCamera translates horizontally
MotionTilt Up / Tilt DownCamera stays put, lens pivots vertically
MotionPedestal Up / Pedestal DownWhole camera moves up / down
MotionArc ShotCamera moves in an arc around the subject
MotionTracking ShotCamera follows a moving subject
MotionStatic ShotPosition and lens both stay still
MotionShake Slightly / Shake StronglySlight / strong camera shake
MotionPOVSubject's own point of view
MotionRoll Clockwise / Roll CounterclockwiseCamera rolls around the lens axis
Amplitudewith small amplitude / with large amplitudeRange of compositional change
Speedat slow speed / at fast speedPacing of the move
The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The camera pans right with large amplitude at fast speed, revealing the open doorway.
The camera holds a static shot as the runner exits the frame.

4. Speakers & dialogue

  • Anyone who speaks, sings, or produces an off-screen voice gets a stable ID β€” (S1), (S2). Two speakers together use a compound ID: (S1,S2). Silent characters get no ID at all.
  • On a speaker's first appearance, establish identity from what's visible/audible: character type, age, gender, on/off-screen, pitch, timbre, speaking rate, accent.
  • Keep the speaker's identity phrase, ID, action and delivery outside <d>. Inside <d>, only the language tag and the exact spoken content β€” preserve every word and punctuation mark verbatim, never translate or rewrite it.
  • Voiceover uses the exact phrase says in an off-screen voiceover, and the very next sentence must confirm the on-screen character's lips stay closed.
  • Dialogue crossing a cut uses <scenetrans> at both connecting points plus a continuity phrase (continues seamlessly across the cut, etc.); speech cut off by the video ending uses <cutoff>.
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>

The two children (S1,S2) shout together, <d>[English] Wait for us!</d>

The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.

5. overall_soundscape & non_diegetic_music

overall_soundscape

1–4 English sentences, one continuous paragraph, covering ambient sound, physical action sound, and non-verbal human sound across the whole video (wind, rain, traffic, footsteps, impacts, breathing, laughter…). Dialogue, singing and diegetic music already live in the description β€” don't repeat them here. Use N/A only if the user explicitly wants total silence.

overall_soundscape: Steady rain taps against the cafΓ© windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.

non_diegetic_music

1–3 English sentences on instrumentation, tempo, rhythm and dynamic changes only β€” no abstract mood words, no explaining what the score is "meant to feel like." Anything the characters themselves can hear (radio, a phone, live singing) is diegetic and belongs in the description instead. Use N/A when there's no score.

non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.

6. Full-reference mode (Ref2VA)

Use this whenever more than one image/video/audio asset needs to be tracked through the video, or when the roles are richer than a plain keyframe (character sheets, style references, an edited source video, a reused voice or soundtrack). Six sections, always in this order:

SectionPurpose
subject_definitionsDefines every referenced item and its label
summaryOne paragraph: task type + target video + main reference relationships
retention_analysisHow each referenced item is preserved / transferred / reused
detailed_descriptionShot-by-shot visuals, actions, sound and dialogue β€” the main body
overall_soundscapeSame rules as Β§5
non_diegetic_musicSame rules as Β§5

The four reference labels

LabelMeaning
<Subject N>Reusable visible content β€” a person, animal, object, scene, costume, prop, style, action or pose. One subject can be defined by multiple assets, and one asset can define multiple subjects.
<Picture N>A reference image used as a concrete frame or shot-planning anchor (first frame, last frame, storyboard). If an image only defines a character/scene/style, cite it inside that <Subject N> instead of giving it its own picture entry.
<Video N>Whole-video relationships only β€” editing a source video, continuing from its end, or referencing its camera/cut/rhythm structure. A person or object taken from the video still gets its own <Subject N>.
<Audio N>A standalone audio asset or a reference video's audio track β€” copied signal, voice-timbre reference, reused dialogue/lyrics/SFX, or referenced beat/rhythm.

Once a label is assigned, it means the same thing everywhere in the rewrite β€” subject_definitions, summary, retention_analysis, detailed_description, and the audio sections all reuse it unchanged. Don't invent new reference labels once you reach summary β€” they're all defined up front.

summary task-type prefix

Starts with a bracketed task type (combine with + when more than one applies, never repeat a type):

Task typeWhen
keyframe completionAn image is a concrete frame anchor (first/last/keyframe)
reference generationAn asset guides a character/scene/style/action/camera/storyboard without being a concrete frame or edited source
video editingAn existing source video is directly modified
video continuationNew content continues/extends/resumes from a source video
audio reuseThe same audio signal is reused in full or in part
audio referenceOnly style/timbre/content/beat is referenced, not the signal itself

retention_analysis relationship markers

Visible contentMeaning
fully_preservedThe defined role is fully kept
partially_preservedStill used, but some defined traits changed or only partly retained
attribute_transferTraits transferred onto a different identifiable subject
weak_referenceOnly broad style/category/composition/atmosphere similarity kept
AudioMeaning
fully_copyThe complete source audio is the video's complete final track
partially_copyOnly part copied, or layers added/removed/replaced after copying
referenceTimbre/rhythm/style/content/texture referenced, signal not copied
weak_referenceOnly broad category/atmosphere similarity kept

detailed_description follows the same shot/camera/speaker/dialogue rules as the base modes (Β§2–4), plus: open with one or two sentences naming the overall style before [Shot 1], and weave in reference labels at each item's first clear appearance rather than front-loading them all. Generation-task descriptions typically run 350–500 English words β€” dialogue-dense scenes prioritize fitting the spoken timeline over hitting that count.

7. Worked examples

One per mode, straight from the reference guide.

T2VA β€” bakery, no reference image

No reference asset at all β€” the full timeline is built directly from text.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.

overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.

non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.
I2VA β€” train window, develops forward from Picture 1
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.

overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.

non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
FL2VA β€” umbrella opening, single shot bridging two frames
How the reference pictures align with the target video β€” Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.

overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.

non_diegetic_music: N/A
L2VA β€” breaking glass, converges onto the last frame
How the reference pictures align with the target video β€” <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.

overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.

non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.
Ref2VA β€” coffee shop, three subjects + a voice-timbre reference

Full six-section rewrite: two picture references define the location and the dog, two video references define the two human subjects, and one audio reference guides a voice timbre without copying the original line.

subject_definitions:
<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.

summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.

detailed_description:
The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.

overall_soundscape:
Soft indoor coffee-shop room tone continues throughout the scene.

non_diegetic_music:
N/A

8. Found on X β€” real H3 prompts

Two working H3 prompts shared publicly by Kōda (@aimikoda), a creative technologist who posts AI video prompts and workflows. Reposted here for reference with attribution β€” these are their prompts, not this page's.

Anime fashion film β€” dance moves, T2VA-style with a character reference

Posted as a reusable template: swap in your own @[character reference], keep the rest. Good study piece for how much physical/kinetic detail H3 accepts in one continuous block β€” dance beats, VFX color language, and camera moves are all interleaved in plain prose rather than split into numbered shots.

Copied!
15-second experimental anime fashion film. Use @[character reference] for the character's exact identity, outfit and rendering style. A pristine, uniform lemon-yellow cyclorama with a seamless matte floor. Flat, clean yellow throughout, free of grain, gradients and set dressing. 16:9. No text or logos.

The character performs one fast, unconventional dance phrase: ankles weave while the torso stays suspended, a wrist flick ripples through both elbows into a chest pop, knees fold into a low corkscrew pivot, then a heel swivel unwinds the body into an off-balance editorial pose. Continue through asymmetric shoulder rolls, angular arm threading around the face and a sideways glide with the head counter-rotating toward the lens. Every sharp lock immediately releases into another surprising movement. Convincing weight, precise foot contact and expressive follow-through.

Electric cyan and hot pink dominate the motion effects. Bold, opaque silhouette afterimages peel away from opposite sides of the body, replay the previous gesture a fraction late and slam back into alignment on each pose. Arm sweeps stretch them into broad curved ribbons; crossing steps weave pink and cyan trails around the ankles; rapid turns fan out overlapping screen-print impressions. At the strongest accents, oversized cyan and pink cutouts of the current pose slide across the foreground, briefly masking the lens. Keep the real character crisp and recognizable, with clean yellow visible between effect bursts.

The camera dances in counterpoint: skim backward at ankle height through advancing footwork, whip upward around an outstretched wrist into an extreme face close-up, then use the passing forearm to cut into a rotating overhead view of the low pivot. Dive to a tilted low-angle full-body composition as the character rises. Orbit opposite the sideways glide, exaggerating hands and feet near the wide lens. Accelerate between compositions and brake hard on pose locks.

Finish with a sudden forward lean toward the lens while enormous pink and cyan echoes continue leaning farther, sweeping past both sides of the camera. Snap wide: the character lands a final asymmetric pose as the colored echoes collapse precisely into the silhouette. Hard cut to black.

170 BPM broken beat, syncopated accents, fluid animation and razor-fast camera transitions. No slow motion or prolonged holds.

Source: x.com/aimikoda/status/2097001707683119410 β€” "MiniMax H3 Prompt Share," Sep 2026.

"The Stone Garden" β€” six-image Ref2VA trailer with dialogue

Six Midjourney v8.2 images, each locked to one subject or prop by index ("Image #3 defines the young rescuer…"), then a full ten-shot dark-fantasy trailer with two lines of accented dialogue and a title card. Worth studying for the reference-mapping opener and for how tightly each shot is timed (0-1s, 1-2s…) with an explicit cut motivation on every beat. Posted alongside the note: "if you give H3 high-quality references, you get high-quality results."

Copied!
Image #3 defines the young rescuer and forbidden red fruit; Image #2 defines his captive beloved; Image #1 defines the garden keeper and his potted flowers; Image #4 defines the hooded hunter; Image #6 defines the seer and her raven; Image #5 defines the stone god. Preserve their separate identities, costumes and distinctive props. Use the foliage and roses as connected areas of one nocturnal garden, with new cinematic viewpoints. Preserve the references' illustrated contours, painterly surfaces, deep ultramarine shadows, violet skin tones and coral-red accents.

A kinetic, artistic 15-second dark-fantasy trailer for THE STONE GARDEN. A young man steals forbidden fruit to free his beloved, awakening the garden's stone god. Ten cinematic animated shots followed by one title card. A rapid opening montage expands into readable two-second action and revelation beats. Animate bodies, fabric and foliage with physical depth while retaining the painted style.

0-1s: Close-up through blue leaves, a brief push toward the captive's frightened face. She reaches toward the rescuer offscreen left and whispers, "Please." Her voice is breathy, with a distinct upper-class British English accent. One fragile string note. Cut on her reaching gesture.

1-2s: Extreme close-up of her outstretched wrist. A living blue vine coils tight and jerks her hand back toward the foliage. Its leaves scrape her sleeve. Hard cut on the sudden recoil.

2-3s: Macro side view of the forbidden red fruit cradled in the rescuer's hands, still attached to a low red-leafed branch. He twists and pulls; show the stem snapping and the fruit coming free. The amplified snap kills the music. Cut on the break.

3-4s: Low medium shot tracking beside the rescuer as he clutches that detached fruit to his chest and bursts through hanging red branches toward the captive's clearing, screen right. Branches whip behind him. Cut as he clears the foliage.

4-5s: Overhead insert into the keeper's pot, securely supported by both hands. Its purple flowers abruptly wilt, their stems buckling inward as a magical consequence of the theft. A dry floral crackle. Cut on the collapse.

5-6s: Tight frontal close-up of the keeper's sunglasses and clenched jaw. He snaps his head toward the thief's route; a brief lateral camera slide catches the rose wall reflected in his glasses. A low bass pulse begins. Hard reaction cut.

6-8s: Low wide lateral tracking shot along a narrow passage through the rose field. The hooded hunter launches into a sprint toward screen right, following the rescuer's route. Roses streak past the foreground; his patterned coat trails behind him. Heavy percussion erupts with his first running strides. Cut during the pursuit, before he catches up.

8-10s: Three-quarter medium close-up of the seer. Her raven pushes off her raised hand as she looks toward the disturbance. A quick upward tilt follows its wings; her feathered head ornament stays attached. In a low, ominous voice with a pronounced aristocratic British English accent, precise consonants and non-rhotic delivery, she says, "Save her..." Cut at 10s, carrying her voice across the edit.

10-12s: Tight side-on rescue shot: the young man is left, the captive right, separated by vines. He reaches the same stolen fruit through a gap. She closes her hand over his on the fruit; at contact, the vines loosen and uncoil from her wrist. A short push keeps their joining hands and the releasing vine visible together. He still supports the fruit. The seer's offscreen warning completes, "...and you wake him." Cut on the release.

12-14s: Low-angle wide shot reveals the monumental stone god above the clearing, with the reunited pair small beneath it. A forceful push toward the god accompanies its red eyes flaring and its heavy carved head turning down toward them for the first time. It remains stone. Grinding rock overwhelms the score. Smash cut to black at 14s before it attacks.

14-15s: Hold a static title card for the full final second: THE STONE GARDEN, exact spelling, large centered coral-red uppercase serif lettering on deep ultramarine-black. The complete title appears immediately, with a single resonant impact and a short decaying tail. No other text.

Sharp cuts, urgent natural-speed action, readable hands and reactions. Preserve the rescuer's rightward travel, the hunter's ongoing pursuit and the stolen fruit's identity. The lovers remain together holding the fruit in the final reveal. Keep music beneath the dialogue. Both spoken lines are English with deliberate British accents. No subtitles, additional narration or slow motion. The final film title is the only onscreen text.

Source: x.com/aimikoda/status/2097112285827268887 (video) and the prompt reply β€” "MiniMax H3 + Midjourney v8.2," Sep 2026.

Note: this prompt is written in Kōda's own plain-English shot-list style, not the strict labeled subject_definitions / summary / retention_analysis schema from Β§6 β€” H3 clearly accepts both. The schema exists to keep multi-asset reference tracking unambiguous at scale; a shorter reference set (six images here) reads fine as plain numbered "Image #N defines…" mapping followed by a shot list, which is easier to write by hand.

9. Best practices

  • Reference quality sets the ceiling. Kōda's own framing for the six-image trailer: give H3 high-quality references and you get high-quality results. Low-res, inconsistent, or ambiguous reference images cost you more downstream than a slightly weaker text description does.
  • Describe what's visible or audible, never plot. Both base and full-reference modes explicitly warn against reducing the description to a plot summary or a list of reference relationships β€” every sentence should map to something the model can actually render: a composition, an action, a sound, a line of dialogue.
  • Map references before you write the shots. The Stone Garden prompt spends its first paragraph purely on "Image #N defines X" before a single shot is described β€” resolving identity/costume/prop assignment up front means the shot list never has to re-explain who's who.
  • Pin down what must NOT drift. Illustrated contour style, shadow color, skin tone, accent, character identity β€” name the specific traits that have to survive every cut, the way both examples pin exact British-accent delivery and painterly rendering across all ten shots.
  • One cut = one new piece of information. If only the camera distance or angle needs to change, that's camera motion, not a cut β€” reserve cuts for a genuine change in subject, space, state, viewpoint or time.
  • Give every cut a reason. Both real-world examples state why each cut happens ("Cut on the break," "Hard cut on the sudden recoil," "Cut during the pursuit, before he catches up") β€” this reads as directing, not just listing shots.
  • Keep dialogue exact and separated from delivery. Identity, action and vocal quality go outside <d>; only the language tag and literal words go inside it β€” never paraphrase or translate what's inside the tag.
  • overall_soundscape and non_diegetic_music are not a dumping ground. If a sound is tied to a specific visible moment (a snap, a footstep, a line reading), it usually belongs inline in the shot description instead β€” the two summary fields are for what plays underneath the whole video.
  • State negative constraints plainly. "No text or logos," "No slow motion or prolonged holds," "No subtitles, additional narration" β€” both found prompts close with explicit exclusions rather than assuming the model will infer them.
  • Match your mode to what you actually have. One image you want as the opening frame β†’ I2VA. Two images, opening and closing β†’ FL2VA, one shot, don't over-describe either static frame β€” describe the path. Several unrelated reference assets with different jobs (character, location, voice, edited source video) β†’ Ref2VA, so nothing's identity gets lost mid-prompt.

10. Use this with Claude

The h3-prompt-writing Agent Skill encodes this entire page β€” install it once and Claude writes correctly structured H3 prompts (any of the five modes) without needing to be re-briefed on the field names, label rules or timing notation every session.

⬇ Download h3-prompt-writing.skill ~15 KB Β· SKILL.md + 2 reference files
  • Claude Code: a .skill file is a zip. Unpack it into .claude/skills/ in your project (or ~/.claude/skills/ to have it everywhere) so you end up with skills/h3-prompt-writing/SKILL.md plus its references/ folder. It loads on the next session.
  • Claude apps: upload the .skill file wherever Skills are managed in your settings.
  • Once installed, describe what you want β€” "write me an H3 prompt for…", "rewrite this as a full-reference H3 prompt with these three images", "port this to L2VA" β€” and it triggers on its own.
  • It's portable beyond Claude too: no external API calls or MiniMax-specific tooling required, just local reference files, so it also runs under agents that can read local files (an OpenAI/Codex metadata file ships alongside for that case).

11. Sources

Prompt structure, field rules and worked examples come from the h3-prompt-writing skill's own reference guides (base modes + full-reference mode).
β†’ Found prompts: Kōda (@aimikoda) on X β€” anime fashion film, "The Stone Garden".