MiniMax's H3 video model reads a structured, field-labeled prompt rather than a loose paragraph β it separates what's on screen (integrated_multimodal_description / detailed_description) from ambient sound (overall_soundscape) and score (non_diegetic_music), and in full-reference mode it tracks every image, video and audio input through explicit <Subject N> / <Picture N> / <Video N> / <Audio N> labels. This page covers the five input modes, the exact field structure for each, two real prompts pulled from a working H3 creator on X, and the paired h3-prompt-writing Claude skill that writes these for you.
H3 takes one of five inputs. The first four share one prompt structure (the base modes); the fifth β full-reference β uses a longer six-section rewrite format because it has to track multiple separate source assets at once.
| Mode | Input | What the prompt does |
|---|---|---|
| T2VA | Text only | Builds the complete audiovisual timeline from nothing but the text. |
| I2VA | First frame | Anchors <Picture 1> as the actual 0.00s frame, then develops forward from it. |
| FL2VA | First + last frame | Describes the continuous path connecting the two given frames β usually one single shot. |
| L2VA | Last frame only | Infers a plausible opening, then converges onto the given last frame by the final shot. |
| Ref2VA | Any mix of images / videos / audio | Full-reference mode β every asset gets a tracked label and a six-section rewrite (see Β§6). |
The four base modes always use the same three core fields, just with a different opening instruction line β see Β§2.
Every base-mode prompt has two parts: an instruction line (skipped entirely for T2VA), then a blank line, then the three core fields in this fixed order.
N is the actual shot index; S.SS is the target duration to exactly two decimal places.
T2VA β no instruction line at all, prompt starts directly at "integrated_multimodal_description:" I2VA: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. FL2VA: How the reference pictures align with the target video β Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video. L2VA: How the reference pictures align with the target video β <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
integrated_multimodal_description: [Shot 1] ... overall_soundscape: ... non_diegetic_music: ...
| Mode | Recommended arc |
|---|---|
| I2VA | first-frame anchor β action onset β continuous development β result or reaction |
| FL2VA | first-frame state β observable intermediate changes β progressively narrowing differences β last-frame state |
| L2VA | plausible preceding state β explicit action/transition β gradual convergence in the final shot β last-frame landing |
FL2VA generally favors a single shot so the model can interpolate continuously between the two given frames β only split into multiple shots when explicitly asked for. In L2VA, the reference picture belongs to the last shot, not the first β don't describe Shot 1 as if it already matches the image.
No timestamp on [Shot 1]. Every later shot starts with a strictly increasing cut time inside the video's duration:
[Shot 2] At 00:03.500, the camera cuts to...
Use the camera cuts to / the shot cuts to / transitions to / changes to / switches to for an ordinary cut. Cross-dissolve, fade or wipe only when the user explicitly asks for one. A cut should introduce genuinely new information (subject, space, state, viewpoint, time) β if only distance or a slight angle needs to change, use camera motion instead of a cut.
Any banner, sign, label, subtitle, or neon text actually visible in the shot goes in English double quotes, original wording and punctuation preserved verbatim, never translated:
A red neon sign reading "θ₯δΈδΈ" glows above the doorway.
A complete camera move has up to three parts β motion type (required), amplitude, and speed. Only add amplitude/speed when they matter; medium amplitude and normal speed are usually left unstated. Write the move as a natural sentence inside the shot, not as stacked labels at the end.
| Dimension | Expression | Meaning |
|---|---|---|
| Motion | Zoom In / Zoom Out | Focal length changes, camera body stays put |
| Motion | Push In / Pull Out | Camera physically moves forward / backward |
| Motion | Pan Left / Pan Right | Camera stays put, lens pivots horizontally |
| Motion | Truck Left / Truck Right | Camera translates horizontally |
| Motion | Tilt Up / Tilt Down | Camera stays put, lens pivots vertically |
| Motion | Pedestal Up / Pedestal Down | Whole camera moves up / down |
| Motion | Arc Shot | Camera moves in an arc around the subject |
| Motion | Tracking Shot | Camera follows a moving subject |
| Motion | Static Shot | Position and lens both stay still |
| Motion | Shake Slightly / Shake Strongly | Slight / strong camera shake |
| Motion | POV | Subject's own point of view |
| Motion | Roll Clockwise / Roll Counterclockwise | Camera rolls around the lens axis |
| Amplitude | with small amplitude / with large amplitude | Range of compositional change |
| Speed | at slow speed / at fast speed | Pacing of the move |
The camera pushes in with small amplitude at slow speed toward the folded letter in her hands. The camera pans right with large amplitude at fast speed, revealing the open doorway. The camera holds a static shot as the runner exits the frame.
(S1), (S2). Two speakers together use a compound ID: (S1,S2). Silent characters get no ID at all.<d>. Inside <d>, only the language tag and the exact spoken content β preserve every word and punctuation mark verbatim, never translate or rewrite it.says in an off-screen voiceover, and the very next sentence must confirm the on-screen character's lips stay closed.<scenetrans> at both connecting points plus a continuity phrase (continues seamlessly across the cut, etc.); speech cut off by the video ending uses <cutoff>.The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d> The two children (S1,S2) shout together, <d>[English] Wait for us!</d> The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.
1β4 English sentences, one continuous paragraph, covering ambient sound, physical action sound, and non-verbal human sound across the whole video (wind, rain, traffic, footsteps, impacts, breathing, laughterβ¦). Dialogue, singing and diegetic music already live in the description β don't repeat them here. Use N/A only if the user explicitly wants total silence.
overall_soundscape: Steady rain taps against the cafΓ© windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.
1β3 English sentences on instrumentation, tempo, rhythm and dynamic changes only β no abstract mood words, no explaining what the score is "meant to feel like." Anything the characters themselves can hear (radio, a phone, live singing) is diegetic and belongs in the description instead. Use N/A when there's no score.
non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.
Use this whenever more than one image/video/audio asset needs to be tracked through the video, or when the roles are richer than a plain keyframe (character sheets, style references, an edited source video, a reused voice or soundtrack). Six sections, always in this order:
| Section | Purpose |
|---|---|
subject_definitions | Defines every referenced item and its label |
summary | One paragraph: task type + target video + main reference relationships |
retention_analysis | How each referenced item is preserved / transferred / reused |
detailed_description | Shot-by-shot visuals, actions, sound and dialogue β the main body |
overall_soundscape | Same rules as Β§5 |
non_diegetic_music | Same rules as Β§5 |
| Label | Meaning |
|---|---|
<Subject N> | Reusable visible content β a person, animal, object, scene, costume, prop, style, action or pose. One subject can be defined by multiple assets, and one asset can define multiple subjects. |
<Picture N> | A reference image used as a concrete frame or shot-planning anchor (first frame, last frame, storyboard). If an image only defines a character/scene/style, cite it inside that <Subject N> instead of giving it its own picture entry. |
<Video N> | Whole-video relationships only β editing a source video, continuing from its end, or referencing its camera/cut/rhythm structure. A person or object taken from the video still gets its own <Subject N>. |
<Audio N> | A standalone audio asset or a reference video's audio track β copied signal, voice-timbre reference, reused dialogue/lyrics/SFX, or referenced beat/rhythm. |
Once a label is assigned, it means the same thing everywhere in the rewrite β subject_definitions, summary, retention_analysis, detailed_description, and the audio sections all reuse it unchanged. Don't invent new reference labels once you reach summary β they're all defined up front.
Starts with a bracketed task type (combine with + when more than one applies, never repeat a type):
| Task type | When |
|---|---|
keyframe completion | An image is a concrete frame anchor (first/last/keyframe) |
reference generation | An asset guides a character/scene/style/action/camera/storyboard without being a concrete frame or edited source |
video editing | An existing source video is directly modified |
video continuation | New content continues/extends/resumes from a source video |
audio reuse | The same audio signal is reused in full or in part |
audio reference | Only style/timbre/content/beat is referenced, not the signal itself |
| Visible content | Meaning |
|---|---|
fully_preserved | The defined role is fully kept |
partially_preserved | Still used, but some defined traits changed or only partly retained |
attribute_transfer | Traits transferred onto a different identifiable subject |
weak_reference | Only broad style/category/composition/atmosphere similarity kept |
| Audio | Meaning |
|---|---|
fully_copy | The complete source audio is the video's complete final track |
partially_copy | Only part copied, or layers added/removed/replaced after copying |
reference | Timbre/rhythm/style/content/texture referenced, signal not copied |
weak_reference | Only broad category/atmosphere similarity kept |
detailed_description follows the same shot/camera/speaker/dialogue rules as the base modes (Β§2β4), plus: open with one or two sentences naming the overall style before [Shot 1], and weave in reference labels at each item's first clear appearance rather than front-loading them all. Generation-task descriptions typically run 350β500 English words β dialogue-dense scenes prioritize fitting the spoken timeline over hitting that count.
One per mode, straight from the reference guide.
No reference asset at all β the full timeline is built directly from text.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot. overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced. non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease. overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands. non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
How the reference pictures align with the target video β Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot. overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes. non_diegetic_music: N/A
How the reference pictures align with the target video β <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video. integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>. overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor. non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.
Full six-section rewrite: two picture references define the location and the dog, two video references define the two human subjects, and one audio reference guides a voice timbre without copying the original line.
subject_definitions: <Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table. <Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail. <Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves. <Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings. <Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer. summary: [reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh. retention_analysis: <Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained. <Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained. <Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained. <Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained. <Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal. detailed_description: The target video uses a realistic multi-camera sitcom style with warm indoor lighting. [Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back. [Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur. [Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame. overall_soundscape: Soft indoor coffee-shop room tone continues throughout the scene. non_diegetic_music: N/A
Two working H3 prompts shared publicly by KΕda (@aimikoda), a creative technologist who posts AI video prompts and workflows. Reposted here for reference with attribution β these are their prompts, not this page's.
Posted as a reusable template: swap in your own @[character reference], keep the rest. Good study piece for how much physical/kinetic detail H3 accepts in one continuous block β dance beats, VFX color language, and camera moves are all interleaved in plain prose rather than split into numbered shots.
15-second experimental anime fashion film. Use @[character reference] for the character's exact identity, outfit and rendering style. A pristine, uniform lemon-yellow cyclorama with a seamless matte floor. Flat, clean yellow throughout, free of grain, gradients and set dressing. 16:9. No text or logos. The character performs one fast, unconventional dance phrase: ankles weave while the torso stays suspended, a wrist flick ripples through both elbows into a chest pop, knees fold into a low corkscrew pivot, then a heel swivel unwinds the body into an off-balance editorial pose. Continue through asymmetric shoulder rolls, angular arm threading around the face and a sideways glide with the head counter-rotating toward the lens. Every sharp lock immediately releases into another surprising movement. Convincing weight, precise foot contact and expressive follow-through. Electric cyan and hot pink dominate the motion effects. Bold, opaque silhouette afterimages peel away from opposite sides of the body, replay the previous gesture a fraction late and slam back into alignment on each pose. Arm sweeps stretch them into broad curved ribbons; crossing steps weave pink and cyan trails around the ankles; rapid turns fan out overlapping screen-print impressions. At the strongest accents, oversized cyan and pink cutouts of the current pose slide across the foreground, briefly masking the lens. Keep the real character crisp and recognizable, with clean yellow visible between effect bursts. The camera dances in counterpoint: skim backward at ankle height through advancing footwork, whip upward around an outstretched wrist into an extreme face close-up, then use the passing forearm to cut into a rotating overhead view of the low pivot. Dive to a tilted low-angle full-body composition as the character rises. Orbit opposite the sideways glide, exaggerating hands and feet near the wide lens. Accelerate between compositions and brake hard on pose locks. Finish with a sudden forward lean toward the lens while enormous pink and cyan echoes continue leaning farther, sweeping past both sides of the camera. Snap wide: the character lands a final asymmetric pose as the colored echoes collapse precisely into the silhouette. Hard cut to black. 170 BPM broken beat, syncopated accents, fluid animation and razor-fast camera transitions. No slow motion or prolonged holds.
Source: x.com/aimikoda/status/2097001707683119410 β "MiniMax H3 Prompt Share," Sep 2026.
Six Midjourney v8.2 images, each locked to one subject or prop by index ("Image #3 defines the young rescuerβ¦"), then a full ten-shot dark-fantasy trailer with two lines of accented dialogue and a title card. Worth studying for the reference-mapping opener and for how tightly each shot is timed (0-1s, 1-2sβ¦) with an explicit cut motivation on every beat. Posted alongside the note: "if you give H3 high-quality references, you get high-quality results."
Image #3 defines the young rescuer and forbidden red fruit; Image #2 defines his captive beloved; Image #1 defines the garden keeper and his potted flowers; Image #4 defines the hooded hunter; Image #6 defines the seer and her raven; Image #5 defines the stone god. Preserve their separate identities, costumes and distinctive props. Use the foliage and roses as connected areas of one nocturnal garden, with new cinematic viewpoints. Preserve the references' illustrated contours, painterly surfaces, deep ultramarine shadows, violet skin tones and coral-red accents. A kinetic, artistic 15-second dark-fantasy trailer for THE STONE GARDEN. A young man steals forbidden fruit to free his beloved, awakening the garden's stone god. Ten cinematic animated shots followed by one title card. A rapid opening montage expands into readable two-second action and revelation beats. Animate bodies, fabric and foliage with physical depth while retaining the painted style. 0-1s: Close-up through blue leaves, a brief push toward the captive's frightened face. She reaches toward the rescuer offscreen left and whispers, "Please." Her voice is breathy, with a distinct upper-class British English accent. One fragile string note. Cut on her reaching gesture. 1-2s: Extreme close-up of her outstretched wrist. A living blue vine coils tight and jerks her hand back toward the foliage. Its leaves scrape her sleeve. Hard cut on the sudden recoil. 2-3s: Macro side view of the forbidden red fruit cradled in the rescuer's hands, still attached to a low red-leafed branch. He twists and pulls; show the stem snapping and the fruit coming free. The amplified snap kills the music. Cut on the break. 3-4s: Low medium shot tracking beside the rescuer as he clutches that detached fruit to his chest and bursts through hanging red branches toward the captive's clearing, screen right. Branches whip behind him. Cut as he clears the foliage. 4-5s: Overhead insert into the keeper's pot, securely supported by both hands. Its purple flowers abruptly wilt, their stems buckling inward as a magical consequence of the theft. A dry floral crackle. Cut on the collapse. 5-6s: Tight frontal close-up of the keeper's sunglasses and clenched jaw. He snaps his head toward the thief's route; a brief lateral camera slide catches the rose wall reflected in his glasses. A low bass pulse begins. Hard reaction cut. 6-8s: Low wide lateral tracking shot along a narrow passage through the rose field. The hooded hunter launches into a sprint toward screen right, following the rescuer's route. Roses streak past the foreground; his patterned coat trails behind him. Heavy percussion erupts with his first running strides. Cut during the pursuit, before he catches up. 8-10s: Three-quarter medium close-up of the seer. Her raven pushes off her raised hand as she looks toward the disturbance. A quick upward tilt follows its wings; her feathered head ornament stays attached. In a low, ominous voice with a pronounced aristocratic British English accent, precise consonants and non-rhotic delivery, she says, "Save her..." Cut at 10s, carrying her voice across the edit. 10-12s: Tight side-on rescue shot: the young man is left, the captive right, separated by vines. He reaches the same stolen fruit through a gap. She closes her hand over his on the fruit; at contact, the vines loosen and uncoil from her wrist. A short push keeps their joining hands and the releasing vine visible together. He still supports the fruit. The seer's offscreen warning completes, "...and you wake him." Cut on the release. 12-14s: Low-angle wide shot reveals the monumental stone god above the clearing, with the reunited pair small beneath it. A forceful push toward the god accompanies its red eyes flaring and its heavy carved head turning down toward them for the first time. It remains stone. Grinding rock overwhelms the score. Smash cut to black at 14s before it attacks. 14-15s: Hold a static title card for the full final second: THE STONE GARDEN, exact spelling, large centered coral-red uppercase serif lettering on deep ultramarine-black. The complete title appears immediately, with a single resonant impact and a short decaying tail. No other text. Sharp cuts, urgent natural-speed action, readable hands and reactions. Preserve the rescuer's rightward travel, the hunter's ongoing pursuit and the stolen fruit's identity. The lovers remain together holding the fruit in the final reveal. Keep music beneath the dialogue. Both spoken lines are English with deliberate British accents. No subtitles, additional narration or slow motion. The final film title is the only onscreen text.
Source: x.com/aimikoda/status/2097112285827268887 (video) and the prompt reply β "MiniMax H3 + Midjourney v8.2," Sep 2026.
subject_definitions / summary / retention_analysis schema from Β§6 β H3 clearly accepts both. The schema exists to keep multi-asset reference tracking unambiguous at scale; a shorter reference set (six images here) reads fine as plain numbered "Image #N definesβ¦" mapping followed by a shot list, which is easier to write by hand.<d>; only the language tag and literal words go inside it β never paraphrase or translate what's inside the tag.The h3-prompt-writing Agent Skill encodes this entire page β install it once and Claude writes correctly structured H3 prompts (any of the five modes) without needing to be re-briefed on the field names, label rules or timing notation every session.
.skill file is a zip. Unpack it into .claude/skills/ in your project (or ~/.claude/skills/ to have it everywhere) so you end up with skills/h3-prompt-writing/SKILL.md plus its references/ folder. It loads on the next session..skill file wherever Skills are managed in your settings.
Prompt structure, field rules and worked examples come from the h3-prompt-writing skill's own reference guides (base modes + full-reference mode).
β Found prompts: KΕda (@aimikoda) on X β anime fashion film, "The Stone Garden".