T2VA
Begin directly with the three shared fields.
Build the complete visual and audio timeline from text. MiniMax can infer supporting detail, but every addition should preserve the creator's intent.
Write a MiniMax H3 prompt that directs shots, motion, dialogue, sound, music, and references as one timeline. This independent guide turns the official MiniMax format into a practical method with original examples you can copy and adapt.
MiniMax H3 can generate visual and audio material together. Review identity, speech, music, rights, and safety before publishing any MiniMax H3 result.
For a base MiniMax H3 prompt, write an optional reference-alignment instruction, then integrated_multimodal_description, overall_soundscape, and non_diegetic_music. The first field holds the visible MiniMax H3 timeline, dialogue, and sounds occurring inside the scene. The second summarizes ambience and physical sound. The third describes background scoring heard only by the audience.
Scan these checks after the first draft. Each line should be easy to confirm from the finished audiovisual brief.
Start with intent
MiniMax H3 is designed around a shared visual and audio context. A useful prompt is therefore closer to a compact production brief than a list of visual adjectives. Tell MiniMax H3 what the opening composition contains, what changes over time, how the camera observes that change, who speaks, what sound belongs inside the scene, and whether the audience hears separate music.
MiniMax documentation recommends stating the whole scene before breaking a longer idea into timed shots. Each cut should reveal new information. When only distance or angle changes, a camera move is usually clearer than adding another cut to the prompt.
Treat every reference as a constraint with one defined job. MiniMax cannot know whether an image controls identity, colour, framing, style, or only the opening pose unless the MiniMax H3 instructions say so. Explicit scope also helps avoid copying scenery or logos that were never meant to transfer.
Match the input
MiniMax uses different opening instructions for text, first-frame, first-and-last-frame, and last-frame generation. Choose the MiniMax H3 mode before drafting the body so the prompt describes the correct direction of travel.
Begin directly with the three shared fields.
Build the complete visual and audio timeline from text. MiniMax can infer supporting detail, but every addition should preserve the creator's intent.
Place the first-frame alignment instruction before the three fields.
Anchor style, identity, composition, colour, objects, and spatial relationships to Picture 1, then describe continuous forward motion.
Map Picture 1 to 0.00 seconds and Picture 2 to the effective end time.
Describe an observable path between the two frames. MiniMax generally benefits from a continuous single shot unless cuts are required.
Map Picture 1 to the effective final time, not to Shot 1 by default.
Infer a plausible earlier state and make action, camera, lighting, and composition converge on the supplied last frame.
Direct playback
Begin with [Shot 1] and no timestamp. Give each later shot a strictly increasing time inside the actual MiniMax H3 duration. MiniMax expects camera language to read as a natural action, not as disconnected tags at the end of the prompt.
A complete move can name the motion type, amplitude, and speed. Omit medium amplitude and normal speed when they add nothing. This keeps MiniMax instructions precise without forcing MiniMax H3 to resolve redundant modifiers.
| Need | Clear wording | Why it helps |
|---|---|---|
| Slow approach | The camera pushes in with small amplitude at slow speed. | Defines physical movement and pacing. |
| Horizontal reveal | The camera pans right with large amplitude at fast speed. | Separates rotation from camera translation. |
| Follow action | A tracking shot follows the cyclist through the gate. | Connects the lens to subject movement. |
| No movement | The camera holds a static shot as the steam clears. | Stops MiniMax from inventing an unwanted move. |
| New shot | [Shot 2] At 00:04.500, the camera cuts to a close-up. | Introduces a real edit at a valid time. |
Audio direction
MiniMax H3 generates stereo audio with the video, so sound should not be an afterthought. Give MiniMax H3 one home for each audible element and avoid repeating the same event in multiple fields.
Keep them inside the integrated timeline. Assign stable IDs such as (S1), preserve the supplied words inside the <d> block, and identify off-screen voiceover explicitly.
Summarize ambience, footsteps, impacts, fabric, weather, breathing, and other physical sounds in one compact paragraph. Do not duplicate spoken lines here.
Describe instrumentation, tempo, rhythm, and dynamic change. Use N/A when no background score is wanted. Music audible to characters belongs in the scene timeline instead.
Reference-to-video
In R2V, connect and name each authorised asset in order: <Picture 1>, <Video 1>, <Audio 1>, and so on. Then tell MiniMax what each source controls. An image may define identity, a video may define motion or camera rhythm, and an audio file may define voice. MiniMax H3 should not be asked to infer these roles from file order alone.
The full-reference MiniMax guide also distinguishes reusable subjects from source assets. A person, object, scene, action, or effect reused in the target can be tracked as a subject, while a video label identifies the temporal or editing source. Consistent labels make a complex MiniMax H3 prompt easier to review.
Never use a reference merely because it is available. Remove assets without a clear purpose, replace real logos and locations when they are outside the creative brief, and confirm permission before any MiniMax generation.
<Picture 1> defines the fictional adult subject's identity and clothing only.
<Video 1> defines body timing and the tracking camera only.
<Audio 1> defines the authorised voice timbre only.
State what transfers and what must be replaced. Narrow assignments give MiniMax a more testable target and reduce accidental carry-over into the MiniMax H3 output.
Copy and adapt
Each example is fictional and follows the official MiniMax field separation. Replace subjects, references, dialogue, timing, and aspect ratio with authorised project details. A copied MiniMax H3 prompt is a starting point, not a guarantee of an identical result.
Text-to-video
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a fictional night market after rain. A middle-aged tea seller with a low, measured voice (S1) lifts a copper kettle while the camera trucks left with small amplitude at slow speed. Steam catches the red lantern light. The seller says: <d>[English] The last cup is always the quietest.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of tea filling a blue ceramic cup as the final word carries over from the previous shot. overall_soundscape: Water drips from canvas awnings over low market ambience. The kettle lid clicks, liquid pours, and a bicycle bell passes in the distance. non_diegetic_music: Sparse plucked strings at a slow tempo, fading beneath the last pour.
Image-to-video
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Preserve the fictional adult subject's face, silver jacket, seated pose, window framing, and cool blue lighting from <Picture 1>. The subject takes one quiet breath and turns toward a warm light outside while the camera pushes in with small amplitude at slow speed. No new jewellery, text, logo, or person appears. overall_soundscape: A low ventilation hum, light jacket movement, and distant city traffic. non_diegetic_music: N/A
First and last frame
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video. integrated_multimodal_description: [Shot 1] Live-action studio film. Preserve the fictional amber bottle, stone platform, black background, and left-side key light established by Picture 1. The camera arcs right with small amplitude at slow speed as a narrow ribbon of mist crosses behind the bottle. The bottle rotates steadily, the label remains readable as ‘NORTH’, and the lighting warms gradually until object angle, mist position, framing, and highlights match Picture 2 exactly. overall_soundscape: Subtle glass resonance and a soft mechanical turntable. non_diegetic_music: A restrained original electronic pulse that rises once and ends cleanly.
Reference-to-video
Reference assignments: <Picture 1> defines the fictional adult subject's identity and green coat only. <Video 1> defines walking rhythm and the slow rear tracking camera only. <Audio 1> defines the authorised original voice timbre only. integrated_multimodal_description: [Shot 1] Place the subject in a new glass greenhouse at dawn. Preserve identity and clothing from <Picture 1>, follow the movement and camera relationship from <Video 1>, and replace all source scenery, props, signage, and logos. The subject walks between wet plants and says: <d>[English] We begin again here.</d> using the voice quality of <Audio 1>. overall_soundscape: Light rain on glass, footsteps on stone, leaves brushing the coat. non_diegetic_music: N/A
9:16 social video
integrated_multimodal_description: [Shot 1] 9:16 live-action craft film for a fictional studio. A top-down close shot shows adult hands folding cream paper on a clean oak desk. The camera remains static while every crease is pressed firmly. [Shot 2] At 00:03.500, the shot cuts to a low macro angle as the paper opens into a geometric lamp shade and warm light switches on inside it. Keep the same hands, paper colour, tools, and desk grain. No platform interface, watermark, real logo, or subtitle. overall_soundscape: Crisp paper folds, fingertip movement, one wooden tool tap, and quiet room tone. non_diegetic_music: Three original marimba notes followed by a soft sustained pad.
Last-frame video
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video. integrated_multimodal_description: [Shot 1] Live-action, a close shot begins with loose blue tiles and a blank wall under neutral daylight. Two adult hands place the tiles one by one while the camera pushes in with small amplitude at slow speed. The arrangement narrows toward the supplied design; dust is brushed away, the final tile settles, and the hands move into the exact positions shown in <Picture 1>. At the end, tile spacing, wall texture, shadows, framing, and colour match the final frame. overall_soundscape: Ceramic clicks, light scraping, a dry brush, and distant workshop ambience. non_diegetic_music: Soft original piano notes at a moderate tempo, ending on the final placement.
Troubleshooting
When MiniMax misses an instruction, the solution is often subtraction. MiniMax H3 must resolve every subject, action, cut, camera move, reference, spoken line, and sound inside a short duration. Reduce competing goals, preserve the strongest constraint, and run another preview with one controlled change.
Repeat stable appearance anchors, narrow each reference role, and remove conflicting style sources.
Use one observable action path with a clear beginning, middle, and final state.
Choose one camera action per shot and state static framing when no movement is wanted.
Keep dialogue in the timeline, physical sounds in the soundscape, and scoring in the music field.
Check that every later timestamp increases and remains inside the effective MiniMax H3 duration.
Describe progressive convergence instead of repeating a static description of the target image.
Final review
Read the complete MiniMax brief once as a timeline and once as a rights document. Then render a smaller MiniMax H3 preview before increasing resolution or duration. Keep the prompt, seed, model file, graph, references, and edit notes with the approved output.
Open the MiniMax H3 workflow guideFor the base modes, MiniMax documents three shared fields in this order: integrated_multimodal_description, overall_soundscape, and non_diegetic_music. I2VA, FL2VA, and L2VA add an image-alignment instruction before those fields. The best structure is the one that matches the actual inputs and effective duration in your H3 workflow.
Use no timestamp on Shot 1. For each later cut, use a strictly increasing time inside the effective video duration, such as ‘[Shot 2] At 00:03.500’. Camera motion within one shot does not require a new timestamp unless a real cut or transition occurs.
Give each speaking or singing character a stable speaker ID such as (S1). Put the identifying description, action, and delivery outside the dialogue tag, and keep only the language marker and exact spoken words inside <d>. Preserve supplied dialogue without rewriting it.
The soundscape covers ambience, physical action sounds, and non-verbal human sounds that belong to the world of the scene. Non-diegetic music is background scoring heard by the audience but not the characters. Music from a radio, phone, or visible performer belongs in the multimodal timeline instead.
For R2V, tag assets in their connection order, such as <Picture 1>, <Video 1>, and <Audio 1>, then assign each asset one explicit purpose: identity, style, motion, camera, voice, or another narrow role. Use only references you own or are authorised to process.
The official writing guide says visible banners, labels, signs, subtitles, and neon text should be placed in English double quotation marks and preserved verbatim. Generated text still needs visual review; do not assume spelling, brand treatment, or legal clearance is automatic.
There is no useful universal word count. Write enough to establish the initial frame, action path, camera, sound, and ending without repeating adjectives or creating contradictory instructions. A short single-shot brief may need less text than a multi-shot audiovisual sequence.
The brief may contain too many competing subjects, cuts, references, camera moves, or sound instructions for the selected duration. Simplify the target, remove decorative adjectives, assign reference roles explicitly, and test one change at a time at preview resolution.
This independent MiniMax H3 prompt guide was checked on 4 August 2026 against the official MiniMax base-mode writing guide, full-reference format guide, model repository, launch material, and ComfyUI documentation. MiniMax capability statements are treated as provider documentation rather than independent performance benchmarks. Interfaces and guidance may change, so verify the current H3 files before production use.
Compare MiniMax H3 and LTX, or visit LTX.dev to explore a hosted creative workflow for text, image, audio, and video. Next Vibe AI is independent and is not affiliated with MiniMax, Lightricks, or LTX.