Official structure explained · Updated 4 August 2026

MiniMax H3 Prompt Guide for Better Audiovisual Video

Write a MiniMax H3 prompt that directs shots, motion, dialogue, sound, music, and references as one timeline. This independent guide turns the official MiniMax format into a practical method with original examples you can copy and adapt.

MiniMax H3 can generate visual and audio material together. Review identity, speech, music, rights, and safety before publishing any MiniMax H3 result.

The MiniMax H3 prompt formula

For a base MiniMax H3 prompt, write an optional reference-alignment instruction, then integrated_multimodal_description, overall_soundscape, and non_diegetic_music. The first field holds the visible MiniMax H3 timeline, dialogue, and sounds occurring inside the scene. The second summarizes ambience and physical sound. The third describes background scoring heard only by the audience.

21-point writing audit

Scan these checks after the first draft. Each line should be easy to confirm from the finished audiovisual brief.

  • MiniMax H3 prompt: establish the location.
  • MiniMax H3 prompt: define the main subject.
  • MiniMax H3 prompt: describe one clear action.
  • MiniMax H3 prompt: direct the camera naturally.
  • MiniMax H3 prompt: time every later cut.
  • MiniMax H3 prompt: name each speaking character.
  • MiniMax H3 prompt: preserve supplied dialogue exactly.
  • MiniMax H3 prompt: separate physical scene sound.
  • MiniMax H3 prompt: describe audience-only music.
  • MiniMax H3 prompt: quote visible text exactly.
  • MiniMax H3 prompt: assign every reference role.
  • MiniMax H3 prompt: protect keyframe anchors.
  • MiniMax H3 prompt: state the final condition.
  • MiniMax H3 prompt: remove conflicting detail.
  • MiniMax H3 prompt: review the MiniMax exported result.
  • MiniMax H3 prompt: use authorised H3 references.
  • MiniMax H3 prompt: preview before H3 scaling.
  • MiniMax H3: change one variable at a time.
  • MiniMax H3: record every production setting.
  • MiniMax H3: verify every rendered word.
  • MiniMax H3: review rights and consent.

Start with intent

What makes a useful MiniMax prompt?

MiniMax H3 is designed around a shared visual and audio context. A useful prompt is therefore closer to a compact production brief than a list of visual adjectives. Tell MiniMax H3 what the opening composition contains, what changes over time, how the camera observes that change, who speaks, what sound belongs inside the scene, and whether the audience hears separate music.

MiniMax documentation recommends stating the whole scene before breaking a longer idea into timed shots. Each cut should reveal new information. When only distance or angle changes, a camera move is usually clearer than adding another cut to the prompt.

Treat every reference as a constraint with one defined job. MiniMax cannot know whether an image controls identity, colour, framing, style, or only the opening pose unless the MiniMax H3 instructions say so. Explicit scope also helps avoid copying scenery or logos that were never meant to transfer.

Match the input

MiniMax H3 prompt modes compared

MiniMax uses different opening instructions for text, first-frame, first-and-last-frame, and last-frame generation. Choose the MiniMax H3 mode before drafting the body so the prompt describes the correct direction of travel.

T2VA

Begin directly with the three shared fields.

Build the complete visual and audio timeline from text. MiniMax can infer supporting detail, but every addition should preserve the creator's intent.

I2VA

Place the first-frame alignment instruction before the three fields.

Anchor style, identity, composition, colour, objects, and spatial relationships to Picture 1, then describe continuous forward motion.

FL2VA

Map Picture 1 to 0.00 seconds and Picture 2 to the effective end time.

Describe an observable path between the two frames. MiniMax generally benefits from a continuous single shot unless cuts are required.

L2VA

Map Picture 1 to the effective final time, not to Shot 1 by default.

Infer a plausible earlier state and make action, camera, lighting, and composition converge on the supplied last frame.

Direct playback

Write shots, cuts, and camera motion

Begin with [Shot 1] and no timestamp. Give each later shot a strictly increasing time inside the actual MiniMax H3 duration. MiniMax expects camera language to read as a natural action, not as disconnected tags at the end of the prompt.

A complete move can name the motion type, amplitude, and speed. Omit medium amplitude and normal speed when they add nothing. This keeps MiniMax instructions precise without forcing MiniMax H3 to resolve redundant modifiers.

NeedClear wordingWhy it helps
Slow approachThe camera pushes in with small amplitude at slow speed.Defines physical movement and pacing.
Horizontal revealThe camera pans right with large amplitude at fast speed.Separates rotation from camera translation.
Follow actionA tracking shot follows the cyclist through the gate.Connects the lens to subject movement.
No movementThe camera holds a static shot as the steam clears.Stops MiniMax from inventing an unwanted move.
New shot[Shot 2] At 00:04.500, the camera cuts to a close-up.Introduces a real edit at a valid time.

Audio direction

Separate dialogue, scene sound, and music

MiniMax H3 generates stereo audio with the video, so sound should not be an afterthought. Give MiniMax H3 one home for each audible element and avoid repeating the same event in multiple fields.

Dialogue and singing

Keep them inside the integrated timeline. Assign stable IDs such as (S1), preserve the supplied words inside the <d> block, and identify off-screen voiceover explicitly.

Overall soundscape

Summarize ambience, footsteps, impacts, fabric, weather, breathing, and other physical sounds in one compact paragraph. Do not duplicate spoken lines here.

Non-diegetic music

Describe instrumentation, tempo, rhythm, and dynamic change. Use N/A when no background score is wanted. Music audible to characters belongs in the scene timeline instead.

Reference-to-video

Assign every MiniMax reference one job

In R2V, connect and name each authorised asset in order: <Picture 1>, <Video 1>, <Audio 1>, and so on. Then tell MiniMax what each source controls. An image may define identity, a video may define motion or camera rhythm, and an audio file may define voice. MiniMax H3 should not be asked to infer these roles from file order alone.

The full-reference MiniMax guide also distinguishes reusable subjects from source assets. A person, object, scene, action, or effect reused in the target can be tracked as a subject, while a video label identifies the temporal or editing source. Consistent labels make a complex MiniMax H3 prompt easier to review.

Never use a reference merely because it is available. Remove assets without a clear purpose, replace real logos and locations when they are outside the creative brief, and confirm permission before any MiniMax generation.

Reference assignment pattern

<Picture 1> defines the fictional adult subject's identity and clothing only.

<Video 1> defines body timing and the tracking camera only.

<Audio 1> defines the authorised voice timbre only.

State what transfers and what must be replaced. Narrow assignments give MiniMax a more testable target and reduce accidental carry-over into the MiniMax H3 output.

Copy and adapt

Six original MiniMax H3 prompt examples

Each example is fictional and follows the official MiniMax field separation. Replace subjects, references, dialogue, timing, and aspect ratio with authorised project details. A copied MiniMax H3 prompt is a starting point, not a guarantee of an identical result.

Text-to-video

T2VA — night market dialogue

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a fictional night market after rain. A middle-aged tea seller with a low, measured voice (S1) lifts a copper kettle while the camera trucks left with small amplitude at slow speed. Steam catches the red lantern light. The seller says: <d>[English] The last cup is always the quietest.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of tea filling a blue ceramic cup as the final word carries over from the previous shot. overall_soundscape: Water drips from canvas awnings over low market ambience. The kettle lid clicks, liquid pours, and a bicycle bell passes in the distance. non_diegetic_music: Sparse plucked strings at a slow tempo, fading beneath the last pour.

Image-to-video

I2VA — portrait comes alive

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Preserve the fictional adult subject's face, silver jacket, seated pose, window framing, and cool blue lighting from <Picture 1>. The subject takes one quiet breath and turns toward a warm light outside while the camera pushes in with small amplitude at slow speed. No new jewellery, text, logo, or person appears. overall_soundscape: A low ventilation hum, light jacket movement, and distant city traffic. non_diegetic_music: N/A

First and last frame

FL2VA — first-to-last product move

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video. integrated_multimodal_description: [Shot 1] Live-action studio film. Preserve the fictional amber bottle, stone platform, black background, and left-side key light established by Picture 1. The camera arcs right with small amplitude at slow speed as a narrow ribbon of mist crosses behind the bottle. The bottle rotates steadily, the label remains readable as ‘NORTH’, and the lighting warms gradually until object angle, mist position, framing, and highlights match Picture 2 exactly. overall_soundscape: Subtle glass resonance and a soft mechanical turntable. non_diegetic_music: A restrained original electronic pulse that rises once and ends cleanly.

Reference-to-video

R2V — identity, motion, and voice

Reference assignments: <Picture 1> defines the fictional adult subject's identity and green coat only. <Video 1> defines walking rhythm and the slow rear tracking camera only. <Audio 1> defines the authorised original voice timbre only. integrated_multimodal_description: [Shot 1] Place the subject in a new glass greenhouse at dawn. Preserve identity and clothing from <Picture 1>, follow the movement and camera relationship from <Video 1>, and replace all source scenery, props, signage, and logos. The subject walks between wet plants and says: <d>[English] We begin again here.</d> using the voice quality of <Audio 1>. overall_soundscape: Light rain on glass, footsteps on stone, leaves brushing the coat. non_diegetic_music: N/A

9:16 social video

Vertical creator clip

integrated_multimodal_description: [Shot 1] 9:16 live-action craft film for a fictional studio. A top-down close shot shows adult hands folding cream paper on a clean oak desk. The camera remains static while every crease is pressed firmly. [Shot 2] At 00:03.500, the shot cuts to a low macro angle as the paper opens into a geometric lamp shade and warm light switches on inside it. Keep the same hands, paper colour, tools, and desk grain. No platform interface, watermark, real logo, or subtitle. overall_soundscape: Crisp paper folds, fingertip movement, one wooden tool tap, and quiet room tone. non_diegetic_music: Three original marimba notes followed by a soft sustained pad.

Last-frame video

L2VA — land on the final frame

How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video. integrated_multimodal_description: [Shot 1] Live-action, a close shot begins with loose blue tiles and a blank wall under neutral daylight. Two adult hands place the tiles one by one while the camera pushes in with small amplitude at slow speed. The arrangement narrows toward the supplied design; dust is brushed away, the final tile settles, and the hands move into the exact positions shown in <Picture 1>. At the end, tile spacing, wall texture, shadows, framing, and colour match the final frame. overall_soundscape: Ceramic clicks, light scraping, a dry brush, and distant workshop ambience. non_diegetic_music: Soft original piano notes at a moderate tempo, ending on the final placement.

Troubleshooting

Fix the brief before adding more words

When MiniMax misses an instruction, the solution is often subtraction. MiniMax H3 must resolve every subject, action, cut, camera move, reference, spoken line, and sound inside a short duration. Reduce competing goals, preserve the strongest constraint, and run another preview with one controlled change.

Identity drifts

Repeat stable appearance anchors, narrow each reference role, and remove conflicting style sources.

Motion feels random

Use one observable action path with a clear beginning, middle, and final state.

Camera is unstable

Choose one camera action per shot and state static framing when no movement is wanted.

Audio is crowded

Keep dialogue in the timeline, physical sounds in the soundscape, and scoring in the music field.

Cuts arrive too early

Check that every later timestamp increases and remains inside the effective MiniMax H3 duration.

Final frame is missed

Describe progressive convergence instead of repeating a static description of the target image.

Final review

MiniMax H3 prompt checklist

Read the complete MiniMax brief once as a timeline and once as a rights document. Then render a smaller MiniMax H3 preview before increasing resolution or duration. Keep the prompt, seed, model file, graph, references, and edit notes with the approved output.

Open the MiniMax H3 workflow guide
  • The correct MiniMax mode is selected
  • Every MiniMax H3 reference has one explicit role
  • Shot times increase within duration
  • Camera wording is physical and unambiguous
  • Speaker IDs remain stable
  • Sound and music use separate fields
  • Visible text is quoted and reviewed
  • Source rights and consent are recorded

Frequently asked questions

What is the best MiniMax H3 prompt structure?+

For the base modes, MiniMax documents three shared fields in this order: integrated_multimodal_description, overall_soundscape, and non_diegetic_music. I2VA, FL2VA, and L2VA add an image-alignment instruction before those fields. The best structure is the one that matches the actual inputs and effective duration in your H3 workflow.

Should a MiniMax H3 prompt include timestamps?+

Use no timestamp on Shot 1. For each later cut, use a strictly increasing time inside the effective video duration, such as ‘[Shot 2] At 00:03.500’. Camera motion within one shot does not require a new timestamp unless a real cut or transition occurs.

How should dialogue be written for H3?+

Give each speaking or singing character a stable speaker ID such as (S1). Put the identifying description, action, and delivery outside the dialogue tag, and keep only the language marker and exact spoken words inside <d>. Preserve supplied dialogue without rewriting it.

What is the difference between soundscape and music?+

The soundscape covers ambience, physical action sounds, and non-verbal human sounds that belong to the world of the scene. Non-diegetic music is background scoring heard by the audience but not the characters. Music from a radio, phone, or visible performer belongs in the multimodal timeline instead.

How do I reference an image, video, or audio file?+

For R2V, tag assets in their connection order, such as <Picture 1>, <Video 1>, and <Audio 1>, then assign each asset one explicit purpose: identity, style, motion, camera, voice, or another narrow role. Use only references you own or are authorised to process.

Can MiniMax H3 preserve on-screen text?+

The official writing guide says visible banners, labels, signs, subtitles, and neon text should be placed in English double quotation marks and preserved verbatim. Generated text still needs visual review; do not assume spelling, brand treatment, or legal clearance is automatic.

How long should a MiniMax H3 prompt be?+

There is no useful universal word count. Write enough to establish the initial frame, action path, camera, sound, and ending without repeating adjectives or creating contradictory instructions. A short single-shot brief may need less text than a multi-shot audiovisual sequence.

Why does my H3 result ignore part of the prompt?+

The brief may contain too many competing subjects, cuts, references, camera moves, or sound instructions for the selected duration. Simplify the target, remove decorative adjectives, assign reference roles explicitly, and test one change at a time at preview resolution.

Official sources and editorial scope

This independent MiniMax H3 prompt guide was checked on 4 August 2026 against the official MiniMax base-mode writing guide, full-reference format guide, model repository, launch material, and ComfyUI documentation. MiniMax capability statements are treated as provider documentation rather than independent performance benchmarks. Interfaces and guidance may change, so verify the current H3 files before production use.

Continue from writing to creation

Want to turn the brief into an original video?

Compare MiniMax H3 and LTX, or visit LTX.dev to explore a hosted creative workflow for text, image, audio, and video. Next Vibe AI is independent and is not affiliated with MiniMax, Lightricks, or LTX.

Explore LTX.dev