Open-weight workflow guide · Updated 4 August 2026

MiniMax H3 Workflow: Text, Image, and Reference-to-Video

Build a practical MiniMax H3 workflow in ComfyUI, choose the correct T2V, I2V, or R2V template, structure synchronized visuals and sound, and review the result before you spend time on a larger render.

Use only media you own or may process. Check the current MiniMax H3 licence, model files, hardware needs, and ComfyUI version before planning production work.

The short answer

A MiniMax H3 workflow combines the model, text encoder, video and audio VAEs, prompt, references, resolution, duration, sampling, and export nodes. ComfyUI currently ships official starting templates for text-to-video, image-to-video, and reference-to-video. Choose the mode from the information you already own, then test at preview resolution before increasing render cost.

What changed

Why MiniMax H3 needs an audiovisual workflow

MiniMax positions H3 as a general-purpose multimodal generation model rather than a collection of isolated video tools. Text, images, footage, and audio can describe a shared context, while the target result can include synchronized voice, effects, ambience, and music.

That changes prompt design. A useful H3 brief does more than describe appearance. It explains what happens over time, how the camera observes it, what every reference is allowed to control, who speaks, which sounds belong inside the scene, and whether the audience should hear separate background music.

The model launch advertises up to 15 seconds, 24 fps, native stereo sound, and output up to 2K. Those are provider claims, not universal results from every local checkpoint. The ComfyUI templates document their own native resolution path, duration grid, and weight variants, so verify the file that a real workflow exports.

Choose the graph

MiniMax H3 workflow modes compared

Pick one mode before downloading large files. T2V and I2V share the fl2va model family; R2V uses separate ref2va weights. Mixing the wrong diffusion model with a template can lead to errors or a workflow that does not behave as documented.

ModeInputWeightsBest starting useTemplate
Text-to-video (T2V)A written audiovisual brief with scene, action, camera, dialogue, effects, and music.fl2vaNew concepts, establishing shots, product ideas, and testing prompt interpretation.Open JSON
Image-to-video (I2V)An authorised first frame, last frame, or first-and-last-frame pair plus instructions.fl2vaAnimating a designed frame while preserving composition, colour, subject, and layout.Open JSON
Reference-to-video (R2V)Tagged image, video, and audio references with a clear job assigned to every asset.ref2vaIdentity, style, motion, camera, or voice continuity across a controlled target shot.Open JSON

Local setup

How to build the workflow

The template library can download the expected files, but you should still understand what the graph loads. Record file names and versions so a successful render remains reproducible after updates.

View ComfyUI model files
  1. 1

    Update ComfyUI

    Use ComfyUI 0.30.0 or later so the native MiniMax H3 nodes and current template definitions are available.

  2. 2

    Choose one workflow mode

    Start with T2V, I2V, or R2V according to the information you already have. Avoid loading references that do not serve a defined purpose.

  3. 3

    Download the matching files

    Use fl2va weights for the official T2V and I2V templates, or ref2va weights for R2V, plus the text encoder and both video and audio VAEs.

  4. 4

    Set a preview resolution

    Select the aspect ratio, use a modest megapixel target for initial tests, and keep dimensions on the 32-pixel resolution grid.

  5. 5

    Write the audiovisual timeline

    Describe visible action, camera movement, dialogue, physical sound, ambience, and non-diegetic music in the order they should occur.

  6. 6

    Generate, review, then scale

    Inspect motion, identity, audio sync, text, rights, and artefacts before increasing resolution or duration. Save the workflow and exact model names.

T2V: write the timeline

Start with overall style and composition, then number later shots and give cuts increasing times. Describe camera motion as an action, not a pile of labels.

I2V: protect the anchors

State whether the image is the first frame, last frame, or both. Explain the continuous motion path and preserve identity, colour, objects, and spatial relationships.

R2V: assign every reference

Use tags in connection order and say which asset controls identity, style, motion, camera, or voice. Explicit jobs reduce accidental copying between references.

Preview before a full-quality render

The official ComfyUI guide recommends a smaller preview size and documents roughly 1344×768 for a one-megapixel 16:9 native canvas. Higher resolution and longer duration increase memory use and generation time, but they do not guarantee a usable result.

First test prompt adherence, subject stability, camera movement, speech, sound timing, and visual artefacts. Only then increase output size. ComfyUI also documents optional Sage Attention acceleration, but installation and benefit depend on compatible CUDA, PyTorch, tensor types, nodes, and hardware; keep the standard graph as a reproducible baseline.

Copy and adapt

Five MiniMax H3 workflow prompts

These original examples follow the official separation between integrated multimodal description, overall soundscape, and non-diegetic music. Replace fictional details with authorised assets, then align timestamps to the effective duration shown by your graph.

T2V product reveal

integrated_multimodal_description: [Shot 1] Live-action cinematic product film. A fictional compact silver radio marked AURORA rests on black stone. The camera pushes in slowly as a warm tuning light switches on. [Shot 2] At 00:04.000, the camera cuts to a macro side angle while one adult hand turns the dial; keep the product shape, controls, and word AURORA consistent. overall_soundscape: Quiet studio room tone, a precise dial click, and a soft analogue tuning sweep. non_diegetic_music: Restrained low electronic pulse, no copyrighted melody.

I2V character moment

For the target video, at 0.00 seconds into the target video, <Picture 1> is fully referenced. integrated_multimodal_description: [Shot 1] Preserve the fictional adult character's face, green coat, framing, rainy station, and colour palette. The character breathes, looks towards an arriving light, and takes one natural step as the camera tracks right slowly. overall_soundscape: Light rain, distant rail vibration, coat fabric movement. non_diegetic_music: None.

First-and-last-frame transition

Use <Picture 1> as the exact opening and <Picture 2> as the exact ending. integrated_multimodal_description: [Shot 1] A continuous single shot moves from the original empty ceramic desk to the finished fictional vase in the final frame. Clay rises gradually under two adult hands; preserve table position, camera angle, lighting direction, and vase colour. The camera remains static and the last pose settles naturally. overall_soundscape: Spinning wheel, wet clay, quiet workshop ambience. non_diegetic_music: Soft original marimba notes.

R2V identity and camera

Reference assignments: <Picture 1> defines the fictional adult subject's identity and clothing only. <Video 1> defines camera rhythm and body timing only. <Audio 1> defines the original voice tone only. integrated_multimodal_description: [Shot 1] Place the subject in a fictional glass observatory at sunrise. Preserve identity from <Picture 1>, follow the slow orbit from <Video 1>, and have the subject say, ‘The signal is finally clear,’ using the voice quality of <Audio 1>. Replace all source locations, props, logos, and background details. overall_soundscape: Low ventilation and distant wind. non_diegetic_music: None.

Vertical social workflow

integrated_multimodal_description: [Shot 1] 9:16 tactile food film for a fictional bakery. Flour falls across a wooden bench in a tight overhead shot. [Shot 2] At 00:02.500, cut to adult hands folding dough, with natural contact and consistent tools. [Shot 3] At 00:05.500, reveal three finished loaves beside a bright window. No real brand, platform interface, captions, or watermark. overall_soundscape: Flour brush, dough movement, gentle oven fan. non_diegetic_music: Simple original acoustic guitar texture.

Compact prompt formula

Reference alignment + integrated timeline + overall soundscape + non-diegetic music + exclusions. Use the official base prompt guide for T2V and keyframes, or the official full-reference guide for R2V.

Quality control

Review the result like production media

Open weights and a successful node graph do not remove editorial responsibility. Check each result at normal playback speed, frame by frame where needed, and with headphones. Preserve the prompt, graph, seed, file names, settings, source permissions, and any post-production changes.

  • Subject and costume consistency
  • Hands, contact, motion, and physics
  • Camera direction and cut timing
  • Dialogue wording and lip movement
  • Stereo placement and audio sync
  • Rendered text and fictional brand marks
  • Source-media rights and consent
  • Export dimensions, fps, codec, and duration

Frequently asked questions

What is a MiniMax H3 workflow?+

A MiniMax H3 workflow is a connected generation graph that loads the model, text encoder, video and audio VAEs, prompt, references, resolution, duration, sampler, and export nodes. In ComfyUI, official templates provide working starting graphs for text-to-video, image-to-video, and reference-to-video.

Which MiniMax H3 workflow should I start with?+

Start with T2V when the idea exists mainly as text, I2V when one or two designed frames should anchor the result, and R2V when identity, style, motion, camera, or voice must come from authorised references. More references are not automatically better; assign each one a specific job.

Is MiniMax H3 available as open weights?+

Yes. MiniMax publishes the original model repository, and Comfy Org provides repackaged files for native ComfyUI workflows. The repositories use the MiniMax H3 Community License Agreement, so review the exact licence before commercial use or redistribution.

Can MiniMax H3 generate audio with video?+

The official model and ComfyUI documentation describe native stereo dialogue, sound effects, ambience, and music generated with the visual result. Treat generated speech and music as material requiring editorial, safety, consent, and rights review before publishing.

Can a ComfyUI MiniMax H3 workflow export 2K?+

MiniMax advertises output up to 2K, but the official ComfyUI templates document a native 768-pixel short edge, with roughly 1344×768 at a one-megapixel 16:9 setting. Do not assume every local template or checkpoint produces the marketed 2K path; verify the actual exported dimensions and workflow stage.

How long can a MiniMax H3 video be?+

The launch material describes clips up to about 15 seconds at 24 fps. ComfyUI notes that duration snaps to a 17k+5 frame grid. Available length, memory use, generation time, and stability still depend on the exact workflow, model variant, resolution, and hardware.

What files does the local workflow need?+

The official templates use a fl2va or ref2va diffusion model, a Qwen3-VL-based H3 text encoder, a video VAE, and an audio VAE. T2V and I2V use fl2va; R2V uses ref2va. Follow current model-card paths rather than renaming files or mixing variants from unrelated templates.

How many references can R2V use?+

The current ComfyUI guide documents up to nine reference images, three reference videos, and three standalone audio references. It recommends tagging them in connection order and explicitly stating whether each controls identity, style, motion, camera, or voice.

Official sources and scope

This independent workflow guide was checked on 4 August 2026 against the MiniMax launch article, official model repository and prompt guides, and ComfyUI documentation and templates. Provider capability and performance statements are labelled rather than presented as independent benchmarks. Hardware, checkpoints, licences, nodes, and interfaces can change, so verify the current files before downloading a large model.

Compare before committing a production pipeline

Need a hosted alternative to the local H3 workflow?

Read the dedicated MiniMax H3 vs LTX comparison, or continue to LTX.dev to explore browser-based text, image, audio, and video creation workflows. Next Vibe AI is independent and is not affiliated with MiniMax, ComfyUI, Lightricks, or LTX.

Explore LTX.dev