MiniMax H3 Max

The world's fastest video model. Describe what happens, add a photo if you like. A 5-second video takes about 5 seconds to generate.

The photo is optional: add one and the video starts from that frame, skip it and the scene you describe is generated. The video is generated without sound.

Reference photos

Up to 5 photos

One photo is enough; you can add up to 5. Show what should stay the same in the video: a person's face, a product's shape, how a place looks.

Audio reference

Up to 3 clips · 15 s

Audio can't be sent on its own. Add at least one reference photo first.

Video length

5seconds

75 stars

5 s15 s

Quality

Format

To start, write what should happen in the video. Adding a photo is optional.

Your video will appear here. H3 Max generates a 5-second video in about 5 seconds with the generation time shown live.

Need an idea?

What is MiniMax H3?

MiniMax H3 is a general-purpose multimodal generation model. It understands text, images, video and audio in a single unified context, and generates video at 2K resolution, up to 15 seconds, with native stereo sound.

Early tests show H3 is ready for commercial content production across a wide range of uses: following instructions to the letter, rendering text and brand elements correctly, and transferring motion from video to video. With precise, controllable multimodal generation and editing, it was built for advertising, brand work, e-commerce, product design, interface design and games.

With its Contextual Omni Representation, H3-VAE, H3-Omni Transformer and In-Context Regeneration technologies it delivers industry-leading price/performance. 2K resolution comes as the default: at 2K, H3's per-second price is less than a third of mainstream models, and at 768p less than half of the same models' 720p price.

2K performance

No separate upscaling module: H3 regenerates its own low-resolution output in context, bringing fine text and small detail back.

Native stereo sound

Sound isn't added afterwards; it's generated together with the video. Speech, effects and music are modelled under one head, not as separate domains.

Multimodal context

Real creative work means blending complex information across modalities: images, audio, video and more as input. Describe the relationship between the context and the target video in words, and H3 handles the complex full-modality understanding on its own.

“Use the Hitchcock camera move from video 1 as the reference, have the character in image 2 sing, and match the vocals to audio 3.”

Where it's being used

Film titles

Multi-shot opening

Product website

Interface and 3D scene

Motion poster

Text stays intact

Ads and e-commerce

Product shoot

The description and sample videos are taken from MiniMax's H3 announcement: minimax.io/blog/minimax-h3

Çerezler

Sana daha iyi bir deneyim sunmak için çerez kullanıyoruz. Dilediğin zaman değiştirebilirsin.