Explore MiniMax's video generation models—from prompt-led Hailuo workflows to H3, a full-modal system designed to understand text, images, video and sound together.
MiniMax overview
A model family spanning focused video tools and multimodal creative systems.
MiniMax develops generative models across language, speech, music, image and video. Its video line began with Hailuo models focused on text-to-video and image-to-video creation, then expanded toward richer motion, camera control and audio-visual production.
H3 marks a broader shift: it is designed to understand a multimodal creative context and generate a coordinated audio-video result. Exact inputs and limits still depend on the endpoint Panify integrates.
Creative capabilities
Choose references and instructions around the creative control your scene needs.
Version guide
Hailuo and H3 serve different workflow needs.
| Model | Family role | Page-safe summary |
|---|---|---|
| Hailuo 01 / 02 | Focused video generation | Text-to-video and image-to-video workflows with increasing motion and camera control. |
| MiniMax H3 | Full-modal generation | Understands text, images, video and audio together and produces coordinated audio-video results. |
Workflows
Use cases
MiniMax FAQ
Review the full-modal workflow and prepare authorized reference media while Panify verifies production access.