MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio
MiniMax H3 treats video generation as true multimodal synthesis—text, images, video, and audio in one unified model, not bolted-on features.

Why it matters
A new contender in the video-generation race ships native audio and longer context as built-in capabilities, not post-processing. Practitioners evaluating video APIs and researchers tracking the frontier labs' convergence on omni-modality should note the technical shift.
The key facts
6 to knowMiniMax H3: general-purpose multimodal generation model
Reads text, images, video, audio as unified context
Outputs: 2K video with native stereo audio
Duration: 4–15 seconds, integer-specified
Omni-modal architecture (not text-to-video + audio bolted on)
Positioned as alternative to Runway, Pika, Flux video competitors
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: MiniMax releases MiniMax H3, a general-purpose multimodal generation model. MiniMax H3 is not a text-to-video model with add-ons. MiniMax describes it as a general-purpose multimodal generation model that reads text, images, video, and audio as one unified context and returns video with native…