MiniMax has launched H3, a general-purpose omni-modal model designed to understand text, images, video and audio while generating video with native stereo sound. The company says H3 can produce clips of up to 15 seconds at resolutions up to 2K, extending its video tools beyond prompt-only generation.

The release is aimed at workflows that combine multiple kinds of creative direction. Users can provide reference media alongside written instructions, which MiniMax positions for motion transfer, video editing, brand presentation and more precise control of visual composition.

What H3 adds

H3 treats text, image, video and audio as context for a single generation process. That allows a creator to specify a subject in an image, motion from another clip, sound characteristics and written constraints without manually translating every input into a text prompt.

Native stereo-audio generation is also part of the model rather than a separately advertised post-production step. In practice, users will still need to assess how reliably dialogue, sound effects and ambience align with the generated action, particularly across longer 15-second clips and complex scenes.

MiniMax highlights instruction following, readable text and brand presentation as areas of focus. Those capabilities are commercially important because video models often struggle to preserve exact wording, logos, product geometry and continuity. The launch examples demonstrate intended behaviour, but they do not replace testing with a user's own assets and difficult prompts.

How MiniMax built the model

The company attributes H3's performance to three main components. Its Contextual Omni Representation is intended to place different input types into a shared context. H3-VAE compresses video and audio into a more efficient representation, while the H3-Omni Transformer performs generation across those representations.

MiniMax says H3-VAE produces an effective sequence length four times shorter than its previous approach and improved training throughput by nearly 30 per cent. These are company-reported engineering results; the announcement does not provide enough information to independently compare end-to-end inference speed or output quality with competing services.

An In-Context Regeneration technique is designed to support editing and controlled transformation. Rather than regenerating an entire concept from a fresh prompt, the model can use existing media as contextual guidance. That could reduce iteration for tasks such as changing motion, adjusting a scene or adapting a campaign asset, provided identity and layout remain consistent.

Availability, cost and open weights

H3 is available through MiniMax's product experience, and the company says model weights will be opened in the coming days after legal and release checks. The precise licence, downloadable formats, hardware requirements and permitted commercial uses were not detailed in the launch post, so developers should wait for the accompanying model card and licence before planning self-hosted deployments.

MiniMax claims its per-second price for 2K output is less than one-third of mainstream alternatives and that 768p output costs less than half the price of mainstream 720p models. The post does not state the comparison set or provide exact price figures, which makes the claim difficult to evaluate on its own. Buyers should compare final rendered cost, retries, generation time and usable-output rate rather than advertised price per second alone.

The promised open weights could make H3 relevant to studios and enterprises that need more control over deployment or customisation. However, a high-resolution video-and-audio model can carry substantial accelerator, storage and serving requirements. Operational cost, safety filters and model provenance will be as important as weight availability.

What teams should test

Creative teams should test prompt adherence, subject consistency, text rendering, lip synchronisation, stereo placement and temporal continuity across the full supported duration. They should also measure whether editing preserves untouched areas of a source clip and whether reference inputs create rights or privacy obligations.

MiniMax acknowledges that model scale, capability and visual detail can still improve. That caution is appropriate: the technical specifications establish a broader input and output envelope, but production value depends on reliability across repeated generations. H3's significance will ultimately be determined by how often it produces an acceptable result, how controllable revisions are and what conditions attach to the open release.