-+ 0.00%
-+ 0.00%
-+ 0.00%

MiniMax (00100) H3 released: video model moving from “video generation” to “commercial-grade production”

Zhitongcaijing·07/31/2026 01:41:04
Listen to the news

The Zhitong Finance App learned that on July 31, MiniMax (00100) released MiniMax H3, a next-generation multi-modal generation model, and announced the recent open source. As an important exploration of MiniMax from a “specialized task model” to “general multi-modal intelligence,” H3 no longer uses a single task such as generating, editing, and referencing images, videos, and sounds, but instead uniformly understands creative intent in a multi-modal context to complete more natural and consistent generation and expression.

MiniMax H3 supports various input formats such as text, images, audio, and video. It supports direct output with 2K resolution, and can generate audio and video content of up to 15 seconds. While improving audiovisual generation capabilities, it also further lowers the threshold for creating and using high-quality video content.

In terms of model effects, MiniMax H3 has commercial-grade multi-scene content generation capabilities. It excels in command compliance, text and brand information presentation, V2V Motion Transfer (video to video motion transfer), etc., and can achieve accurate and controlled multi-modal content editing and generation, and is widely used in commercial scenarios such as advertising, brands, e-commerce, product design, UI/UX, and games.

In the Artificial Analysis video model list, MiniMax H3 ranked first in the world in terms of video editing ability.

Picture3.jpg

Since its establishment, MiniMax has insisted on self-research in parallel with text and multi-modal models, and has been open source for three generations of M-series text models. H3 is the first open source multi-modal generation model launched by MiniMax.

H3 draws on proven methods in text model development, including complex problem solving, scaling context, and Muon optimizer, and applies them to multi-modal understanding, data construction, and model training. Focusing on understanding different tasks and following instructions, MiniMax iterated the data system in multiple rounds to improve the model's generalization ability in the face of open inputs and complex instructions.

The MiniMax H3 video generation price is 0.8 yuan/second (2K resolution), which is only one-third of similar flagship video models in the industry. MiniMax reduces the number of tokens required for video generation through a high-compression tokenizer, optimizing costs and improving efficiency. In response to the characteristics of large differences in video length, resolution, and multi-modal understanding and generation load, MiniMax optimizes the system in terms of heterogeneous training, load balancing, and GPU utilization efficiency to control training and inference costs while pursuing model results.

MiniMax H3 further brings multi-modal model competition in the industry to a comprehensive dimension of effectiveness, cost, openness, and ecological collaboration, and enriches the collaborative ecosystem of model companies and domestic AI chip software and hardware. At the same time, the MiniMax multi-modal route has entered a new stage of collaborative promotion of model capabilities, training efficiency, open source ecosystem, and industry applications.

“With MiniMax H3, we want to lower the threshold for content creation for more businesses, developers, and creators. H3 will be open sourced within the next few days, and enterprises can flexibly deploy locally, combine their own data and business to better meet security compliance requirements; chip manufacturers and developers can also participate in adaptation and optimization to reduce usage costs and expand the scope of application.” The relevant person in charge of MiniMax said, “We will continue to promote excellent text and multi-modal models from closed services in the past to a more open and collaborative industrial ecosystem.”