MiniMax releases the full-modal generation model H3, supporting images, videos, and sound
2026-08-20 10:48:52
According to CoinMeta, MiniMax has released a full-modal generation model, H3, today. This model is capable of understanding text, images, videos, and audio simultaneously, and can generate or edit videos based on natural language instructions. Users can train H3 to learn camera movements and use characters and audio from different images. H3 can generate videos up to 15 seconds in length, supporting 2K resolution and stereo sound. The official price for 2K videos on API is $0.13 per second, so generating a 15-second video would cost approximately $1.95. In addition, MiniMax has also launched a multi-modal creation agent, MiniMax Design, which allows users to create videos directly through natural language commands. The system will optimize the input based on the task requirements and reference materials.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.