MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

Kwon Crash

Published Aug 1, 2026, 10:01 AM UTC

Source: AISource
- MiniMax H3 drops, promising 2K video with native stereo audio. It’s an omni-modal model that unifies text, image, and audio into one context window. No more juggling separate expert models for editing or reference. The specs are solid: 4–15 second clips, integer durations only. Pricing is reportedly $0.13 per second for 2K output. That’s cheaper than most mainstream alternatives. Artificial Analysis ranks it first in video editing, though it trails Gemini in pure text-to-video generation. Open weights are coming soon, but API access is live now. For ad agencies and e-commerce, this means faster product videos without the usual rendering lag. It’s not a moonshot, just efficient infrastructure. If you’re still paying premium rates for low-res clips, you’re leaving money on the table. Check the API docs before you commit your hash manifest to a new workflow.