AI & Tech 🔥 Trending

MiniMax H3 Open-Weights Video Generates Its Own Audio

MiniMax open-sourced H3 on 3 August: an omni-modal model that takes text, image, video and audio in one context and outputs up to 2K at 24fps with native stereo audio. The contribution is not the resolution but the joint generation — voice, sound effects and music are produced in the same forward pass as the picture, removing the alignment step that every two-stage video-then-dub pipeline spends its time on. Quantisation cuts the footprint from 123.6GB to 42.5GB, and ComfyUI shipped day-zero support with four nodes and six workflow templates. Clips cap at roughly 15 seconds.

Read the original — via ComfyUI Blog ↗

← All shorts