On August 3, 2026, MiniMax officially open-sourced its cutting-edge multimodal generative model MiniMax H3. On the same day, Moore Threads completed ultra-fast adaptation and stable deployment of the H3 model on its MTT S5000 AI training and inference GPU, marking a landmark ecosystem breakthrough for domestic general-purpose GPU vendors. The rapid implementation fully demonstrates MUSA’s full-stack collaborative strength and ultra-fast iteration capability for frontier multimodal models, validating Moore Threads’ leading full-featured GPU competency in empowering next-generation generative AI and industrial ecosystem development.

Post Images
Moore Threads Technology

MiniMax H3: High-Performance, Cost-Effective Commercial-Grade Multimodal Foundation Model

As MiniMax’s first open-source universal multimodal generative model, H3 breaks the task boundaries of traditional single-modality AI models and unifies creative intent understanding within complex multimodal contexts. The model supports diversified input forms including text, image, audio and video, and natively generates synchronized audio-video content with up to 2K resolution and a maximum duration of 15 seconds. With outstanding comprehensive generative capabilities and industrial cost advantages, H3 secured the global No.1 ranking in video editing capability on the authoritative Artificial Analysis video model benchmark. Its commercial 2K video generation cost is as low as RMB 0.8 per second, delivering exceptional commercial cost-effectiveness for large-scale industrial deployment.

In addition, the model’s open-source attributes significantly lower the threshold for enterprise-level multimodal AI application. It supports flexible on-premises deployment and customized fine-tuning based on private data, effectively meeting enterprise demands for data security and business compliance while accelerating the popularization of multimodal AI productivity tools across industries.

Industry-Leading Engineering Efficiency: 3-Hour Ultra-Fast Full-Stack Adaptation

Exemplifying top-tier domestic GPU engineering capabilities, Moore Threads completed the full industrial adaptation of MiniMax H3 within only three hours after the model’s official open-source release. The technical team rapidly accomplished model architecture decomposition, core mechanism analysis and key operator sorting, and built a complete and stable adaptation pipeline covering the SGLang-MUSA inference framework, MATE operator optimization engine and muDNN high-performance operator library. The project realized one-stop deployment and stable operation of the H3 model on a single-node 8-card MTT S5000 cluster, setting a new benchmark for domestic GPU’s Day-0 adaptation efficiency for cutting-edge multimodal models.

Post Images
Moore Threads Technology

MTT S5000 High-Density Computing Infrastructure Supports Multimodal High-Load Inference

Compared with traditional single-modal large models, MiniMax H3’s multimodal contextual mechanism expands sequence length variance by three times, bringing a substantial increase in computational load throughout contextual understanding and content generation stages and raising stricter requirements for underlying GPU computing power, throughput and memory bandwidth. Equipped with industry-leading high computing density and high-bandwidth memory architecture, Moore Threads’ MTT S5000 AI training and inference card fully undertakes the full-cycle high-intensity computing demands of H3 multimodal inference. The hardware effectively releases the complete generative performance of the H3 model, delivering stable and high-throughput computing support for complex multimodal contextual tasks.

MUSA Full-Stack Software Ecosystem Forms Closed-Loop Acceleration Capability

The successful Day-0 adaptation of MiniMax H3 relies on the systematic synergy of Moore Threads’ layered MUSA software stack, rather than simple single-operator transplantation. The MUSA ecosystem features high native compatibility with mainstream global AI frameworks and excellent cross-model migration flexibility, enabling rapid identification, matching and verification of new-model operators, and greatly shortening the adaptation cycle of cutting-edge AI architectures.

The self-developed open-source MATE optimization engine undertakes core hardware scheduling, capability detection and backend resource allocation, providing highly optimized underlying implementations and unified calling interfaces for core operators such as Attention and GEMM. Combined with the MTCC compilation toolchain and MUSA Runtime, it forms an efficient closed-loop acceleration chain: SGLang-MUSA → MATE → MTCC → MUSA Runtime, achieving full-link performance optimization for multimodal generative tasks.

At the inference framework level, SGLang-MUSA was officially merged into the upstream SGLang main branch in April 2026, gaining official native backend recognition in the global mainstream inference ecosystem. Developers can natively invoke Moore Threads GPU acceleration across the entire SGLang toolchain, including the core sglang inference framework, high-performance sgl-kernel operator library and SGLang-Diffusion multimodal generation subsystem, achieving extreme end-to-end inference performance.

Outlook: Deepened MiniMax Collaboration Accelerates AGI Industrialization

The ultra-fast adaptation of MiniMax H3 further verifies MUSA ecosystem’s mature “release-to-adapt” rapid iteration capability for frontier AI models. Moore Threads’ complete industrial engineering system covering model adaptation, cluster deployment and performance optimization significantly reduces the R&D and application costs of cutting-edge multimodal AI for industrial clients. Moving forward, Moore Threads will further deepen strategic cooperation with MiniMax, jointly exploring larger-scale, longer-context multimodal model capabilities, and continuously empowering technological innovation and large-scale industrial landing of AGI applications.

Invest It

Choose INVEST IT as your preferred source on Listed Companies of china and never miss a moment from the most trusted name in business news.

Featured Video

CONVERGE