Alibaba's Amap Team Open-Sources DreamX-Creator: A 7B Model That Generates Video and Sound Together at 2K
DreamX-Creator 1.0 jointly denoises audio and video streams in a single 7B model, adds RL with multimodal feedback, and refines output to 2K in one denoising step per chunk.
Most AI video generators today are silent films with a soundtrack bolted on afterward. The video model renders the frames, a separate audio model synthesizes sound, and an alignment step tries to glue the two together. Alibaba’s Amap machine learning team (AMAP-ML, publishing as the DreamX Team) has just published a technical report that takes the other road: DreamX-Creator 1.0, a compact 7-billion-parameter system that generates synchronized audio and video natively, in one model, at 2K resolution — and plans to release the weights under Apache 2.0.
The paper, posted to arXiv as 2608.31106 and surfacing on Hugging Face Daily Papers on September 1, climbed to 84 upvotes within hours and pulled the freshly initialized GitHub repository past 90 stars. For a system whose stated goal is “democratizing native audio-video generation,” the reception suggests the open-source community has been waiting for exactly this.
Why joint generation is hard
The core problem is coupling. Visual dynamics and acoustic events influence each other — a glass shattering on screen dictates a specific sound at a specific moment, and the rustle of off-screen movement should be inferable from what the scene implies. When audio is synthesized in a second stage, that reciprocal modeling is severed: the audio model watches the finished video but can never influence it, and fine-grained temporal synchronization becomes a post-processing headache.
DreamX-Creator’s answer starts at the architecture level. Conditioned on a first frame and a text prompt, the 7B generator jointly denoises modality-specialized audio and video streams. The two streams are processed independently through the first half of the network — letting each modality build up its own representations — and then coupled in the latter half through Gated Cross-Modal Attention, a mechanism whose token- and head-wise output gates modulate each active cross-modal attention head. In plain terms: rather than blindly averaging information across modalities, the network learns how much visual information should shape the audio stream and vice versa, at the level of individual attention heads and tokens.
A data system, not just a model
Like most modern generative systems, DreamX-Creator’s reported quality rests on a data pipeline the paper treats as a first-class citizen. The Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools — so that training data isn’t just large, but sorted by what it teaches.
Training itself follows Progressive Joint Training: two audio-video pre-training stages followed by a High-Quality Finetuning phase. Then comes the step that separates this work from a standard diffusion recipe — Audio-Video Reinforcement Learning. The generator is post-trained with what the authors call Modality-Aware Multimodal Feedback, which routes video-side, audio-side, and cross-modal reward signals to the corresponding streams. Visual quality rewards flow to the video stream; audio fidelity rewards flow to the audio stream; synchronization and semantic-consistency rewards touch both. It is a division-of-labor approach to RLHF that mirrors the modality-specialized architecture.
The 2K trick: distill until one step is enough
High-resolution video is where diffusion latency usually explodes. DreamX-Creator’s Autoregressive 1-Step 2K Refinement tackles this with a three-stage pipeline: a bidirectional multi-step teacher is adapted into an autoregressive multi-step refiner, which is then distilled into a student that requires a single denoising evaluation per temporal chunk. The refiner upgrades lower-resolution generations to 2K while preserving content, motion, and audio-aligned timing — meaning the expensive model does the teaching, and the cheap student does the serving.
The team reports performance “competitive with state-of-the-art open-source systems” — a claim that independent benchmarks will need to test once weights are public. The GitHub roadmap shows the repository initialization and technical report release are complete, with validated model weights, inference code, configurations, and evaluation tools still flagged as pending. The acknowledgments cite Alibaba’s Wan video team and Fudan’s OpenMOSS MOVA project as foundations, placing DreamX-Creator firmly in the fast-growing lineage of Chinese open-source generative models.
Why it matters
A compact 7B joint audio-video generator with an Apache 2.0 license would be a meaningful shift for developers. Today, “video with sound” in production usually means orchestrating two or three APIs — a video model, a music/sfx model, a stitching layer — with costs and failure modes multiplying at each seam. A single native model collapses that stack, and at 7B it is small enough to fine-tune and serve on hardware well below frontier-lab scale. That is the “democratizing” in the title made concrete.
It also signals where the frontier is moving. Google’s Veo line made native audio a flagship feature; open-source answers like Wan 2.6 and MOVA pushed synchronized sound into the open ecosystem. DreamX-Creator raises the open-source bar by combining three things that rarely ship together: joint generation, reinforcement learning on multimodal feedback, and 2K refinement at one denoising step per chunk.
The caveats are the usual ones for a day-old technical report: no public weights yet, quantitative claims awaiting third-party evaluation, and no demonstration yet of how the system handles long multi-shot sequences or dialogue-heavy content where lip sync is unforgiving. But the architecture write-up is detailed, the code repository is live, and the team has a track record through Alibaba’s Amap unit of shipping production-scale multimodal systems.
For anyone building video tooling, this is one to watch — and, once the weights land, one to download.