Alibaba Officially Launches Wan3.0, a Document-to-Video AI Model, One Day After Its Record $10 Billion Share Sale
Alibaba Cloud's Wan3.0 turns decks, spreadsheets and web pages into 30-second videos — and lands just as the company raises a record HK$80 billion to fund its AI buildout.
One day after raising US$10.2 billion in the largest primary follow-on offering ever completed by a Hong Kong-listed company, Alibaba has put that money straight to work. On Monday, August 24, 2026, the company officially rolled out Wan3.0, the newest version of its AI video generation model — and the first frontier video model built to read office documents and turn them into finished video.
The timing is not subtle. The share sale, announced Sunday, and the model launch, announced Monday, are two moves in the same strategy: pour capital into the AI race at any near-term cost, and ship the models that justify the spend.
What Wan3.0 Actually Does
Wan3.0, developed by Alibaba’s Tongyi Lab, generates up to 30 seconds of video in a single pass at up to 1080p. That headline number matters less than what sits behind it. The model unifies what previously required four separate models in the Wan 2.7 generation — text-to-video, image-to-video, reference-to-video, and video editing — into a single set of weights that handles every mode.
The genuinely new capability is on the input side. Wan3.0 accepts documents — Word files, Excel spreadsheets, PowerPoint decks, PDFs, plain text, and even Apple’s Pages and Keynote formats — as creative source material, alongside text, image, audio, and video references. According to Alibaba Cloud’s announcement on WeChat, the model can generate 30-second videos directly from documents, spreadsheets, slides, and web pages.
That collapses a pipeline that used to require several manual steps. Before Wan3.0, a team converting a quarterly results deck into a recap video had to read the deck, decide which points mattered, write a script, break it into a storyboard, and then generate scenes one at a time. Wan3.0 compresses that to: upload the deck, describe the tone, and generate. The model’s language understanding parses headings, bullet points, and table values natively rather than treating the document as a flat image.
The reference system is generous in scope: up to 10 images, 5 videos, and 5 audio files can condition a single generation, and audio is generated by default on every clip rather than as an opt-in extra. When a video reference is included, the reference duration plus the requested output is capped at 30 seconds combined.
From Beta to Full Release in Eighteen Days
Wan3.0 has had a fast track. The model entered public beta on August 6, 2026, on Alibaba Cloud Model Studio (model ID wan3.0-video) with a parallel release on Qwen Cloud, initially gated to invited accounts. Alibaba now says the model has already been used in short drama and film production, advertising and marketing, tourism promotion, and music video creation during the beta period.
Access runs through Alibaba Cloud Model Studio’s Beijing and Singapore endpoints. Pricing is published by resolution tier: roughly ¥0.30, ¥0.60, and ¥1.20 per second in Beijing at 480p, 720p, and 1080p respectively, with Singapore rates about 20% higher. At international rates, a 30-second 1080p clip costs about US$6 — or US$12 per finished minute.
Notably, Wan3.0 ships without open weights, a departure from Wan 2.1 and 2.2, which Alibaba released openly and which became fixtures of the open-source video generation ecosystem. The company has not said whether weights will follow.
The Context: A Record Raise and a 75% Earnings Plunge
The share sale tells the story of why Alibaba is shipping so aggressively. The company sold 710 million new ordinary shares at HK$112.70 each, raising HK$80 billion (US$10.2 billion) — the largest primary follow-on offering on record for a Hong Kong-listed company. The offering was priced at a 3.6% discount to Friday’s close, and the stock slid as much as 10% on Monday morning, its steepest drop since April.
The raise exists because the spending is enormous. Just last week, Alibaba reported a 75% year-over-year plunge in quarterly earnings, driven by soaring AI-related capital expenditure. The company is effectively trading near-term profitability for compute capacity, data centre construction, and model development — and then returning to the equity market to replenish the war chest.
A Different Axis of Competition
The AI video market has spent 2025 and 2026 in a clip-length arms race. Google’s Veo 3.1 sits at 8 seconds per pass, Runway’s Gen-4 at 12, and Kling 3.0 at 15. ByteDance’s Seedance 2.5 broke 30 seconds first, in June. Wan3.0 matches that 30-second ceiling but pointedly does not claim to beat it — Alibaba’s materials make no direct comparative claims against Seedance, Kling, or Veo.
Instead, Alibaba is competing on a dimension the rest of the field has not touched: none of the other frontier video models accept office documents as native input. Sora 2 and Veo 3.1 take text and image references. Kling 3.0 and Seedance 2.5 add video references. Wan 3.0 is alone in taking a .pptx file.
That positioning makes sense for Alibaba’s enterprise base. Sales teams sitting on pitch decks, training teams with onboarding manuals, and finance teams with quarterly reports represent a vast backlog of source material that never became video because someone had to convert it into a prompt first. The widely-cited (if methodologically fuzzy) industry estimate of roughly 35 million PowerPoint presentations given daily hints at the scale of that backlog.
What to Watch
Three open questions will determine whether Wan3.0’s launch becomes a durable shift or a beta footnote. First, whether the promised consumer-facing wan.video surface ships and access opens beyond invited enterprise accounts. Second, whether the model appears on independent leaderboards — its predecessor Wan 2.7 ranks fourth on the Artificial Analysis video arena, and until Wan 3.0 gets its own placement, quality claims rest on Alibaba’s demonstration material. And third, whether document parsing holds up on messy real-world decks with inconsistent formatting and embedded charts, rather than the clean four-bullet slides used in demos.
What is already clear is the strategic picture. Alibaba has raised a record amount of money, accepted a historic earnings hit, and shipped a model with a genuinely novel capability — all within the space of a week. In an AI market where Sora shut down in the spring and Seedance spent months frozen by an IP dispute, Alibaba is betting that aggression and capital, applied consistently, win the video generation race.
Sources are listed in the article metadata.
Sources
- [1] https://www.channelnewsasia.com/business/alibaba-launches-wan30-ai-video-model-after-10-billion-share-sale-6337366
- [2] https://www.reuters.com/business/retail-consumer/alibaba-proposes-hong-kong-share-placement-worth-10-billion-2026-08-23/
- [3] https://www.alibabacloud.com/blog/wan3-0-30-second-ai-video-generation-from-any-input_603452
- [4] https://www.ngram.com/blog/wan-3-0-document-to-video-ai-model
- [5] https://www.businesstimes.com.sg/companies-markets/alibaba-raises-us10-billion-record-hong-kong-share-sale