Skip to content

cloud

2 posts with the tag “cloud”

Alibaba's Zhenwu V900: China's Most Powerful AI Chip and a 10-Trillion-Parameter Roadmap

Alibaba’s Zhenwu V900: China’s Most Powerful AI Chip and a 10-Trillion-Parameter Roadmap

Section titled “Alibaba’s Zhenwu V900: China’s Most Powerful AI Chip and a 10-Trillion-Parameter Roadmap”

Alibaba opened its annual Apsara conference in Hangzhou on September 22 with a full-stack AI announcement: a new AI chip, a 20-gigawatt data center target, and a Qwen model roadmap that reaches 10 trillion parameters (The Next Web, 2026). CEO Eddie Wu called the chip “the most powerful AI chip in China today” (NBC News, 2026). The announcement lands days before a U.S.-China summit where AI leadership is a stated theme (NBC News, 2026).

The V900 is the successor to the Zhenwu M890, which launched in May 2026 (FinanceFeeds, 2026). Wu said the V900 delivers three times the performance of the M890 (The Next Web, 2026).

Two numbers define the scale ambition:

  • Clusters can connect up to 500,000 V900 chips for training runs (FinanceFeeds, 2026).
  • Mass production and commercial release are planned for the first quarter of 2027 (TrendForce, 2026).

The Zhenwu series already serves more than 650 enterprise customers across autonomous driving, finance, large language models, embodied AI, energy, and manufacturing (TechNode, 2026).

Alibaba Cloud plans to run more than 20 gigawatts of data center capacity worldwide by 2032 (CNBC, 2026). The scale shows how much power Alibaba expects next-generation AI systems to consume (FinanceFeeds, 2026). Hong Kong-listed shares of Alibaba rose more than 3% on the announcement (CNBC, 2026).

The model roadmap: Qwen goes to 10 trillion

Section titled “The model roadmap: Qwen goes to 10 trillion”

Alibaba’s next-generation Qwen 4 model is currently in training (Reuters, 2026). The company projects that Qwen 4.5 and Qwen 5 series models will reach 5 trillion to 10 trillion parameters (Asia Tech Review, 2026).

For scale, the current flagship Qwen 3.8 Max has 2.4 trillion parameters (Reuters, 2026). The planned 10-trillion model would be roughly two to four times larger (Reuters, 2026).

Alibaba’s proprietary M890 AI supernode already handles inference for models above 2 trillion parameters (Reuters, 2026). Wu said only “a handful” of systems can do this today (Reuters, 2026).

T-Head, Alibaba’s semiconductor arm, also mapped its server CPU line (TrendForce, 2026). The Yitian 720 and Yitian 730 server CPUs are scheduled to launch in the third quarter of 2027 (TrendForce, 2026). A later Yitian 750 adds ICN-link direct attach to Zhenwu AI accelerators (Pandaily, 2026).

The full stack now covers four chip classes: Zhenwu AI accelerators, Yitian CPUs, Panmai smart NICs, and ICN interconnect chips (TechNode, 2026).

Alibaba is building the AI stack from silicon to deployed model, and that changes three planning assumptions:

  1. GPU supply is diversifying. When a hyperscaler ships its own accelerator, CUDA dependence becomes a choice, not a default (The Next Web, 2026). Teams should keep workloads portable across accelerator vendors.
  2. Cluster scale is the new metric. A 500,000-chip training cluster means orchestration, networking, and fault-tolerance at a size most operators have not scheduled for (FinanceFeeds, 2026).
  3. Power is the constraint. Twenty gigawatts by 2032 forces site selection, cooling, and grid contracts to the front of AI infrastructure planning (CNBC, 2026).

Watch the Q1 2027 mass-production window for the V900 and the Qwen 4 release (TrendForce, 2026). Both dates will test whether the full-stack claim holds under real load (Asia Tech Review, 2026). For teams adopting Qwen models, plan for the parameter jump now: the difference between 2.4 trillion and 10 trillion parameters is not a bigger GPU, it is a different infrastructure class (Reuters, 2026).

Microsoft's Missing AI Chips: The $280B Buildout That Can't Plug In

On August 17, 2026, the Guardian published an investigation into Microsoft’s AI buildout. Its reporters reviewed internal Microsoft documents (The Guardian). The documents show about 2.2 million AI chips installed globally, against roughly $280 billion spent since 2022. The gap between announced capacity and working hardware is now the central question in AI infrastructure (The Guardian).

Microsoft reported 5GW of data-centre capacity added over two years. It set an internal target of 1.8 million installed chips by the end of 2024 (SightsIn Plus). The installed count today only modestly exceeds that two-year-old target. That is not the picture the spending suggested (BERI).

The investigation is not about a chip shortage. It is about how little of the purchased hardware can actually run (SightsIn Plus).

Shaolei Ren, a professor at the University of California, Riverside, read Microsoft’s audited sustainability reports. He estimated the company’s 2024 AI capacity at closer to 1.2GW. He concluded that, combined with the reported 5GW addition, Microsoft would need roughly 4 million chips to fill that footprint (SightsIn Plus). The ~2.2 million installed is less than half that figure.

One Nvidia analyst told the Guardian the count looked wrong. “They’re low to me. They’re less than I expected Microsoft would have,” the analyst said (Inside Telecom).

Microsoft says the arithmetic is wrong, but it does not dispute the mechanism behind it (BERI).

CEO Satya Nadella described the constraint bluntly. “You may actually have a bunch of chips sitting in inventory that I can’t plug in,” he said. “In fact, that is my problem today. It’s not a supply issue of chips. It’s actually the fact that I don’t have warm shells to plug into” (The Guardian).

A warm shell is a completed data-centre building. It has power, cooling, and rack space ready for hardware. Nadella made the same point months earlier: “The biggest issue we are now having is not a compute glut, but it’s power” (BERI).

Servers need three things that are not chips: power, cooling, and completed buildings. A company can secure processors and leave them unused if a data centre cannot connect to the grid (Inside Telecom).

Delays compound the problem. The Guardian’s investigation also flagged questions around Microsoft’s Fairwater data-centre project and how much announced capacity is truly online (TechStartups). Microsoft has rejected the investigation’s calculations (Inside Telecom).

Why this matters for anyone provisioning AI

Section titled “Why this matters for anyone provisioning AI”

Announced capacity is not live capacity. That distinction is the reason provisioned-throughput orders get rejected (BERI).

A cloud that has bought millions of chips cannot sell compute it cannot power. The wall has moved downstream from silicon to electricity and construction (BERI).

This matters beyond Microsoft. Every major AI buildout hits the same three walls. Getting GPUs is the easy part. Turning them into working capacity requires grid power and finished facilities (Inside Telecom).

Microsoft is pushing its own chip to cut its dependence on Nvidia. It plans to unveil the next-generation Maia 300 accelerator as soon as September (AI Weekly).

The company is negotiating with TSMC for more than 300,000 units, with delivery targeted for 2027. Its longer-term ambition is capacity for over one million chips (AI Weekly).

Every AI accelerator depends on a single packaging process that Nvidia largely controls. That packaging queue is a real obstacle for any custom chip program (TechTimes).

Andrew Wall, general manager for Azure Maia, said Microsoft “continues to invest in custom silicon as part of our long-term AI infrastructure strategy.” He added that the production figures reported “don’t reflect the scale of our program” (Quartz via Yahoo Finance).

The 300,000-unit figure is still a negotiation, not a signed order. The exact number is a moving target, not a confirmed plan (AI Weekly).

  1. Audit real capacity, not announced capacity. A vendor’s GPU count means little without power and facilities behind it (Inside Telecom).
  2. Treat power as the scheduling constraint. The biggest AI issue is no longer compute supply. It is power and finished buildings (BERI).
  3. Plan long lead times for capacity. If your provisioned throughput gets rejected, the vendor’s hardware may be sitting unplugged (BERI).
  4. Watch for the bringing-down-own-silicon shift. When a cloud runs its own chip, every Maia workload is one it does not run on Nvidia at Nvidia’s margins. That is a future cost driver for AI services (TechTimes).

The AI buildout has hit its physical wall. Microsoft has spent $280 billion and installed 2.2 million chips, but the machines it can actually switch on are far fewer (The Guardian). Power, cooling, and warm shells now decide when the next wave of capacity arrives. Buy delivery. Do not buy capex (BERI).