Huawei Unveils Ascend 960 SuperPoD: Optical Efficiency Still Needs Chips and Supply
TL;DR
Huawei details optical interconnects and an accelerated 2027 chip schedule. Claimed power savings exceed 550 kW, but testing and supply constraints still separate the architecture from delivered capacity.
Huawei unveiled its Ascend 960 SuperPoD with near-package optics (NPO) on 2026-09-17, claiming a power reduction exceeding 550 kW. The announcement does not compare electricity consumed per completed training run with the same model and numerical precision. That missing evidence limits what the savings figure establishes: more efficient interconnects could improve the whole system, but do not yet demonstrate lower model-training costs. Huawei announcement
This is a new announcement from that day’s Huawei Connect conference in Shanghai, and this article also carries September 17 as its Taipei publication date. The official pages provide a date without a publication time. Reuters first published at 11:33 Taipei time and updated at 17:07, both before this run’s 20:46 cutoff. Reuters confirms that Huawei now schedules the 960DT for the first quarter of 2027 and the 960PR for the third quarter, respectively three quarters and one quarter earlier than planned. These remain future schedules. Reuters report
Optical engines reduce the burden of connecting chips
Huawei uses UnifiedBus to connect computing, memory and storage components, with unified memory addressing. The design aims to help processors work together while reducing the burden of moving data between devices. Architecture announcement The new SuperPoD scales to 4096 cards. Huawei says 5500 Hi-ONE optical engines replace 48000 800G optical modules; its claimed reduction of more than 550 kW refers to that replacement. Huawei announcement
Delivering data to working chips requires connecting equipment and electricity. Huawei is trying to reduce this overhead so that expanding a system does not require the same increase in optical modules as before. I favor this design direction because it addresses the burden of making many chips work together. How much it saves for a particular model, however, still depends on how much of that workload involves data exchange, as well as the price of the new equipment.
Huawei’s other announcement that day distinguishes the delivery stages: the NPO-based Atlas 960 system remains under testing, while the 256000-card cluster being deployed uses the preceding Atlas 950 generation. This is a concrete large-scale deployment example, but the company does not name the customer or publish acceptance results. Deployment of the earlier generation cannot be counted as shipments of the new system. Architecture announcement
An earlier chip schedule does not remove supply constraints
Reuters quotes rotating chairman Eric Xu saying that Huawei’s AI computing equipment capacity cannot yet satisfy demand within China. The company therefore is not pursuing a broad overseas expansion, although it makes limited deliveries abroad. Reuters also notes that Huawei supplied no data supporting its estimate that Ascend’s market share in China exceeds NVIDIA’s. Reuters report The architecture is announced; how much equipment buyers can obtain remains a separate constraint.
I consequently see this announcement primarily as an improvement in how computing resources work together, rather than a basis for immediately forecasting displacement of competitors. Suppose a model provider is already constrained by its electricity allocation. Lower interconnect consumption could leave more of that allocation available for useful computation. If the provider is waiting for chips that have not shipped, however, efficient optical engines cannot supply the missing production capacity. Those situations require different remedies.
A product-cost judgment should follow the same distinction. Lower electricity consumption can support lower computing costs only if total consumption falls for the same model and precision, including waiting, retries and cooling. If additional equipment costs offset the savings, service prices need not decline. This is a comparison proposed from the announcement; the sources lack complete testing conditions and prices, so there is no basis for assigning a return figure.
Once Atlas 960 completes testing and ships, I would examine total electricity consumption and completion time for a fixed workload, alongside the quantity of equipment buyers actually receive. The former can test whether optical interconnects improve work efficiency; the latter establishes how much available computing capacity the accelerated schedule adds.
The cover reuses this site’s Huawei HarmonyOS 7 launch image. It does not show Ascend 960 hardware or this conference; downloading the new official photograph failed.
Sources:
Related Articles
a16z Raises $1.1B Machine Age Fund for AI Hardware, Data Centers, and Power
Andreessen Horowitz has raised the $1.1B Machine Age Fund for chips, memory, networking, storage, data centers, robotics, and home AI devices, while its deployment pace and portfolio remain undisclosed.
OpenAI Publishes the First Jalapeño Benchmarks: The Comparison Conditions Matter
OpenAI says its Jalapeño system delivered 1.5 to 1.9 times the peak throughput with lower latency in InferenceX tests, but newer GPUs, speculative decoding, and system power were not included.