The single-chip narrative is a distraction. The real war is being fought in the interconnect fabric, the software stack, and the memory supply chain.
Beijing's stated goal of training frontier AI models on domestic hardware by 2028 is not a technology roadmap. It is a declaration of war against a supply chain architecture that has been optimized for two decades around one company: NVIDIA. The market reads this as a chip race. It is not. It is a systems engineering problem disguised as a semiconductor policy.
The headline numbers are seductive. Huawei's Ascend 910B delivers roughly 320 TFLOPS in FP16, marginally edging out the A100's 312 TFLOPS. The 910C is projected to reach 70-80% of H100 performance. Cambricon's Siyuan 590 approaches A100-level energy efficiency. On paper, the gap is closing faster than most Western analysts anticipated. But paper does not train models. Clusters do.
The Cluster Is the Unit of Computation
Here is the uncomfortable truth that gets lost in the spec-sheet theater: frontier model training is not a single-chip problem. It is a distributed systems problem. A GPT-4-class model requires roughly 10^25 FLOPs. By 2028, that figure will likely reach 10^26 to 10^27. No single accelerator can deliver that. You need tens of thousands of chips working in lockstep, with minimal communication overhead, sustained fault tolerance, and near-linear scaling efficiency.
This is where the Chinese ecosystem hits its hardest wall.
NVIDIA's dominance is not the GPU. It is the fabric β NVLink for chip-to-chip communication, InfiniBand for node-to-node, and a software stack that has been battle-hardened across a decade of hyperscale deployments. The result is a cluster that scales to 100,000 GPUs with a Model FLOPs Utilization (MFU) of 50-60%. That MFU number is the dirty secret of the AI arms race. It represents the percentage of theoretical compute actually used during training. Everything else is wasted cycles, idle silicon, and heat.
Chinese domestic solutions β Huawei's HCCS for chip-to-chip, RoCE for node-to-node β deliver roughly 400-500 GB/s of interconnect bandwidth versus NVIDIA's 900 GB/s+. Industry estimates place domestic cluster scaling efficiency at 70-85% of NVIDIA's equivalent. The 2028 target demands at least 90%. That gap is not a chip problem. It is a physics problem, a networking problem, and a software problem rolled into one.
The MFU gap is the hidden tax on domestic compute. At 30-40% MFU versus NVIDIA's 50-60%, a Chinese cluster needs 50-70% more raw hardware to deliver the same effective training throughput. That is not a minor inefficiency. That is a structural cost disadvantage that ripples through every downstream application.

The CUDA Gravity Well
The second bottleneck is less visible but arguably more intractable: software ecosystem lock-in.

CUDA is not just a programming model. It is a gravitational field that has captured the entire machine learning research community. PyTorch, TensorFlow, Megatron-DeepSpeed, FSDP β every major framework has been optimized, debugged, and battle-tested against NVIDIA hardware for years. The operator libraries are rich. The distributed training tools are mature. The developer knowledge base is enormous.
Huawei's CANN platform and MindSpore framework are improving. The Ascend community claims over 2 million developers. But developer count is not ecosystem maturity. The question is whether a researcher can take a state-of-the-art model architecture, port it to Ascend hardware, and achieve comparable training efficiency without weeks of manual optimization. Today, the answer is often no. The migration cost β learning curves, compatibility issues, performance losses β is the hidden tax on domestic compute.
This is where the "code-first" lens matters. Where the code forks, we find the fold. The fork between CUDA and CANN is not a technical divergence. It is a strategic chokepoint. NVIDIA's moat is not the silicon; it is the accumulated engineering hours embedded in the software stack. That is not something a policy directive can replicate in three years.
The HBM Supply Chain Trap
The third constraint is the one that keeps supply chain analysts awake at night: High Bandwidth Memory.

Domestic AI chips depend on HBM2E and HBM3, manufactured primarily by Samsung and SK Hynix. Both are subject to US export controls. China's domestic HBM efforts β led by ChangXin Memory Technologies β are in early stages. The gap between early-stage HBM development and production-grade, high-yield, high-bandwidth memory suitable for frontier-scale training is measured in years, not quarters.
This is the critical vulnerability. You can design a world-class AI chip, but without HBM, it is a paperweight. The 2028 timeline assumes either a breakthrough in domestic HBM production or a relaxation of export controls. Both are uncertain. The first is a physics and chemistry problem. The second is a geopolitical problem. Neither yields to policy mandates.
The "Frontier" Definition Problem
There is a strategic ambiguity embedded in the 2028 target that deserves scrutiny: what exactly does "frontier AI model" mean?
If it means "matching the global state-of-the-art at that moment" β a moving target that will be defined by whatever OpenAI, Google, and Anthropic are training in 2028 β then the goal is extraordinarily aggressive. The compute requirements for frontier models are doubling every few months. Catching a moving target while running on constrained hardware is a different problem than closing a static gap.
If it means "approaching or matching GPT-4-class capabilities" β the 2024 state-of-the-art β then the target is far more realistic. A 2028 model trained on domestic hardware that achieves 2024-level frontier performance would be a significant achievement, even if it lags the 2028 global frontier.
This definitional flexibility is not an oversight. It is a political necessity. It gives policymakers room to declare victory under either scenario. The market should not mistake this ambiguity for a commitment to absolute parity.
The Contrarian Angle: What the Market Misses
The conventional reading of this story is straightforward: China is building a domestic AI supply chain to counter US export controls. The market prices this as a binary outcome β either China succeeds and NVIDIA loses its largest overseas market, or China fails and the status quo holds.
The contrarian view is more nuanced. The real impact of the 2028 plan is not the Chinese domestic market. It is the global fragmentation of AI compute standards.
If China demonstrates that frontier-adjacent models can be trained on non-NVIDIA hardware, it validates an alternative path for every country that fears US technology leverage. Russia, Iran, and a growing list of Global South nations are watching. The "compute sovereignty" concept β analogous to data sovereignty β will become a mainstream policy framework. The result will not be a binary win-loss for NVIDIA. It will be a multi-polar compute ecosystem where the NVIDIA+CUDA stack is one option among several.
This is the deeper strategic play. Governance is not a vote; it is a vector. The 2028 plan is not just about training models. It is about establishing an alternative vector for global AI development that does not route through US-controlled infrastructure.
The second contrarian insight concerns the investment angle. The market is pricing domestic compute stocks β Cambricon, Hygon, and others β as if policy support guarantees commercial success. Cambricon trades at over 50x price-to-sales, roughly double NVIDIA's multiple. This pricing embeds an assumption that policy-driven demand will translate into sustainable commercial margins. That assumption ignores the unit economics problem: domestic chips currently have higher per-compute costs due to process node disadvantages, higher power consumption, and lower MFU. The total cost of ownership may be competitive when accounting for import difficulty, but the margin structure will be thinner than the narrative suggests.
Floor cracks reveal the foundation's weight. The foundation of the domestic compute story is not chip design. It is the entire stack β manufacturing, packaging, memory, networking, software, and operational expertise. Each layer has its own cracks. The market is pricing the top layer while ignoring the structural weaknesses below.
The 2028 Reality Check
What is the realistic outcome for 2028?
The most probable scenario is a "usable but not optimal" outcome. China will likely field domestic clusters capable of training models that approach β but do not match β the global frontier. The gap will be measurable but not disqualifying. The models will be sufficient for domestic applications, government use cases, and a growing export market to allied nations. They will not be sufficient to challenge NVIDIA's dominance in the global AI research ecosystem.
The more interesting question is what happens after 2028. The plan's real value is not the specific milestone. It is the institutional learning that occurs during the attempt. Every failed optimization, every interconnect bottleneck, every software migration pain point generates knowledge that cannot be acquired any other way. The 2028 target is a forcing function for building capabilities that will pay dividends in 2030 and beyond.
Hedging is the art of profiting from fear. The fear here is US technological leverage. The hedge is a parallel compute ecosystem. Whether the 2028 target is fully met is almost beside the point. The attempt itself reshapes the global compute landscape.
The Signal in the Noise
The market should track three signals between now and 2028.
First, the Ascend 910C production timeline and real-world cluster performance. Spec sheets are marketing. MFU numbers from actual training runs are truth.
Second, the HBM supply chain. Any progress from ChangXin Memory Technologies on production-grade HBM is more significant than any single chip announcement.
Third, the developer migration curve. Watch the number of production models β not demos β trained on domestic hardware. That is the real ecosystem maturity metric.
The 2028 plan is not a technology roadmap. It is a stress test of China's ability to build complex systems under constraint. The chips are the visible surface. The invisible layers β interconnect, software, memory, operations β will determine the outcome.
The ledger remembers what the market forgets. The market is pricing a chip race. The ledger will record a systems engineering marathon. Those are different events with different winners.
The question is not whether China can build a competitive AI chip by 2028. It is whether China can build a competitive AI system β and whether the global market is prepared for a world where the answer is a qualified yes.