china ai hardware decoupling notes

published: August 21, 2025updated: September 8, 2026
on this page

current observations (august 2025)

  • the format: ue8m0 emerges as key technical differentiator - 8-bit float with zero mantissa for ai inference

  • recent catalyst: lutnick’s july 2025 “addiction” comment accelerated existing regulatory shifts against nvidia

  • hardware ecosystem: moore threads (ex-nvidia china leadership) positioned as domestic ue8m0 partner after years of development

  • manufacturing pressure: export controls restrict access to leading equipment and overseas foundries, but public evidence does not establish a fixed domestic node limit

  • strategic split: training may still require nvidia hardware, but ue8m0 chips enable large-scale inference with existing or “borrowed” model weights

  • market response: nvidia’s august 2025 ue8m0 support suggests acceptance of parallel ecosystems

note: this is an evolving collection of observations on china-us ai hardware decoupling that began accelerating in 2019-2020. this page documents recent developments in a multi-year strategic divergence, with updates added as new information becomes available.

Policy update (September 2026). The United States relaxed one part of the export regime; it did not reverse the wider split. A Bureau of Industry and Security rule published in January 2026 moved license applications for Nvidia H200, AMD MI325X, and similar chips to case-by-case review. Applicants must meet supply, customer-screening, and independent-testing conditions. See the BIS announcement.

Two forward indicators listed at the end of this page have also resolved. Moore Threads began trading on the STAR Market in December 2025. Cambricon reported CNY 6.00 billion of first-half 2026 revenue and CNY 2.31 billion of net income. The analysis below remains a dated August 2025 snapshot unless a later note says otherwise.

recent developments (july-august 2025)

the lutnick comment in july 2025 represents a continuation of tensions that have been building since the october 2022 export controls and earlier trade restrictions dating to 2019.

on july 15, 2025, u.s. commerce secretary howard lutnick stated: “you want to sell the chinese enough that their developers get addicted to the american technology stack.”1

this comment catalyzed existing chinese regulatory momentum:1

  • july 22: cyberspace administration issues guidance to halt h20 purchases1
  • july 31: cac summons nvidia executives over “serious security issues”1
  • august: ndrc requests tech groups refrain from nvidia chip purchases1

these actions build on years of preparation for technological independence, including substantial investments in domestic semiconductor capabilities beginning in 2020.

technical divergence strategy

format differentiation

the ue8m0 data format represents the latest phase in a multi-year effort to develop alternative technical standards. deepseek’s explicit statement that ue8m0 fp8 scale is “designed for the upcoming next-generation domestically produced chips”6 reflects years of coordinated development between chinese ai companies and hardware manufacturers.

ue8m0 technical details

  • 8-bit exponent, 0 mantissa design differs from standard fp83

  • optimized for inference over training workloads2

  • reduces memory bandwidth by up to 75%7

  • nvidia added ue8m0 support in ptx isa 9.0 (august 2025)3

parallel software ecosystems

the development of alternative frameworks has been ongoing since at least 2020:

  • moore threads musa (2021-present): cuda-compatible platform with musify migration tool8
  • huawei cann (2019-present): proprietary framework for ascend chips, accelerated after entity list addition
  • deepseek deepep (2024-present): hardware-specific optimizations showing ue8m0 regression issues on non-gb200 hardware9

model-hardware co-evolution

deepseek v3.1 (august 21, 2025) trained with ue8m0 format represents culmination of multi-year collaboration.6 the 840 billion additional training tokens and format-specific optimization create technical lock-in effects that reinforce ecosystem separation.

critically, while model training may still benefit from or require nvidia hardware for optimal performance, the ue8m0-optimized chips enable china to deploy large-scale inference infrastructure using model weights developed domestically or obtained through other channels. this decouples inference capability from training dependency.

key players

companyfocusecosystemkey product
moore threadsconsumer/researchmusa (cuda-compatible)

mtt s90, s400010

huaweienterprise/governmentcann (proprietary)ascend 910c
biren technologydatacentertraditional gpubr100/104
cambriconinference/cloudspecialized npumlu590, mlu370

structural constraints

manufacturing limitations (2020-present)

china’s fabrication constraints have shaped strategy since the 2020 entity list additions:

  • the Netherlands restricts exports of ASML’s extreme-ultraviolet systems to China;
  • United States rules cover specified advanced chips, manufacturing equipment, and related technology;
  • public estimates of SMIC process nodes and yields vary, and neither company filings nor the cited rules establish one fixed limit through 2026.

these limits raise the cost and difficulty of leading-edge production. they do not prove that ue8m0 was adopted for one manufacturing reason alone.

key timeline (selected events)

dateeventcontext
may 2019huawei entity list additioncatalyst for domestic chip development
oct 2020moore threads founded

ex-nvidia china gm starts gpu company10

oct 2022us export controls on advanced chipsrestricts nvidia a100/h100 to china
oct 2023moore threads entity list

blocks access to tsmc, design tools10

dec 2023mtt s4000 launch

notably lacks fp8 support10

feb 2025moore threads-deepseek partnership

hardware-software alignment10

jul 2025lutnick comments

accelerates existing tensions1

aug 2025deepseek v3.1 with ue8m0

format designed for domestic chips6

market implications

global ai chip market projections

amd ceo lisa su expects the ai processor market to exceed $500 billion by 2028.12 asia-pacific region led with 33% market share in 2023, with china as key driver.13

china market impact

august 2025 market response to ue8m0 developments:

  • cambricon shares surged 20% hitting daily trading limits
  • chinese ai chip stocks rallied on hardware-software co-optimization news
  • investors positioned ue8m0 as breakthrough enabling domestic chip competitiveness
  • market response validates years of strategic semiconductor investments

the market forecasts in the 2025 draft described possible futures, not observed separation. the firmer evidence is institutional: china continues to fund domestic chips and software, while united states rules limit access to specified products. the january 2026 h200 policy shows that the boundary can move without ending the wider controls.

nvidia faces strategic dilemma:

  1. support ue8m0 (validates china’s strategy)
  2. ignore ue8m0 (loses china market access)
  3. create compatibility bridges (undermines u.s. policy)

the addition of ue8m0 to ptx isa 9.0 suggests nvidia chose option 1.3

government funding acceleration

china’s big fund phase 3 established may 2024 with ¥344 billion ($47.5 billion) - the largest semiconductor state investment fund globally.14

targeted company funding:

  • cambricon: reported cny 6.00 billion in first-half 2026 revenue
  • moore threads: began star market trading in december 2025
  • domestic chip mandates driving enterprise procurement shifts
  • bytedance and other tech giants increasing domestic chip adoption

observations and analysis

technical tradeoffs (current state)

the ue8m0 approach reflects years of navigating constraints:

  • error tolerance: 7e-4 (ue8m0) vs 1e-5 (standard fp8)9
  • memory reduction: up to 75%7
  • simplified hardware: no mantissa circuits3
  • inference focus: lower-precision formats can reduce memory traffic for supported inference kernels
  • training-inference split: accepts continued nvidia dependency for training while achieving inference independence

indicators to track

ongoing developments to monitor:

  • moore threads post-listing financial and product disclosures
  • cambricon filed revenue, profit, and product disclosures
  • smic yield improvements and 5nm/3nm progress11
  • deepseek model performance on ue8m0 vs standard hardware9
  • patent filings mentioning “8-bit exponent” or “microscaling”
  • ieee p3109 working group standards proposals
  • additional domestic chip announcements supporting ue8m0

evolving dynamics

the ue8m0 format and associated ecosystem represent one visible outcome of multi-year strategic decisions on both sides. what began as trade tensions in 2019 has evolved into technical divergence, with the august 2025 developments marking a new phase rather than an isolated event.

china’s approach - architectural divergence through format incompatibility - reflects constraints imposed since 2019 and investments made in response. the strategy optimizes for specific realities: persistent fabrication limitations,4 large domestic market, and independence imperatives reinforced by successive policy actions.

the strategic insight is the decoupling of training from inference: while cutting-edge model training may continue to benefit from nvidia’s superior hardware, the ue8m0 ecosystem enables china to deploy these models at scale for inference. this creates a sustainable path where model weights - whether developed domestically on nvidia hardware, trained through international collaborations, or obtained through other means - can be efficiently deployed on domestic infrastructure.

nvidia’s addition of ue8m0 support in ptx isa 9.03 suggests recognition that parallel ecosystems may be the new equilibrium, rather than temporary divergence.

future updates: this page will be updated as new information becomes available about technical developments, policy changes, and market evolution in the china-us ai hardware landscape.

references

[1] financial times. (2025, august 20). china turns against nvidia’s ai chip after ‘insulting’ howard lutnick remarks.

[2] deepseek ai. (2025, august 21). deepseek-v3.1 model card. hugging face.

[3] nvidia. (2025, august 1). parallel thread execution isa version 9.0.

[4] wccftech. (2024). smic to limit huawei to 7nm chips until 2026.

[6] investing.com. (2025, august 21). china’s deepseek upgrades ai model to support domestic chips.

[7] autogpt. (2025, august 21). deepseek launches new model with domestic chips.

[8] tom’s hardware. (2024). china’s moore threads polishes homegrown cuda alternative.

[9] deepseek ai. (2025). ue8m0(pr206) features cause severe regression issue. github issue #240.

[10] technode. (2024, november 15). chinese gpu unicorn moore threads files for ipo in china.

[11] granitefirm. (2025, march 8). how is smic after us embargo?

[12] bloomberg. (2025, june 12). amd ceo sees ai processor market exceeding $500 billion by 2028.

[13] globenewswire. (2024, october 28). ai chip market expected to reach usd 621.15 billion by 2032.

[14] china chip industry gets $47.5 billion in new funding | cnn business.

on this page