1 Comment
User's avatar
Leon Liao's avatar

Data infrastructure could become one of China’s most consequential advantages in the global AI competition.

I strongly agree with the argument made in this discussion. The United States is moving to restrict both Chinese open-weight models and cross-border data flows without providing American businesses with comparably affordable models, accessible public datasets or institutions that facilitate data circulation. It may protect a small number of domestic model providers while raising the cost of AI adoption across the rest of the economy and sacrificing long-term productivity gains. The logic resembles the 100% tariff used to protect America’s automobile industry: US companies will be pushed toward expensive domestic closed models, while competitors in Europe, Asia and the Global South continue using Chinese models that offer comparable performance at a fraction of the price.

Early foundation models could be trained largely on public text scraped from the internet. Industrial AI will require smaller quantities of much higher-quality data that are difficult and increasingly expensive to collect, label and verify—sometimes costing $30–50 per hour of specialized human work. Medical records, transport networks, public services, factory operations and scientific research produce precisely the datasets needed for the next generation of vertical models. Individual startups usually cannot generate such data themselves and may struggle even to discover or purchase it. China’s public-data initiatives, industry data pools, exchanges, rights registries and compliance systems are designed to increase this effective supply, lowering training costs and accelerating AI development in healthcare, manufacturing, transportation and energy.

China’s low-cost open models and its data-factor strategy could therefore reinforce each other. Affordable models lower the threshold for adoption; broader access to industrial data improves model quality and specialized applications; widespread deployment then generates another round of valuable data. If this cycle takes hold, China’s AI competitiveness will rest on more than chips, compute and model parameters. It will also draw strength from the breadth of its industrial economy—and from an institutional system increasingly designed to convert that economy’s data into a shared input for innovation.