Ou Shisheng · AI4S Insights | The Data Dilemma and the Path to Self-Driven Innovation: The Rise of Self-Driven Laboratories

Published on: 2026-08-07 16:33
Category: Articles

AI Reconstructs the Scientific Paradigm: An Irreversible Historical Process

In October 2024, the Nobel Prize Committee made a historic decision—awarding the Nobel Prize in Physics to the pioneers of artificial neural networks and the Nobel Prize in Chemistry to a team of scientists who applied AI to protein structure prediction. This is the first time in the history of the Nobel Prize that two awards have been presented simultaneously in AI-related fields. This landmark event signals that AI is not merely a technological tool, but has become a core method for scientific discovery. This declaration did not come out of nowhere. Looking back over the past few years, the breakthroughs in AI within the scientific community have been truly astonishing:

In 2021, AlphaFold 2 dramatically reduced the time required for protein structure prediction, shortening what traditionally took months to years of experimental work to just a few hours. It had completed over 200 million protein structure predictions by that time, completely transforming the research paradigm in structural biology; In 2023, DeepMind’s GNoME system expanded its database of known stable materials from over 40,000 to 421,000, nearly a tenfold increase.

AI is reconstructing the logic of scientific discovery at an unprecedented pace.

China has not been left out of the transformation either. In August 2025, the State Council issued the “Opinions on Deepening the Implementation of the ‘AI+’ Initiative,” which identified “AI + Science and Technology” as a top priority and explicitly called for leveraging AI to accelerate “0-to-1” original innovation; In December of the same year, the proposal for the 15th Five-Year Plan further called for the full implementation of the “AI+” initiative, leveraging artificial intelligence to drive innovation in scientific research paradigms and secure a leading position in AI industrial applications. This dual boost from top-down policies cleared institutional barriers for the development of China’s AI4S industry, and the sector officially entered a phase of rapid growth.

The Data Infrastructure Challenge in AI4S: Bottlenecks in the Production of High-Quality Experimental Data

While AlphaFold 2 completes in a matter of hours what traditionally takes years in structural biology, and GNoME has expanded the number of known stable materials from over 40,000 to 420,000—these achievements seem to suggest that AI is ready to take over scientific discovery. But this optimism obscures a key prerequisite: data.

The performance of deep learning models fundamentally depends on the scale, quality, and structure of the training data. The success of AlphaFold 2 is largely attributable to a high-quality, standardized data infrastructure—Protein Data Bank (PDB), which has archived more than 170,000 experimentally validated protein structures, complete with detailed experimental conditions, resolution, and validation information. It is precisely this data foundation, built up over decades, that has unlocked the potential of these algorithms. However, in the fields of materials science—such as chemistry, materials science, and catalysis—there is no standardized data system comparable to the PDB. Research data faces two challenges:

First, there is a serious mismatch between traditional “peasant economy”-style data production and the industrial-scale demands of AI. Most papers report only “successful experiments,” while negative results, failed conditions, and subtle variations in synthesis parameters are systematically excluded. A catalyst research paper may report the optimal conversion rate, but it may not document key details such as the precipitation rate of the precursor, the shape of the stirrer, or environmental humidity—and these are precisely the latent variables required for AI modeling, meaning the data is incomplete from the source.

Second, the crisis of experimental reproducibility undermines the foundation of data credibility. The Nature survey found that more than 70% of researchers are unable to replicate others’ experiments, and more than 50% are unable to replicate their own experiments. If AI is “fed” large amounts of low-quality literature data, the model may learn noise rather than patterns, making it difficult for its predictions to guide real-world applications.

The deeper issue is that virtual data cannot replace real physical experiments. Data generated by first-principles calculations, molecular dynamics simulations, or generative models are, by their very nature, fits to existing physical models; they cannot cover unexplored chemical space, nor can they reveal non-ideal behavior, side reactions, or kinetic traps that occur under real-world conditions. Training an AI with simulated data to predict real experiments is like training a chess player using chess books without ever letting them touch a real chessboard—They might understand the theory, but they will not develop an intuitive understanding of the sport.

As a result, a long-overlooked fact is gradually coming into focus: the real bottleneck in AI4S lies neither in computing power nor entirely in algorithms, but in the scarcity of high-quality experimental data. Computing power can be procured, and models can be open-sourced, but structured, reproducible, and variable-controlled experimental data cannot be extracted from the literature nor generated in a virtual environment—it can only be produced at scale through real, standardized, high-throughput physical experiments.

This realization is shifting the focus of the global research community from “how to make AI more powerful” to “how to improve experimental data.” If AI is the engine, then experimental data is the fuel. Without high-quality fuel, even the most powerful engine cannot drive scientific progress. To bridge this gap, it is essential to establish a hardware infrastructure capable of producing high-quality, structured experimental data at scale—which is precisely the fundamental reason for the existence of high-throughput automated laboratories.

Explorations in the Global Academic Community: The Rise of Autonomous Laboratories

Research institutions around the world have long recognized the industry challenges caused by delays at the experimental level and have been continuously advancing the research and development of autonomous intelligent laboratory technologies. In 2025, a team led by Jiang Jun and Liu Daobin of the University of Science and Technology of China published a landmark review in Digital Discovery, systematically tracing the development of domestic autonomous laboratories and dividing the technological evolution into three distinct phases:

Phase 1: Iterative Algorithm-Driven Automation Platform (Starting in 2018)

In 2018, China’s first intelligent chemical robotics system, AIR-Chem, was unveiled by Zhu Xi’s team. It used a gradient descent algorithm to iteratively optimize the synthesis conditions for CsPbBr₃ quantum dots. Since then, the integration of high-throughput experiments and machine learning has accelerated even further. Fang et al. used liquid-core waveguide technology to construct a microfluidic photocatalytic reactor, enabling ultra-large-scale screening of up to 10,000 reactions per day. The robotic platform developed by Zhao et al. can automatically synthesize colloidal nanocrystals and use machine learning models to perform inverse design of nanocrystal morphology. This phase is characterized by the initial integration of algorithms into experiments, but the relationship between AI and the experiments remains loosely coupled.

Phase 2: Iterative Autonomous Experiments Driven by a Computational “Brain” (Starting in 2021)

Purely data-driven black-box optimization lacks systematic prior knowledge and has limited exploration efficiency. Combining first-principles calculations with machine learning has become a key breakthrough in enhancing interpretability and efficiency. The “All-in-One AI Chemist” (AI-Chemist) system developed by Jiang Jun’s team integrates three major modules—machine reading, mobile robotics, and a computational brain—to achieve a complete closed-loop process encompassing literature review → theoretical calculations → experimental planning → automated execution → data analysis → model training → and the generation of new solutions. Its extended applications are truly impressive: the system automatically synthesizes oxygen evolution reaction (OER) catalysts and, from over 3 million potential combinations, identifies the optimal catalyst formulation using only approximately 30,000 theoretical datasets and 243 experimental datasets, achieving stable operation for over 550,000 seconds at a current density of 10 mA cm⁻². This phase is characterized by AI evolving from a “support tool” to a “driving brain, forming a closed-loop discovery cycle of “prediction—preparation—measurement.”

Phase 3: Large-Model-Driven End-to-End Intelligent Autonomous Systems (Starting in 2023)

The rise of large language models (LLM) has opened up new possibilities for achieving truly end-to-end autonomous chemical research. The LLM-Research Development Framework (RDF) developed by Ruan et al., using aerobic alcohol oxidation as a case study, demonstrated the applicability of large language model agents throughout the entire end-to-end synthetic development process. Song et al.’s ChemAgents system goes a step further by building a hierarchical multi-agent architecture based on the Llama-3-70B large language model, supporting the coordinated scheduling of a literature database containing millions of entries, a library of 150 experimental protocols, two robots and 20 automated workstations, and 130 machine learning models. With this, AI-driven autonomous laboratories have moved from concept to reality.

Core Contradictions in the Industry: A Capability Gap Between Experimental Systems and AI Computing Power

Although the academic community has completed several generations of iterations in autonomous laboratory technology, the industrial sector and conventional research environments face a significant gap between AI computing capabilities and real-world laboratory operations. Large AI models and computing systems can rapidly generate vast numbers of candidate formulations, new material structures, and catalytic reaction pathways in bulk, often producing hundreds or even tens of thousands of sets of proposals for verification in a single run; however, traditional laboratories are completely unable to handle such large-scale experimental testing tasks.

In manual mode, only a small number of reactions can be conducted simultaneously in a single experiment. Reagent preparation, sample preparation, product detection, and data recording are all performed manually, and verifying dozens of AI-predicted protocols in a single batch can take weeks or even months; Human error resulting from manual operations, environmental interference, and inconsistent operational standards further lead to fragmented experimental data, low reproducibility, and a lack of standardization. As a result, a large number of high-quality predictive approaches generated by AI cannot be implemented and validated, and remain confined to the realm of theoretical simulation.

On one hand, AI computing is undergoing rapid iteration and continuously generating a vast number of candidate solutions; on the other hand, traditional experimental workflows rely on inefficient, fragmented, and low-throughput manual operations. The disparity in speed, throughput, and standardization between the two has created a massive gap. “Computing power is advancing rapidly, but experiments can’t keep up” has become a common pain point across the industry, directly hindering the transition of AI4S technology from laboratory research to large-scale industrial applications. Only by establishing high-throughput, automated, AI-driven laboratories capable of autonomous closed-loop operation—and addressing hardware shortcomings on the experimental side—can we establish a seamless end-to-end AI R&D pipeline.

High-Throughput Automated Experiment Platform: The Core Hardware Platform Enabling the Implementation of AI4S

In the context of AI4S, the core value of the experimental validation platform lies in the deep synergy across the four dimensions of “high-throughput, microscale, automation, and intelligence. High-throughput approaches, achieved through parallel reactor design and robotic cluster scheduling, have increased experimental throughput by 1–2 orders of magnitude—the “Intelligent Scientist” system at the University of Science and Technology of China performs up to 2,000 precise operations per day, enabling AI to explore high-dimensional spaces within a reasonable time frame. Microscale technology utilizes high-precision fluid manipulation techniques—such as piezoelectric inkjet dispensing, acoustic droplet transfer, and microfluidic chips—to reduce reaction volumes to the microliter or even nanoliter range, cutting consumption to 1/100 to 1/1000 of that of traditional methods, thereby making ultra-high-throughput screening feasible both physically and economically. Automation uses closed-loop motion control and constant operating sequences to maintain experimental conditions within a narrow tolerance range, thereby eliminating data noise introduced by human error. The intelligent system uses data-driven control algorithms to perform autonomous regulation across the entire process chain, eliminating the need for manual intervention throughout the process. This ensures safe and orderly closed-loop operation of the equipment while simultaneously enabling the complete archiving and traceable retention of raw process data across all dimensions. These four elements are interlinked: microscale processing reduces the cost per run, making high-throughput sustainable; high-throughput generates massive amounts of data, providing the statistical foundation for intelligent systems; intelligent systems enable secure, orderly, and autonomous regulation of the entire process while simultaneously archiving comprehensive, multi-dimensional process traceability data; and automation frees up human resources while ensuring high-quality data.

Once these capabilities have been established, the key challenge facing the platform is how to apply these general capabilities to specific and diverse chemical systems. Real-world research scenarios involve harsh conditions such as high temperature and pressure, severe corrosion, and multi-field coupling; high-throughput automated platforms serve as “technical tool boxes” deeply customized for different disciplinary paradigms.

The ultimate goal of platform development is to integrate the entire process—from “sample preparation to reaction, characterization, and performance testing”—and enable real-time data feedback; this is the fundamental feature that distinguishes it from traditional automated equipment. The data required for AI4S must not only be “voluminous,” but also “accurate” and “complete.” Fragmented data consisting solely of final results has limited value for model training. The complete end-to-end process begins with sample preparation, proceeds through reaction execution, moves on to characterization, and concludes with performance testing; each stage is automatically linked to the next via standardized containers and robotic arms. Both the “Self-Driven Scientific Experiment System” developed by the Shenyang Institute of Automation, Chinese Academy of Sciences and the “Materivo” platform developed by the University of Science and Technology Beijing have successfully implemented this complete workflow. Researchers only need to define the objectives and parameter space within the software, and the platform will automatically handle the entire process from raw materials to the final report.

In summary, the high-throughput automated experimental platform delivers foundational capabilities that surpass those of human operators through the synergy of the “Four Key Pillars.” Through a differentiated architectural design, it adapts to complex chemical systems; and through end-to-end integration and real-time feedback, it serves as a “key node in the data closed-loop”—these three elements build upon one another, collectively forming a robust bridge from AI computing power to experimental validation.

Building an AI Self-Driving Lab: From Hardware Integration to Autonomous Iteration

The essence of an AI-Driven Laboratory is not simply to assemble automated equipment and AI algorithms, but to establish a new paradigm for scientific research capable of autonomously generating hypotheses, intelligently designing and validating experiments, and performing closed-loop iterative optimization. Its core architecture consists of three major modules—an AI-powered decision-making system, a high-throughput automated experimental platform, and a private database—which form a closed-loop, coupled relationship of “design–execution–evaluation–redesign,” enabling the laboratory to continuously generate high-quality experimental data and scientific insights with minimal human intervention.

The AI-powered decision-making system serves as the cognitive and decision-making hub of the entire self-driven laboratory. Its core mission is to accurately identify the intrinsic mapping relationship between “synthesis parameters and performance metrics” within a multidimensional process parameter space, and to conduct global intelligent optimization based on this relationship. Specifically, the system uses regression and classification models—such as XGBoost, neural networks, and random forests—to model and train the process parameters and performance metrics in the database, thereby establishing reliable performance prediction models. It subsequently employs global optimization strategies—including Bayesian optimization, genetic algorithms, and particle swarm optimization—to automatically search within the multidimensional parameter space and output the optimal formulation and preparation process parameters that meet the performance objectives.

The high-throughput automated experimental platform serves as the physical executor. Leveraging high-throughput material preparation and performance evaluation equipment, it enables automated parallel preparation and performance testing of materials, allowing for simultaneous batch-scale experimental validation. The platform is capable of precisely converting the parameter vectors output by the AI decision-making system—including reagent types, dosages, and preparation process conditions—into a continuous sequence of physical operations, while simultaneously collecting multimodal sensor data throughout the entire experiment to rapidly generate complete, reproducible, and comparable high-quality experimental data. We continuously iterate and optimize AI decision-making models using vast amounts of standardized experimental data, thereby continuously enhancing the predictive and simulation capabilities of the intelligent system.

The private database forms the knowledge repository for the entire laboratory. It combines experimental data with AI-driven literature data collection to convert experimental methods into standardized feature parameters—such as reagent type, dosage, and process conditions—thereby creating a structured database of preparation process parameters. This database undergoes dynamic, iterative upgrades through continuous experimental feedback. Each time new data from an experiment is transmitted back, the database is continuously enriched, providing ongoing support for model updates. Real-world, proprietary data generated using  high-throughput equipment constitutes a company’s core digital asset, offering significant technical and economic value: it can substantially reduce consumables and labor costs associated with repetitive experiments, while shortening new product development cycles; Exclusive coverage of extreme and proprietary process data builds competitive barriers that are difficult for peers to replicate; comprehensive traceability records support patent applications and compliance reviews, protecting the company’s intellectual property; and a unified, standardized set of measured samples continuously drives the iteration of intelligent algorithms, continuously improving process optimization efficiency and amplifying the long-term return on investment in digital R&D. This database undergoes dynamic, iterative upgrades through continuous experimental feedback. Each time new data from an experiment is transmitted back, the database is continuously enriched, providing ongoing support for model updates.

The interaction among the three major modules forms the core operational logic of the self-driven laboratory. During the cold start phase, the AI decision-making system generates the first set of experimental designs covering the parameter space, which are then executed by a high-throughput automated platform for batch validation, and the resulting data is stored in a private database. Thereafter, each new experimental record triggers an incremental update to the machine learning model—the regression and classification models are retrained on the expanded dataset, thereby improving prediction accuracy; subsequently, the global optimization algorithm conducts a new round of optimization based on the updated models, outputs a new optimal formulation, and submits it once again to the automation platform for validation. This closed-loop cycle of “experiment—data—modeling—optimization—re-experiment” operates continuously without human intervention. If the real-time data transmitted by the sensors during the process triggers an anomaly, the system can bypass the preset protocol to adjust parameters in real time or terminate the reaction early, thereby preventing invalid data from contaminating model training.

After integrating the three major modules into a complete system, the fundamental differences between the self-driven labs and traditional labs are evident on three levels: In terms of decision-making, researchers have shifted from being “operators” to “goal definers,” with the algorithm autonomously handling all aspects of experimental design and parameter adjustment; In terms of knowledge accumulation, research experience is no longer tied to individuals but is encoded in databases and model parameters, growing systematically as the number of experimental runs increases; in terms of time scale, the traditional iteration cycles—measured in weeks or months—have been compressed to the hourly level.

Looking back at the development of the AI4S high-throughput automated laboratory, a clear evolutionary thread runs throughout: From the early explorations of AI-assisted scientific research, to the emergence of data bottlenecks, and on to the rise of autonomous laboratories, the core contradiction has consistently centered on the gap between experimental capabilities and AI computing power. This gap is gradually being bridged by high-throughput automated experimental platforms, and the establishment of AI-driven laboratories marks a shift from “replacing humans with machines” to a new paradigm where “AI autonomously drives scientific discovery.” As foundational models and open ecosystems continue to evolve, AI-driven laboratories are poised to become the core form of next-generation scientific research infrastructure, shifting scientific discovery from a “human-led” approach to “human-machine collaboration” and redefining the efficiency and boundaries of exploring the unknown.

Share
  • toolbar
  • toolbar
  • toolbar
  • toolbar