Overcoming the Terascale Design Challenge: Next-Generation Efficient AI for Discoveries in Surface Science and Catalysis 

About the programme 

Every industrial process relies on catalysts. They are the materials that make chemical reactions faster, more efficient, and less energy-intensive. Our goal is to transform catalyst discovery from a slow, linear process into an adaptive, autonomous and self-improving cycle which could transform how we produce energy, reduce industrial emissions and build a lower-carbon economy. 

AI for Catalyst Discovery is building a surface science foundation model –  AI trained on an extensive dataset of surface structures, interfaces and surface molecular interactions across a wide range of inorganic compounds. 

This foundation enables us to develop resource-efficient AI models to predict material properties and catalytic performance at a speed and scale that conventional methods cannot match. 

The challenge 

Materials discovery has always been slow. Researchers test one candidate at a time, building knowledge incrementally through years of laboratory work. Catalyst discovery through traditional trial and error experimentation cannot keep pace with the number of possible material compositions, structures and synthesis conditions. Materials can yield billions of possible configurations, and development cycles often span more than a decade. 

While computational methods have helped, even the most powerful simulations are limited by the sheer scale of the design space. Existing databases of catalytic materials are largely built on idealised, simplified models that do not reflect the complexity of real operating conditions. 

For example, they do not account for the supports, promoters, defects, and dynamic surface behaviour that determine how a catalyst performs in the real world. 

We need a fundamentally different approach, and the impacts could make a significant difference to reducing emissions and energy usage. 

Our research streams 

Our programme is structured around five interconnected research streams, spanning computational data generation, AI model development, and experimental validation. 

Stream 1: Building a Representative Dataset 

Existing databases of catalytic materials contain around 1.2 million entries – but they are built almost entirely on idealised, perfect crystal structures. Real catalysts are messier and more complex: they have supports, promoters, defects and surfaces that restructure dynamically under operating conditions. 

Our programme stream generates over one million additional data points that capture these realistic catalyst environments, using high-throughput quantum mechanical calculations integrated with the National Supercomputing Centre Singapore. 

We build on two powerful computational workflow platforms: CatPlat, developed by our A*STAR partner, and Atomate2, for which our ³Ô¹ÏºÚÁÏ team serves as lead developer. 

Stream 2: The Surface Science Foundation Model 

Our second stream develops a multi-fidelity surface science foundation model. This is an AI system that learns complex relationships across datasets of varying quality and resolution.  

A key innovation is our Equivariant Sparse Structure Learning framework, which allows the model to identify the most relevant structural features for a given prediction task while filtering out noise. 

Our preliminary results show this approach can negate up to 80 per cent of redundant information at each layer of the model, while maintaining or improving predictive accuracy. The target is a model that runs 10 to 50 times faster than conventional methods. 

Stream 3: Agentic AI for Inverse Design 

Once we have a foundation model, we can use it to design new materials. 

This stream develops generative AI models, reinforcement learning, and Bayesian optimisation tools that autonomously propose candidate catalyst structures meeting multiple performance criteria simultaneously. 

These tools operate within an agentic AI workflow – a closed-loop system that continuously coordinates structure generation, evaluation, and model refinement in response to evolving design goals. The result is a discovery pipeline that we expect to accelerate design space coverage by more than 100 times relative to traditional static approaches. 

Stream 4: High-Throughput Catalyst Synthesis 

Our fourth stream experimentally synthesises the catalyst candidates proposed by Streams 2 and 3,  focusing on reactions that directly support the transition to a more sustainable and lower-carbon economy. These include the conversion of carbon dioxide into methanol, reactions using bio-based feedstocks such as glycerol, and more complex organic transformations involving ethanol and propanol. Synthesis is carried out using automated robotic platforms at A*STAR's Institute of Sustainability for Chemicals, Energy and Environment (ISCE2) and the Institute of Materials Research and Engineering (IMRE). Advanced characterisation techniques allow us to observe how catalysts behave under real reaction conditions. 

Stream 5: High-Throughput Catalytic Testing 

Catalytic performance is screened and tested in the at A*STAR's Institute of Sustainability for Chemicals, Energy and Environment (ISCE²). 

Catalyst performance, including Kinetic parameters extracted from these experiments, are fed back into the computational models, continuously improving the foundation model's predictve power. 

This closes the loop between computation and experiment, accelerating catalyst discovery at unprecedented scale.