A research team including Professor Akira Nakayama and Associate Professor Koki Muraoka of the Graduate School of Engineering at the University of Tokyo has developed "Graph ID," a breakthrough identifier capable of integrating vast materials databases scattered across the globe. This development is expected to enable global-scale data integration and deduplication in the search for new materials, such as storage batteries, catalysts, and semiconductors, thereby dramatically boosting development speeds. The findings were published in Nature Communications.
In recent years, the spread of high-throughput computing technologies has led to the daily generation of massive amounts of structural data, including unknown materials. While these data are accumulated in international databases such as the Materials Project and AFLOW, each platform maintains its own proprietary management system. Consequently, it has been difficult to immediately determine "whether a material registered in one database is identical to a material in another."
To bridge this gap, experts have traditionally relied on manual naming. However, processing big data exceeding millions of entries is practically impossible. Furthermore, existing automated naming methods suffer from limitations. Minor numerical errors or discrepancies in coordinate system selections often cause identical structures to be flagged as distinct, or conversely, cause subtly different structures to be erroneously conflated.
The "Graph ID" developed by the research team resolves these issues by treating chemical structures as mathematical graphs. Graph ID iteratively analyzes the surrounding environment of each individual atom to generate a unique hash string that functions as a structural fingerprint.
Validation tests demonstrated that Graph ID possesses outstanding characteristics, including high precision, rapid processing speed, and versatility. It can accurately identify complex crystal structures and surface structures containing adsorbed molecules, both of which have been difficult to distinguish using conventional symmetry-based methods. Additionally, the computational cost required to match new structures within a database is minimal, achieving a massive speedup compared to traditional pairwise comparison methods. The application is not limited to crystals but it is adaptable to a broad spectrum of chemical structures, including molecules and surface interfaces.
Furthermore, using Graph ID to perform an integrated analysis of three of the world's largest materials databases (Materials Project, AFLOW, and OQMD), the team successfully identified overlapping materials across the different platforms. This breakthrough makes it possible to construct cross-database, unified datasets.
To establish this technology as a common foundation for the wider scientific community, the research team has released the program code that generates Graph ID as open-source software. Along with the code, they have published a database featuring IDs assigned to more than 150,000 known structures. Moving forward, as this common identifier becomes widely adopted as a standard registration system for materials data, it is expected to accelerate AI-driven novel materials prediction and the creation of platforms where researchers worldwide can seamlessly share insights.
Journal Information
Publication: Nature Communications
Title: Universal graph-based identifiers of chemical structures for linking large material databases
DOI: 10.1038/s41467-026-74536-5
This article has been translated by JST with permission from The Science News Ltd. (https://sci-news.co.jp/). Unauthorized reproduction of the article and photographs is prohibited.

