The Tohoku Medical Megabank (TMM) Project has significantly expanded the data available for research distribution, including whole-genome information from 69,000 individuals. A wide range of datasets such as genomic data, omics data, long-term health follow-up surveys, and brain MRI images are now accessible for research. This expansion is expected to accelerate progress in personalized medicine, preventive medicine, and drug discovery.
Executive Director Hideo Harigae of the Tohoku Medical Megabank Organization stated "Over the past 15 years, we have accumulated data to become a valuable biobank that includes health surveys. In the future, we want to facilitate various forms of utilization so that this data can serve society."
TMM is Japan's largest biobank, targeting approximately 150,000 general citizens. It has provided a foundation for life-course research, exploratory studies on disease causes and biomarkers, and sex-specific medicine. Data distribution began in 2015 and has been enriched annually. In the latest update, whole-genome analysis data has expanded dramatically from 15,000 to 69,000 individuals.
SNP array data has increased by 14,000 individuals, bringing the total to 120,000, establishing it as a world-class genomic resource. Human Leukocyte Antigen (HLA) polymorphism data, previously unavailable, has now been added for 54,000 individuals.
Because HLA is closely involved in immune response and drug adverse reactions, this large-scale dataset is expected to accelerate research across many diseases. In addition, metabolome data has expanded by 8,000 individuals, reaching a total of 69,000. Whole-blood transcriptome data has also newly become available for 576 individuals.
Microbiome datasets have also expanded, now including oral microbiome data from 2,991 individuals and tongue coating microbiome data from 400 individuals. Longitudinal health survey data now covers both the baseline survey period of the community-based cohort and the three-generation cohort study (May 2013 to March 2017), as well as the second-stage survey period (April 2017 to March 2021).
Follow-up questionnaire data has also been added for the second stage of the community-based cohort, enabling more detailed tracking of health changes over time. The scope of laboratory tests and dental data, previously limited to individuals aged 20 and older, has been expanded to all age groups, including participants under 20.
For the three-generation cohort, new data from infancy, school entry, and school health checkups has been added, strengthening life-course research from the fetal stage through adulthood. Additional data now includes disease onset questionnaires for Kawasaki disease and congenital anomalies, as well as adjudicated onset data for stroke, myocardial infarction, and angina pectoris, reviewed by cardiovascular specialists across both survey stages.
The coverage of administrative data such as medical receipts and long-term care insurance records has also been significantly expanded from the first to the second stage.
Regarding brain MRI data from the brain and mental health survey, DICOM-format images from the second-period survey (October 2019 to March 2024, approximately 7,400 individuals) have been added to the first-period dataset (July 2014 to October 2019, approximately 12,000 individuals). In addition, derived imaging measures such as brain region volumes processed using FreeSurfer are also provided for the second-period cohort.
To improve usability and support integrated analysis, the project has launched "dbTMM2026." This database consolidates 33 previous separate releases into a single system and significantly expands the range of distributable data.
To help researchers navigate the expanded resources, more than 50 types of information have been organized into 9 categories and over 20 subcategories, enabling category-based data access.
With this expanded genomic and omics resource, integrated analyses including longitudinal changes will make it possible to trace health trajectories across diverse genetic backgrounds. This is expected to help identify genetic variants and lifestyle factors associated with disease onset and further refine genomic medicine tailored to the Japanese population.
In addition, multi-layered omics data combined with Mendelian randomization analysis is expected to support the identification of novel drug discovery targets.
Finally, the dataset is expected to advance personalized prevention across the life course from fetal development to old age, supporting research such as disease prevention using the three-generation cohort and dementia prevention using large-scale brain imaging data from approximately 12,000 individuals.
This article has been translated by JST with permission from The Science News Ltd. (https://sci-news.co.jp/). Unauthorized reproduction of the article and photographs is prohibited.

