Latest News

sciencenews.png

"Genome map" of Japanese and Saudi Arabian population created by National Institute of Global Health and Medicine

2025.10.08

Deputy Project Manager Yosuke Kawai, Specially Appointed Researcher Saeideh Ashouri, and Project Director Katsushi Tokunaga of the Genome Medical Science Project at the National Institute of Global Health and Medicine, together with King Abdullah University of Science and Technology (KAUST) (Saudi Arabia), the Database Center for Life Science, Kyushu University, the University of Tokyo, and the National Institute of Genetics, announced the determination of complete genome sequences of Japanese and Saudi Arabians and their registration in public databases. They also developed a pangenome graph called "JaSaPaGe" that reflects the genomic diversity of both countries and enables high-precision detection of genetic variants. This serves as a foundation for region-specific precision medicine and genetic diagnosis and is expected to advance more accurate and equitable genome research that includes Japan and Middle Eastern regions. The results were published in Scientific Data on August 12.

In human genome analysis, specific individual genome sequences have been used as reference sequences, presenting challenges in reflecting structural variants with large individual differences and population-specific genetic information.

In recent years, pangenome graphs have emerged as a new data representation method that comprehensively handles and compares genome information from multiple individuals. While international efforts to create pangenome graphs are underway, these initiatives include very few samples from Japanese or Arab populations, leaving populations corresponding to approximately 600 million people worldwide excluded from the "genome map."

Therefore, the research group aimed to provide a foundation for more accurate genome analysis rooted in these regions by constructing a unique pangenome graph that reflects the genomic diversity of Japanese and Saudi Arabian populations.

Based on high-quality genome data obtained from 10 Japanese and 9 Saudi Arabian individuals, they performed "diploid genome assemblies" containing sequences derived from two sets of chromosomes. Based on these, they created a pangenome graph consisting of Japanese sequences, Saudi Arabian sequences, and conventional reference sequences.

When using these for variant calling from short-read sequences, it was found that approximately 3-10% more genetic variants could be detected compared with the use of conventional reference sequences.

Furthermore, it was demonstrated that population-specific structural polymorphisms could be clearly identified in the copy number variation (CNV) of CYP2D6, an important gene involved in drug metabolism.

Data obtained from the research has been registered in the DNA Data Bank of Japan (DDBJ) and is freely available for use and reanalysis by anyone.

Kawai commented: "Using the precise genome data released through this study, it will be possible to analyze regions missing from conventional references (such as centromeres and repeat regions) and population-specific sequences, and we expect advances in discovering pathogenic variants and researching rare diseases. We plan to proceed with applied research using this data in the future."

Journal Information
Publication: Scientific Data
Title: Phased genome assemblies and pangenome graphs of human populations of Japan and Saudi Arabia
DOI: 10.1038/s41597-025-05652-y

This article has been translated by JST with permission from The Science News Ltd. (https://sci-news.co.jp/). Unauthorized reproduction of the article and photographs is prohibited.

Back to Latest News

Latest News

Recent Updates

    Most Viewed