MARATTO

preprint

Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation

202441 citationsOpen accessUniversity of the Witwatersrand

In plain language

Standard clinical genetic tests fail to provide a definitive molecular diagnosis for more than half of individuals suspected of having a Mendelian disorder. While long-read sequencing can capture complex genetic variation, the lack of reference control datasets has restricted its use for variant filtering in diagnostics. To resolve this bottleneck, nanopore sequencing was applied to samples from diverse global populations within the 1000 Genomes Project. Analysis of the initial one hundred genomes detected tens of thousands of structural variants per individual, successfully uncovering pathogenic repeat expansions and gene-disrupting alterations that short-read methods miss. The dataset also profiles base modifications, identifying expected and novel DNA methylation patterns across the genome. This publicly accessible catalog provides a critical baseline of standard human genetic variation to support clinical genomic testing and disease research.

Key takeaways

  • Long-read nanopore sequencing of the first one hundred diverse reference samples revealed an average of 24,543 high-confidence structural variants per genome.
  • The approach captured pathogenic repeat expansions and functional gene disruptions that were previously undetected by short-read sequencing.
  • The dataset maps both genetic variations and methylation signatures, highlighting imprinted loci and novel differentially methylated regions.

Why it matters

Identifying the genetic causes of rare inherited disorders remains difficult with standard technologies. By creating a comprehensive reference catalogue of structural and epigenetic variations across diverse human populations, this work equips medical geneticists with the baseline data needed to distinguish benign natural variations from disease-causing mutations.

Commercialisation angle

This research provides reference data to advance clinical genomics and diagnostic test development. The direct users are diagnostic laboratories, clinical genetics services, and bioinformatics tool developers who require baseline datasets to filter and prioritise variants. As an open-access foundational resource with early data already public, it is immediately usable for tertiary analysis pipelines, though clinical diagnostic integration remains an ongoing, applied process.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Less than half of individuals with a suspected Mendelian condition receive a precise molecular diagnosis after comprehensive clinical genetic testing. Improvements in data quality and costs have heightened interest in using long-read sequencing (LRS) to streamline clinical genomic testing, but the absence of control datasets for variant filtering and prioritization has made tertiary analysis of LRS data challenging. To address this, the 1000 Genomes Project ONT Sequencing Consortium aims to generate LRS data from at least 800 of the 1000 Genomes Project samples. Our goal is to use LRS to identify a broader spectrum of variation so we may improve our understanding of normal patterns of human variation. Here, we present data from analysis of the first 100 samples, representing all 5 superpopulations and 19 subpopulations. These samples, sequenced to an average depth of coverage of 37x and sequence read N50 of 54 kbp, have high concordance with previous studies for identifying single nucleotide and indel variants outside of homopolymer regions. Using multiple structural variant (SV) callers, we identify an average of 24,543 high-confidence SVs per genome, including shared and private SVs likely to disrupt gene function as well as pathogenic expansions within disease-associated repeats that were not detected using short reads. Evaluation of methylation signatures revealed expected patterns at known imprinted loci, samples with skewed X-inactivation patterns, and novel differentially methylated regions. All raw sequencing data, processed data, and summary statistics are publicly available, providing a valuable resource for the clinical genetics community to discover pathogenic SVs.

Research topics

  • Genomics and Rare Diseases
  • Genetic Syndromes and Imprinting
  • Genomic variations and chromosomal abnormalities

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1101/2024.03.05.24303792

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.