This repository provides workflows and analysis scripts to build a structural variant control set from 1,725 long-read genomes and to use it for phenome-wide association studies in biobanks.

Stars

3

7-day growth

No data

Forks

0

Open issues

0

License

No data

Last updated

2026-07-29

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It combines a large, multi-platform long-read SV resource (HPRC, HGSVC, UW-ONT, IB-ONT) with rigorous multi-caller integration and demonstrates utility by identifying SV-disease associations in 230,000 All of Us participants.

Who it is for

  • Researchers studying structural variants in human genomes
  • Bioinformaticians building SV callsets or genotyping pipelines
  • Clinical genomics groups seeking rare SV-disease associations
  • Biobank analysts using long-read references to interpret short-read data

Use cases

  • Creating a high-quality SV control set from diverse long-read sequencing data
  • Benchmarking and combining multiple SV callers (assembly-based and read-based)
  • Genotyping SVs in large short-read cohorts to discover disease associations
  • Fine-mapping causal SVs with credible set analysis using SuSiE

Strengths

  • Large sample size (1,725 genomes) across multiple long-read platforms (HiFi, ONT R9/R10)
  • Multi-caller integration per genome followed by cohort-level collapse with Truvari (≥90% sequence/size similarity)
  • Provides both GRCh38 and T2T-CHM13 callsets with estimated allele frequencies
  • Demonstrated real-world use: 2,752 high-impact SVs genotyped in 230K AoU samples for PheWAS

Considerations

  • Some data download links (e.g., Zenodo, index file) are marked as TBA (to be added)
  • Focuses only on insertions and deletions 50 bp – 100 kbp; other SV types (inversions, duplications) are not included
  • Missing genotypes are imputed as reference based on callable regions, which may introduce bias in rare variants

README quick start

1KG_LongRead_SV

This repository contains workflows and analysis scripts for creating a SV control set from 1725 1KG long-read genomes. These genomes are sequenced by HPRC (Lucas et al. Biorxiv 2026), HGSVC (Logsdon et al. Nature 2025), UW-ONT and IB-ONT.

The UW-ONT sequenced a total of 500 genomes, 400 are novel and 100 genomes were published by Gustafson et al. 2025. The IB-ONT genomes are published by Siegfried Schloissnig et.al. (Nature 2025). The SV control set has been used to identify potential pathogenic variants and new SV-disease associations in biobanks.

Genomes

The table below lists the source of 1,218 genomes used to create the callset in Lin et al. Note that the number below counts for the unique genomes.

HG002 and HG005 are included as part of the HPRC dataset. We also included NA12877 and NA12878. NA12878 is also one of the genomes sequenced by GIAB and HGSVC. NA12877 is sequenced by IB-ONT but we used the assembly and data published in David Porubsky et al. Nature 2025

DatasetGenomesPlatform
HPRC232HiFi, UL-ONT
HGSVC61HiFi, UL-ONT
1KG-ONT480ONT (R9, R10)
IB-ONT445ONT (R9)

SV callset

Individual genomes

Multiple caller merged callset for each genome. Please check the index file (TBA) for download.

Integrated callset

We provide both GRCh38 and T2T-CHM13 callsets (zenodo link, TBA) for 293 HPRC+HGSVC genomes and 1,218 genomes.

CHM13_INSDEL_HGSVC_HPRC_wAF.vcf.gz: Integrated SVs from HGSVC/HPRC genomes with estimated allele frequency.

CHM13_INSDEL_1218_wAF.vcf.gz: Integrated SVs from all dataset containing 1,218 genomes with estimated allele frequency.

GRCh38_INSDEL_HGSVC_HPRC_wAF.vcf.gz: Integrated SVs from HGSVC genomes with estimated allele frequency.

GRCh38_INSDEL_1218_wAF.vcf.gz: Integrated SVs from all dataset containing 1,218 genomes with estimated allele frequency.

Genome assembly

HPRC

The HPRC genomes are assembled with hifiasm and verkko. Please refer to HPRC [release](https://gith

Description

The repository for 1KG long-read sequencing based SV discovery

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482