该仓库提供了从一个包含1,725个长读长基因组的结构变异控制集的构建流程和分析脚本,并用于生物库的全表型组关联研究。

Stars

3

7 天增长

暂无数据

Fork 数

0

开放 Issue

0

开源协议

暂无数据

最近更新

2026-07-29

AI 仓库情报摘要
FR-AI / ANALYSIS

为什么值得关注

它结合了多种长读长平台(HPRC、HGSVC、UW-ONT、IB-ONT)的大规模结构变异资源,通过严格的多检测器整合,并在23万名All of Us参与者中展示了识别SV-疾病关联的实用价值。

适合谁使用

  • 研究人类基因组结构变异的科研人员
  • 构建SV调用集或基因分型管道的生物信息学家
  • 寻求罕见SV-疾病关联的临床基因组学团队
  • 使用长读长参考解读短读长数据的生物库分析人员

典型使用场景

  • 从多种长读长测序数据构建高质量结构变异控制集
  • 基准测试并整合多个结构变异检测器(基于组装和基于读长)
  • 在大规模短读长队列中对结构变异进行基因分型以发现疾病关联
  • 使用SuSiE进行可信集分析实现因果结构变异的精细定位

项目优势

  • 跨多种长读长平台(HiFi、ONT R9/R10)的大样本量(1,725个基因组)
  • 每个基因组进行多检测器整合,随后使用Truvari进行队列级合并(≥90%序列/大小相似性)
  • 提供GRCh38和T2T-CHM13两个参考基因组的调用集并附带等位基因频率
  • 实际应用示例:2,752个高影响结构变异在23万个AoU样本中进行了基因分型并开展全表型组关联分析

使用前须知

  • 部分数据下载链接(如Zenodo、索引文件)标记为待添加(TBA)
  • 仅关注50 bp至100 kbp的插入和缺失,不包括其他类型结构变异(倒位、重复)
  • 缺失基因型基于可调用区域被填充为参考基因型,可能对稀有变异引入偏倚

README 快速开始

1KG_LongRead_SV

This repository contains workflows and analysis scripts for creating a SV control set from 1725 1KG long-read genomes. These genomes are sequenced by HPRC (Lucas et al. Biorxiv 2026), HGSVC (Logsdon et al. Nature 2025), UW-ONT and IB-ONT.

The UW-ONT sequenced a total of 500 genomes, 400 are novel and 100 genomes were published by Gustafson et al. 2025. The IB-ONT genomes are published by Siegfried Schloissnig et.al. (Nature 2025). The SV control set has been used to identify potential pathogenic variants and new SV-disease associations in biobanks.

Genomes

The table below lists the source of 1,218 genomes used to create the callset in Lin et al. Note that the number below counts for the unique genomes.

HG002 and HG005 are included as part of the HPRC dataset. We also included NA12877 and NA12878. NA12878 is also one of the genomes sequenced by GIAB and HGSVC. NA12877 is sequenced by IB-ONT but we used the assembly and data published in David Porubsky et al. Nature 2025

DatasetGenomesPlatform
HPRC232HiFi, UL-ONT
HGSVC61HiFi, UL-ONT
1KG-ONT480ONT (R9, R10)
IB-ONT445ONT (R9)

SV callset

Individual genomes

Multiple caller merged callset for each genome. Please check the index file (TBA) for download.

Integrated callset

We provide both GRCh38 and T2T-CHM13 callsets (zenodo link, TBA) for 293 HPRC+HGSVC genomes and 1,218 genomes.

CHM13_INSDEL_HGSVC_HPRC_wAF.vcf.gz: Integrated SVs from HGSVC/HPRC genomes with estimated allele frequency.

CHM13_INSDEL_1218_wAF.vcf.gz: Integrated SVs from all dataset containing 1,218 genomes with estimated allele frequency.

GRCh38_INSDEL_HGSVC_HPRC_wAF.vcf.gz: Integrated SVs from HGSVC genomes with estimated allele frequency.

GRCh38_INSDEL_1218_wAF.vcf.gz: Integrated SVs from all dataset containing 1,218 genomes with estimated allele frequency.

Genome assembly

HPRC

The HPRC genomes are assembled with hifiasm and verkko. Please refer to HPRC [release](https://gith

项目描述

The repository for 1KG long-read sequencing based SV discovery

相关仓库与替代方案

根据分类、Topic 和编程语言匹配的相似项目。

lopopolo
精选
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI 与机器学习AI 智能体
2,390
slvDev
精选
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI 与机器学习大语言模型
1,960
littledivy
精选
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI 与机器学习
1,482