This track shows segmental duplications in this Human Pangenome Reference Consortium (HPRC) Release 2 assembly. Segmental duplications are large blocks of genomic sequence (typically longer than 1 kb) that occur in two or more nearly identical copies. They are hotspots for recurrent structural rearrangements and harbour many genes that have expanded in the human lineage, so their exact positions differ from person to person and are of particular interest in a pangenome. Each entry marks one copy of such a block together with the paralogous copy it aligns to.
Each item is one segmental-duplication region. The item is labelled with, and links conceptually to, the paralogous region it aligns to (shown as the partner field on the details page, in the form sequence:start-end). The fraction of identical bases in the alignment, the alignment length, and other per-duplication measures from the source are listed on the details page. The score is the fraction identity scaled to 0–1000.
Segmental duplications were detected with SEDEF, which finds pairs of homologous genomic segments by seeding on shared k-mers and extending and refining the alignments, then reports each duplicated segment pair with a set of alignment statistics (see reference below). The calls were produced by the Eichler laboratory as part of the HPRC assembly annotation.
The annotation files were obtained from the HPRC Release 2 data collection on the public s3://human-pangenomics bucket, indexed at the hprc_intermediate_assembly data tables. The per-assembly SEDEF output was converted to a UCSC bigBed file, keeping the region coordinates, the paralog partner, and the main alignment statistics. The build scripts are in the kent source tree.
For automated analysis, the annotation is stored in a bigBed-format file (segdups.bb) that can be read with the UCSC tool bigBedToBed. The original files are available from the HPRC S3 bucket linked above.
Annotations were generated by the Human Pangenome Reference Consortium and the Eichler laboratory. Thanks to the HPRC production team for making these data available.
Numanagic I, Gökkaya AS, Zhang L, Berger B, Alkan C, Hach F. Fast characterization of segmental duplications in genome assemblies. Bioinformatics. 2018 Sep 1;34(17):i706-i714. PMID: 30423092; PMC: PMC6129265