INTRODUCTION
Methods for partitioning single nucleotide polymorphisms (SNPs) into blocks followed by haplotype analysis are among the approaches increasing the power of genome-wide association studies (GWAS). However, its applicability and coverage of genomic loci depend on the correlation between SNPs, determined by their density and location. We investigated SNP block partitioning in genomic data represented by a limited number of markers using our previously proposed method based on the determinant of linkage disequilibrium (LD) matrix to identify loci associated with ischemic stroke (IS).
MATERIAL AND METHODS
Genotypic data from patients with ischemic stroke (N=923) and the control group (N=305) were analyzed for 67.925 SNPs. The LD matrix determinant was used to group SNPs and assess their association within a block. Blocks with recovered haplotypes were tested for association with ischemic stroke using the χ2 test for independence. Functional analysis of the identified candidate genes for ischemic stroke was performed using the DAVID online service.
RESULTS
SNPs were divided into blocks of correlated SNPs using the determinant of LD matrix. It was found that the maximum number of blocks with the greatest variability in their sizes was observed at similar values of the determinant’s rounding threshold to zero (ε=0.001) in data with low and high SNP density (883.908 SNPs). Eight blocks associated with IS were identified. These blocks contained 102 genomic loci, three of which — IGLC3, IGLC6, and IGLC7 — were overrepresented in the immunoglobulin complex of Gene Ontology.
CONCLUSION
It was found that the determinant of LD matrix as a measure of SNP connectivity allows one to identify SNP blocks in genomic data with both high and low SNP density, which expands the scope of applicability of GWAS based on the use of haplotype tests.