Significance tests for R2 of out-of-sample prediction using polygenic scores

MM Momin, S Lee, NR Wray, SH Lee - The American Journal of Human …, 2023 - cell.com
MM Momin, S Lee, NR Wray, SH Lee
The American Journal of Human Genetics, 2023cell.com
The coefficient of determination (R 2) is a well-established measure to indicate the predictive
ability of polygenic scores (PGSs). However, the sampling variance of R 2 is rarely
considered so that 95% confidence intervals (CI) are not usually reported. Moreover, when
comparisons are made between PGSs based on different discovery samples, the sampling
covariance of R 2 is required to test the difference between them. Here, we show how to
estimate the variance and covariance of R 2 values to assess the 95% CI and p value of the …
Summary
The coefficient of determination (R2) is a well-established measure to indicate the predictive ability of polygenic scores (PGSs). However, the sampling variance of R2 is rarely considered so that 95% confidence intervals (CI) are not usually reported. Moreover, when comparisons are made between PGSs based on different discovery samples, the sampling covariance of R2 is required to test the difference between them. Here, we show how to estimate the variance and covariance of R2 values to assess the 95% CI and p value of the R2 difference. We apply this approach to real data calculating PGSs in 28,880 European participants derived from UK Biobank (UKBB) and Biobank Japan (BBJ) GWAS summary statistics for cholesterol and BMI. We quantify the significantly higher predictive ability of UKBB PGSs compared to BBJ PGSs (p value 7.6e−31 for cholesterol and 1.4e−50 for BMI). A joint model of UKBB and BBJ PGSs significantly improves the predictive ability, compared to a model of UKBB PGS only (p value 3.5e−05 for cholesterol and 1.3e−28 for BMI). We also show that the predictive ability of regulatory SNPs is significantly enriched over non-regulatory SNPs for cholesterol (p value 8.9e−26 for UKBB and 3.8e−17 for BBJ). We suggest that the proposed approach (available in R package r2redux) should be used to test the statistical significance of difference between pairs of PGSs, which may help to draw a correct conclusion about the comparative predictive ability of PGSs.
cell.com
以上显示的是最相近的搜索结果。 查看全部搜索结果