Comparison of missing data approaches in linkage analysis.

Chao Xing; Fredrick R. Schumacher; David V. Conti; John S. Witte

doi:10.1186/1471-2156-4-s1-s44

Comparison of missing data approaches in linkage analysis.

Chao Xing, Fredrick R. Schumacher, David V. Conti, John S. Witte

Research output: Contribution to journal › Article › peer-review

3 Scopus citations

Abstract

Observational cohort studies have been little used in linkage analyses due to their general lack of large, disease-specific pedigrees. Nevertheless, the longitudinal nature of such studies makes them potentially valuable for assessing the linkage between genotypes and temporal trends in phenotypes. The repeated phenotype measures in cohort studies (i.e., across time), however, can have extensive missing information. Existing methods for handling missing data in observational studies may decrease efficiency, introduce biases, and give spurious results. The impact of such methods when undertaking linkage analysis of cohort studies is unclear. Therefore, we compare here six methods of imputing missing repeated phenotypes on results from genome-wide linkage analyses of four quantitative traits from the Framingham Heart Study cohort. We found that simply deleting observations with missing values gave many more nominally statistically significant linkages than the other five approaches. Among the latter, those with similar underlying methodology (i.e., imputation- versus model-based) gave the most consistent results, although some discrepancies remained. Different methods for addressing missing values in linkage analyses of cohort studies can give substantially diverse results, and must be carefully considered to protect against biases and spurious findings.

Original language	English (US)
Journal	BMC genetics
Volume	4 Suppl 1
DOIs	https://doi.org/10.1186/1471-2156-4-s1-s44
State	Published - 2003

ASJC Scopus subject areas

Genetics
Genetics(clinical)

Access to Document

10.1186/1471-2156-4-s1-s44

Cite this

@article{3ec26f54ada746f28926f6f85862a424,

title = "Comparison of missing data approaches in linkage analysis.",

abstract = "Observational cohort studies have been little used in linkage analyses due to their general lack of large, disease-specific pedigrees. Nevertheless, the longitudinal nature of such studies makes them potentially valuable for assessing the linkage between genotypes and temporal trends in phenotypes. The repeated phenotype measures in cohort studies (i.e., across time), however, can have extensive missing information. Existing methods for handling missing data in observational studies may decrease efficiency, introduce biases, and give spurious results. The impact of such methods when undertaking linkage analysis of cohort studies is unclear. Therefore, we compare here six methods of imputing missing repeated phenotypes on results from genome-wide linkage analyses of four quantitative traits from the Framingham Heart Study cohort. We found that simply deleting observations with missing values gave many more nominally statistically significant linkages than the other five approaches. Among the latter, those with similar underlying methodology (i.e., imputation- versus model-based) gave the most consistent results, although some discrepancies remained. Different methods for addressing missing values in linkage analyses of cohort studies can give substantially diverse results, and must be carefully considered to protect against biases and spurious findings.",

author = "Chao Xing and Schumacher, {Fredrick R.} and Conti, {David V.} and Witte, {John S.}",

year = "2003",

doi = "10.1186/1471-2156-4-s1-s44",

language = "English (US)",

volume = "4 Suppl 1",

journal = "BMC genetics",

issn = "1471-2156",

publisher = "BioMed Central",

}

TY - JOUR

T1 - Comparison of missing data approaches in linkage analysis.

AU - Xing, Chao

AU - Schumacher, Fredrick R.

AU - Conti, David V.

AU - Witte, John S.

PY - 2003

Y1 - 2003

N2 - Observational cohort studies have been little used in linkage analyses due to their general lack of large, disease-specific pedigrees. Nevertheless, the longitudinal nature of such studies makes them potentially valuable for assessing the linkage between genotypes and temporal trends in phenotypes. The repeated phenotype measures in cohort studies (i.e., across time), however, can have extensive missing information. Existing methods for handling missing data in observational studies may decrease efficiency, introduce biases, and give spurious results. The impact of such methods when undertaking linkage analysis of cohort studies is unclear. Therefore, we compare here six methods of imputing missing repeated phenotypes on results from genome-wide linkage analyses of four quantitative traits from the Framingham Heart Study cohort. We found that simply deleting observations with missing values gave many more nominally statistically significant linkages than the other five approaches. Among the latter, those with similar underlying methodology (i.e., imputation- versus model-based) gave the most consistent results, although some discrepancies remained. Different methods for addressing missing values in linkage analyses of cohort studies can give substantially diverse results, and must be carefully considered to protect against biases and spurious findings.

AB - Observational cohort studies have been little used in linkage analyses due to their general lack of large, disease-specific pedigrees. Nevertheless, the longitudinal nature of such studies makes them potentially valuable for assessing the linkage between genotypes and temporal trends in phenotypes. The repeated phenotype measures in cohort studies (i.e., across time), however, can have extensive missing information. Existing methods for handling missing data in observational studies may decrease efficiency, introduce biases, and give spurious results. The impact of such methods when undertaking linkage analysis of cohort studies is unclear. Therefore, we compare here six methods of imputing missing repeated phenotypes on results from genome-wide linkage analyses of four quantitative traits from the Framingham Heart Study cohort. We found that simply deleting observations with missing values gave many more nominally statistically significant linkages than the other five approaches. Among the latter, those with similar underlying methodology (i.e., imputation- versus model-based) gave the most consistent results, although some discrepancies remained. Different methods for addressing missing values in linkage analyses of cohort studies can give substantially diverse results, and must be carefully considered to protect against biases and spurious findings.

UR - http://www.scopus.com/inward/record.url?scp=34248652911&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=34248652911&partnerID=8YFLogxK

U2 - 10.1186/1471-2156-4-s1-s44

DO - 10.1186/1471-2156-4-s1-s44

M3 - Article

C2 - 14975112

AN - SCOPUS:34248652911

SN - 1471-2156

VL - 4 Suppl 1

JO - BMC genetics

JF - BMC genetics

ER -

Comparison of missing data approaches in linkage analysis.

Abstract

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this