Literature DB >> 33400061

Best Practices for Binary and Ordinal Data Analyses.

Brad Verhulst1, Michael C Neale2.   

Abstract

The measurement of many human traits, states, and disorders begins with a set of items on a questionnaire. The response format for these questions is often simply binary (e.g., yes/no) or ordered (e.g., high, medium or low). During data analysis, these items are frequently summed or used to estimate factor scores. In clinical applications, such assessments are often non-normally distributed in the general population because many respondents are unaffected, and therefore asymptomatic. As a result, in many cases these measures violate the statistical assumptions required for subsequent analyses. To reduce the influence of the non-normality and quasi-continuous assessment, variables are frequently recoded into binary (affected-unaffected) or ordinal (mild-moderate-severe) diagnoses. Ordinal data therefore present challenges at multiple levels of analysis. Categorizing continuous variables into ordered categories typically results in a loss of statistical power, which represents an incentive to the data analyst to assume that the data are normally distributed, even when they are not. Despite prior zeitgeists suggesting that, e.g., variables with more than 10 ordered categories may be regarded as continuous and analyzed as if they were, we show via simulation studies that this is not generally the case. In particular, using Pearson product-moment correlations instead of maximum likelihood estimates of polychoric correlations biases the estimated correlations towards zero. This bias is especially severe when a plurality of the observations fall into a single observed category, such as a score of zero. By contrast, estimating the ordinal correlation by maximum likelihood yields no estimation bias, although standard errors are (appropriately) larger. We also illustrate how odds ratios depend critically on the proportion or prevalence of affected individuals in the population, and therefore are sub-optimal for studies where comparisons of association metrics are needed. Finally, we extend these analyses to the classical twin model and demonstrate that treating binary data as continuous will underestimate genetic and common environmental variance components, and overestimate unique environment (residual) variance. These biases increase as prevalence declines. While modeling ordinal data appropriately may be more computationally intensive and time consuming, failing to do so will likely yield biased correlations and biased parameter estimates from modeling them.

Entities:  

Keywords:  Odds ratio; Ordinal data; Pearson product-moment correlation; Point biserial correlation; Polychoric correlation; Prevalence; Tetrachoric correlation

Mesh:

Year:  2021        PMID: 33400061      PMCID: PMC8096648          DOI: 10.1007/s10519-020-10031-x

Source DB:  PubMed          Journal:  Behav Genet        ISSN: 0001-8244            Impact factor:   2.805


  15 in total

1.  Squeezing interval change from ordinal panel data: latent growth curves with ordinal outcomes.

Authors:  Paras D Mehta; Michael C Neale; Brian R Flay
Journal:  Psychol Methods       Date:  2004-09

2.  An empirical evaluation of alternative methods of estimation for confirmatory factor analysis with ordinal data.

Authors:  David B Flora; Patrick J Curran
Journal:  Psychol Methods       Date:  2004-12

3.  Problems and pit-falls in testing for G × E and epistasis in candidate gene studies of human behavior.

Authors:  Lindon Eaves; Brad Verhulst
Journal:  Behav Genet       Date:  2014-09-07       Impact factor: 2.805

Review 4.  An Expanded View of Complex Traits: From Polygenic to Omnigenic.

Authors:  Evan A Boyle; Yang I Li; Jonathan K Pritchard
Journal:  Cell       Date:  2017-06-15       Impact factor: 41.582

5.  OpenMx 2.0: Extended Structural Equation and Statistical Modeling.

Authors:  Michael C Neale; Michael D Hunter; Joshua N Pritikin; Mahsa Zahery; Timothy R Brick; Robert M Kirkpatrick; Ryne Estabrook; Timothy C Bates; Hermine H Maes; Steven M Boker
Journal:  Psychometrika       Date:  2015-01-27       Impact factor: 2.500

6.  A polygenic theory of schizophrenia.

Authors:  I I Gottesman; J Shields
Journal:  Proc Natl Acad Sci U S A       Date:  1967-07       Impact factor: 11.205

7.  Genotype × Environment Interaction in Psychiatric Genetics: Deep Truth or Thin Ice?

Authors:  Lindon Eaves
Journal:  Twin Res Hum Genet       Date:  2017-05-24       Impact factor: 1.587

8.  Asymptotically distribution-free methods for the analysis of covariance structures.

Authors:  M W Browne
Journal:  Br J Math Stat Psychol       Date:  1984-05       Impact factor: 3.380

9.  Latent class analysis of ADHD and comorbid symptoms in a population sample of adolescent female twins.

Authors:  R J Neuman; A Heath; W Reich; K K Bucholz; L Sun; R D Todd; J J Hudziak
Journal:  J Child Psychol Psychiatry       Date:  2001-10       Impact factor: 8.982

10.  Multivariate normal maximum likelihood with both ordinal and continuous variables, and data missing at random.

Authors:  Joshua N Pritikin; Timothy R Brick; Michael C Neale
Journal:  Behav Res Methods       Date:  2018-04
View more
  1 in total

1.  Epigenome-wide meta-analysis of PTSD symptom severity in three military cohorts implicates DNA methylation changes in genes involved in immune system and oxidative stress.

Authors:  Seyma Katrinli; Adam X Maihofer; Agaz H Wani; John R Pfeiffer; Elizabeth Ketema; Andrew Ratanatharathorn; Dewleen G Baker; Marco P Boks; Elbert Geuze; Ronald C Kessler; Victoria B Risbrough; Bart P F Rutten; Murray B Stein; Robert J Ursano; Eric Vermetten; Mark W Logue; Caroline M Nievergelt; Alicia K Smith; Monica Uddin
Journal:  Mol Psychiatry       Date:  2022-01-07       Impact factor: 13.437

  1 in total

北京卡尤迪生物科技股份有限公司 © 2022-2023.