T. Jaeger, Cognitive Sciences
2008.11.1JOURNAL OF MEMORY AND LANGUAGE
tlooto Summary
This paper identifies several serious problems with the widespread use of ANOVAs for the analysis of categorical outcome variables, and introduces ordinary logit models (i.e. logistic regression), which are well-suited to analyze categorical data and offer many advantages over ANOVA.
Abstract
This paper identifies several serious problems with the widespread use of ANOVAs for the analysis of categorical outcome variables such as forced-choice variables, question-answer accuracy, choice in production (e.g. in syntactic priming research), et cetera. I show that even after applying the arcsine-square-root transformation to proportional data, ANOVA can yield spurious results. I discuss conceptual issues underlying these problems and alternatives provided by modern statistics. Specifically, I introduce ordinary logit models (i.e. logistic regression), which are well-suited to analyze categorical data and offer many advantages over ANOVA. Unfortunately, ordinary logit models do not include random effect modeling. To address this issue, I describe mixed logit models (Generalized Linear Mixed Models for binomially distributed outcomes, Breslow & Clayton, 1993), which combine the advantages of ordinary logit models with the ability to account for random subject and item effects in one step of analysis. Throughout the paper, I use a psycholinguistic data set to compare the different statistical methods.
Citation format
JAEGER, T.; SCIENCES, Cognitive. Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models. JOURNAL OF MEMORY AND LANGUAGE, 2008, 59: 434–446.