The most under-used statistical method in corpus linguistics: multi-level (and mixed-effects) models
S. Gries
2015.4.8Corpora
tlooto Summary
This paper is specifically written for corpus linguists to get more information about mixed-effects/multi-level modelling and shows how these models are used practically.
Abstract
Much statistical analysis of psycholinguistic data is now being done with so-called mixed-effects regression models. This development was spearheaded by a few highly influential introductory articles that (i) showed how these regression models are superior to what was the previous gold standard and, perhaps even more importantly, (ii) showed how these models are used practically. Corpus linguistics can benefit from mixed-effects/multi-level models for the same reason that psycholinguistics can – because, for example, speaker-specific and lexically specific idiosyncrasies can be accounted for elegantly; but, in fact, corpus linguistics needs them even more because (i) corpus-linguistic data are observational and, thus, usually unbalanced and messy/noisy, and (ii) most widely used corpora come with a hierarchical structure that corpus linguists routinely fail to consider. Unlike nearly all overviews of mixed-effects/multi-level modelling, this paper is specifically written for corpus linguists to get more o...
Citation format
GRIES, S. The most under-used statistical method in corpus linguistics: Multi-level (and mixed-effects) models. Corpora, 2015, 10: 95–125.