| Literature DB >> 29568820 |
K Bretonnel Cohen1, William A Baumgartner1, Irina Temnikova2.
Abstract
This paper reports SuperCAT, a corpus analysis toolkit. It is a radical extension of SubCAT, the Sublanguage Corpus Analysis Toolkit, from sublanguage analysis to corpus analysis in general. The idea behind SuperCAT is that representative corpora have no tendency towards closure-that is, they tend towards infinity. In contrast, non-representative corpora have a tendency towards closure-roughly, finiteness. SuperCAT focuses on general techniques for the quantitative description of the characteristics of any corpus (or other language sample), particularly concerning the characteristics of lexical distributions. Additionally, SuperCAT features a complete re-engineering of the previous SubCAT architecture.Entities:
Keywords: corpus; representativeness; sublanguage; toolkit
Year: 2016 PMID: 29568820 PMCID: PMC5860820
Source DB: PubMed Journal: LREC Int Conf Lang Resour Eval