corpora
COCA (Corpus of Contemporary American English)
COHA (Corpus of Historical American English)
Google Books (huge n-grams dataset; has AE and BE)
See also Culturomics
Strathy Corpus of Canadian English
A collection of English corpora
Corpus Concordance English (Lextutor)
More from wikipedia:
- Google N-Grams Corpus – Largest English corpus at 155 billion words.[1] Also has corpora for other languages. To download datasets of this corpus, see [2]
- American National Corpus
- Bank of English
- British National Corpus
- Corpus Juris Secundum
- Corpus of Contemporary American English (COCA) 425 million words, 1990–2011. Freely searchable online.
- Brown Corpus, forming part of the “Brown Family” of corpora, together with LOB, Frown and F-LOB.
- International Corpus of English
- Oxford English Corpus
- Scottish Corpus of Texts & Speech
- Corpus Resource Database (CoRD), more than 80 English language corpora.[3]
HC Corpora — various langauges, download (not searchable online)
USENET ARCHIVES:
The best way to search Usenet is with ordinary Google search, by adding the term site:groups.google.org and then clicking on “Search tools” to set a date range.
Corpora on Github — JSONs for programmers






