The Russian National Corpus is a representative collection of texts in Russian, counting more than 17 bln tokens and completed with linguistic annotation and search tools
Search in corpora
News
Show allThe collections of spoken language in the Accentological and Spoken corpora have been expanded. The additions include recordings of lectures, television programs, radio and TV interviews, as well as a large collection of literary prose readings. The regional collections have been supplemented with recordings of conversations based on documentary films from the series "Letters from the Province", samples of everyday conversational speech collected by students of Voronezh State University and the Kazakhstan Branch of Lomonosov Moscow State University, materials from folklore expeditions, and interviews collected as part of a project on the history of everyday life in Karelia. The total size of the update is approximately 300,000 word tokens.
The Spoken Corpus now features more than 15 million word tokens, while the total size of the Accentological Corpus, including naive poetry, amounts to 136.6 million word tokens.
We continue to expand the functionality of the Church Slavonic Corpus. As in the other historical corpora of the RNC, users can now search by modern Russian, Late and Early Old East Slavic lemmas. The modern Russian lemma is displayed in the search form by default alongside the Church Slavonic lemma.
Church Slavonic lemmas are linked to the Church Slavonic–Russian Reference Dictionary, which brings together data from several dictionaries and provides grammatical information and word definitions.
In the subcorpus text list, it is now possible to open any text in full. In addition, texts from Church Slavonic printed editions are displayed in the Monomakh font, which more closely reflects traditional Church Slavonic typography.
We continue to expand the corpus functionality for teaching Russian at school. The Practice Example Generator has been updated with rules for spelling consonants in prefixes. These include invariant prefixes such as в-, от-, над-, под, меж-, and others; prefixes ending in з-/с-, such as без-/бес-, из-/ис-, раз-/роз and рас-/рос-, and others; and borrowed prefixes such as экс-, суб-. The new rules cover 16 groups of words with prefixes.
You can access the generator page from the RNC for Schools section by clicking on the corresponding banner.