1 - 7 of 7
Number of results to display per page
Search Results
2. Co je v ČNK nového VII: (Zprávy z Českého národního korpusu)
- Creator:
- Křen, Michal
- Format:
- bez média and svazek
- Type:
- model:article and TEXT
- Language:
- Czech
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
3. ORAL2013: balanced corpus of informal spoken Czech (transcriptions & audio)
- Creator:
- Benešová, Lucie, Křen, Michal, and Waclawičová, Martina
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- audio and corpus
- Subject:
- balanced corpus, spoken language, and speech corpus
- Language:
- Czech
- Description:
- ORAL2013 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus comprises 835 recordings from 2008–2011 that contain 2 785 189 words (i.e. 3 285 508 tokens including punctuation) uttered by 2 544 speakers, out of which 1 297 speakers are unique. ORAL2013 is balanced in the main sociolinguistic categories of the speakers (gender, age group, education, region of childhood residence). The (anonymized) transcriptions are provided in the Transcriber XML format, audio (with corresponding anonymization beeps) is in uncompressed 16-bit PCM WAV, mono, 16 kHz format. Another format option of the transcriptions is also available under less restrictive CC BY-NC-SA license at http://hdl.handle.net/11234/1-1847
- Rights:
- License Agreement for Czech National Corpus Data, https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc-data, and ACA
4. ORAL2013: balanced corpus of informal spoken Czech (transcriptions)
- Creator:
- Benešová, Lucie, Křen, Michal, and Waclawičová, Martina
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- text and corpus
- Subject:
- balanced corpus and spoken language
- Language:
- Czech
- Description:
- ORAL2013 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus comprises 835 recordings from 2008–2011 that contain 2 785 189 words (i.e. 3 285 508 tokens including punctuation) uttered by 2 544 speakers, out of which 1 297 speakers are unique. ORAL2013 is balanced in the main sociolinguistic categories of speakers (gender, age group, education, region of childhood residence). The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query engine to registered users of the CNC at http://www.korpus.cz Please note: this item includes only the transcriptions, audio is available under more restrictive non-CC license at http://hdl.handle.net/11234/1-1848
- Rights:
- Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB
5. SYN v4: large corpus of written Czech
- Creator:
- Křen, Michal, Cvrček, Václav, Čapka, Tomáš, Čermáková, Anna, Hnátková, Milena, Chlumská, Lucie, Jelínek, Tomáš, Kováříková, Dominika, Petkevič, Vladimír, Procházka, Pavel, Skoumalová, Hana, Škrabal, Michal, Truneček, Petr, Vondřička, Pavel, and Zasina, Adrian
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- text and corpus
- Subject:
- corpus and written language
- Language:
- Czech
- Description:
- Corpus of contemporary written (printed) Czech sized 3.6 GW (i.e. 4.3 billion tokens). It covers mostly the period of 1990–2014 and it is a traditional corpus (as opposed to the web-crawled corpora) with rich metadata containing bibliographical information etc. Although it contains a wide range of text types (fiction, non-fiction, newspapers), the newspapers prevail noticeably. The corpus is lemmatized and morphologically annotated by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query interface to registered users of the CNC at http://www.korpus.cz with one important exception: the corpus are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) with ordering randomized within the given document.
- Rights:
- Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA
6. SYN v9: large corpus of written Czech
- Creator:
- Křen, Michal, Cvrček, Václav, Henyš, Jan, Hnátková, Milena, Jelínek, Tomáš, Kocek, Jan, Kováříková, Dominika, Křivan, Jan, Milička, Jiří, Petkevič, Vladimír, Procházka, Pavel, Skoumalová, Hana, Šindlerová, Jana, and Škrabal, Michal
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- text and corpus
- Subject:
- corpus and written language
- Language:
- Czech
- Description:
- Corpus of contemporary written (printed) Czech sized 4.7 GW (i.e. 5.7 billion tokens). It covers mostly the 1990-2019 period and features rich metadata including detailed bibliographical information, text-type classification etc. SYN v9 contains a wide variety of text types (fiction, non-fiction, newspapers), but the newspapers prevail noticeably. The corpus is lemmatized and morphologically tagged by the new CNC tagset first utilized for the annotation of the SYN2020 corpus. SYN v9 is provided in a CoNLL-U-like vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query interface to the registered users of CNC at http://www.korpus.cz with one important exception: the corpus is shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) with ordering randomized within the given document.
- Rights:
- Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA
7. SYN2015: representative corpus of written Czech
- Creator:
- Křen, Michal, Cvrček, Václav, Čapka, Tomáš, Čermáková, Anna, Hnátková, Milena, Chlumská, Lucie, Kováříková, Dominika, Jelínek, Tomáš, Petkevič, Vladimír, Procházka, Pavel, Skoumalová, Hana, Škrabal, Michal, Truneček, Petr, Vondřička, Pavel, and Zasina, Adrian
- Publisher:
- Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
- Type:
- text and corpus
- Subject:
- representative corpus and written language
- Language:
- Czech
- Description:
- Representative corpus of contemporary written Czech sized 100 MW. It was created as a representation of printed language from 2010–2014 containing a wide range of text types (fiction, professional literature, newspapers etc.). The corpus is lemmatized, morphologically and syntactically annotated by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) with ordering randomized within the given document.
- Rights:
- Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA