Subject: corpus - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Subject corpus Date Unknown

41. Prague Dependency Treebank 2.0 (PDT 2.0)

Creator:: Hajič, Jan, Panevová, Jarmila, Hajičová, Eva, Sgall, Petr, Pajas, Petr, Štěpánek, Jan, Havelka, Jiří, Mikulová, Marie, Žabokrtský, Zdeněk, Ševčíková-Razímová, Magda, and Urešová, Zdeňka
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: corpus, Czech, treebank, and PDT
Language:: Czech
Description:: The Prague Dependency Treebank 2.0 (PDT 2.0) contains a large amount of Czech texts with complex and interlinked morphological (two million words), syntactic (1.5 MW) and complex semantic annotation (0.8 MW); in addition, certain properties of sentence information structure and coreference relations are annotated at the semantic level. PDT 2.0 is based on the long-standing Praguian linguistic tradition, adapted for the current Computational Linguistics research needs. The corpus itself uses the latest annotation technology. Software tools for corpus search, annotation and language analysis are included. Extensive documentation (in English) is provided as well. and 1ET101120413 (Data a nástroje pro informační systémy) MSM 0021620838 (Moderní metody, struktury a systémy informatiky) 1ET101120503 (Integrace jazykových zdrojů za účelem extrakce informací z přirozených textů) 1P05ME752 (Vícejazyčný valenční a predikátový slovník přirozeného jazyka) LC536 (Centrum komputační lingvistiky)
Rights:: PDT 2.0 License, https://lindat.mff.cuni.cz/repository/xmlui/page/license-pdt2, and ACA

42. Prague Dependency Treebank of Spoken Language (PDTSL) 0.5

Creator:: Hajič, Jan, Pajas, Petr, Mareček, David, Mikulová, Marie, Urešová, Zdeňka, and Podveský, Petr
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: audio and corpus
Subject:: corpus and spoken language
Language:: Czech and English
Description:: The first edition of a speech corpus with a speech reconstruction layer (edited transcript). The project of speech reconstruction of Czech and English has been started at UFAL together with the PIRE project in 2005, and has gradually grown from ideas to (first) annotation specification, annotation software and actual annotation. It is part of the Prague Dependency Treebank family of annotated corpus resources and tools, to which it adds the spoken language layer(s). and LC536; MSM0021620838; IST-034344; ME838
Rights:: PDTSL, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-pdtsl, and ACA

43. Preamble 1.0

Creator:: Hladká, Barbora and Mírovský, Jiří
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: corpus, multilingual, and subjects
Language:: Czech, English, French, and Polish
Description:: Preamble 1.0 is a multilingual annotated corpus of the preamble of the EU REGULATION 2020/2092 OF THE EUROPEAN PARLIAMENT AND OF THE COUNCIL. The corpus consists of four language versions of the preamble (Czech, English, French, Polish), each of them annotated with sentence subjects. The data were annotated in the Brat tool (https://brat.nlplab.org/) and are distributed in the Brat native format, i.e. each annotated preamble is represented by the original plain text and a stand-off annotation file.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

44. Proměny prózy v letech 1992 až 2018

Creator:: Poukarová, Petra and Cvrček, Václav
Format:: bez média and svazek
Type:: model:article and TEXT
Subject:: beletrie, próza, registr, multidimenzionální analýza, korpus, překlad, fiction, prose, register, multidimension analysis, corpus, and translation
Language:: Czech
Description:: This study summarizes a corpus-based analysis of tendencies in register variation of Czech-written fiction texts in the period from 1992 to 2018. The analysis is based on projection of the results from a large sample of Czech prose texts (1070 texts, 12.7 mil. words) on a general register model (established by previous research using multidimensional analysis). The major tendencies found in the material are a decrease of cohesion level, addressee coding and retrospective narration, and increased polythematicity/lexical richness. These findings are supplemented by additional analyses of the role of translation, the position of a text excerpt in the original text (beginning, middle and end) and type of text in the results
Rights:: http://creativecommons.org/licenses/by-nc-sa/4.0/ and policy:public

45. Regionenkorpus (C4-Korpus)

Publisher:: Berlin-Brandenburg Academy of Sciences and Humanities
Format:: application/tei+xml
Type:: corpus
Subject:: corpus
Language:: German
Description:: The C4 corpus is a joined effort of the project Digitales Wörterbuch der deutschen Sprache (DWDS), the Austrian Academy Corpus (AAC), the Korpus Südtirol and the Schweizer Textkorpus (CHTK). The Corpus is composed of corpora of all four partner institutions.
Rights:: Not specified

46. Some current problems of corpus and computational linguistics, or Fifteen commandments and general truths

Creator:: Čermák, František
Format:: bez média and svazek
Type:: model:article and TEXT
Subject:: corpus, corpus lingustics, computational linguistics, methodology, type of data, type of information, representativeness of corpora, systems of tagging, lemmatizers, ir/regularity in language, collocations, meaning, aligners, korpus, korpusová lingvistika, komputační lingvistika, metodologie, typy dat, typy informace, reprezentativnost korpusu, systémy taggování, lemmatizátory, ne/pravidelnost v jazyce, kolokace, význam, and alignery
Language:: Czech
Description:: This contribution, which in a brief, succint and almost aphoristic way, critically brings forward to the reader a number of problems of today’s corpus and computational linguistics as well as their unsatisfactory solutions, is trying, at the same time, to do away with a number of myths and simplified opinions in the field. and Příspěvek ve stručné a téměř aforizované podobě připomíná řadu kritizovaných problémů a jejich neuspokojivých řešení v dnešní korpusové a komputační lingvistice a snaží se tak odstranit řadu mýtů a zjednodušujících představ.
Rights:: http://creativecommons.org/publicdomain/mark/1.0/ and policy:public

47. Srovnání žánrů v korpusu na základě syntaktických funkcí substantiv

Creator:: Jelínek, Tomáš
Format:: bez média and svazek
Type:: model:article and TEXT
Subject:: syntax, syntaktická funkce, korpus, žánr, reprezentativnost, syntactic function, corpus, genre, and representativeness
Language:: Czech
Description:: Large synchronic textual corpora of the Czech National Corpus are built as representative: they contain a balanced quantity of texts of various styles, divided into three genre subcorpora: fiction, technical/scientific literature and journalism. Comparisons of these genres have been performed on phonological and morphological level; in this paper, I deal with differences between genres on the surface-syntactic level. I use an automatic syntactic annotation of the SYN2005 corpus in the formalism of the analytical layer of the Prague Dependency Treebank. I compare the frequencies of syntactic functions of nouns in the three genres represented by the corresponding subcorpora of SYN2005. I also present a more detailed analysis of four syntactic phenomena: subtypes of the function of attribute in non-prepositional genitive; frequencies of groups of the type pan Novák (Mr. Novák); frequencies of the function of agent in passive constructions expressed by nouns in non-prepositional instrumental and the ratio of the expression of the nominal part of a verbal-nominal predicate by nominative and instrumental. Significant differences found between genres in all the syntactic phenomena analyzed show that in comparing corpora one should carefully monitor their genre composition.
Rights:: http://creativecommons.org/publicdomain/mark/1.0/ and policy:public

48. SYN v4: large corpus of written Czech

Creator:: Křen, Michal, Cvrček, Václav, Čapka, Tomáš, Čermáková, Anna, Hnátková, Milena, Chlumská, Lucie, Jelínek, Tomáš, Kováříková, Dominika, Petkevič, Vladimír, Procházka, Pavel, Skoumalová, Hana, Škrabal, Michal, Truneček, Petr, Vondřička, Pavel, and Zasina, Adrian
Publisher:: Charles University, Faculty of Arts, Institute of the Czech National Corpus
Type:: text and corpus
Subject:: corpus and written language
Language:: Czech
Description:: Corpus of contemporary written (printed) Czech sized 3.6 GW (i.e. 4.3 billion tokens). It covers mostly the period of 1990–2014 and it is a traditional corpus (as opposed to the web-crawled corpora) with rich metadata containing bibliographical information etc. Although it contains a wide range of text types (fiction, non-fiction, newspapers), the newspapers prevail noticeably. The corpus is lemmatized and morphologically annotated by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query interface to registered users of the CNC at http://www.korpus.cz with one important exception: the corpus are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) with ordering randomized within the given document.
Rights:: Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA

49. SYN v9: large corpus of written Czech

Creator:: Křen, Michal, Cvrček, Václav, Henyš, Jan, Hnátková, Milena, Jelínek, Tomáš, Kocek, Jan, Kováříková, Dominika, Křivan, Jan, Milička, Jiří, Petkevič, Vladimír, Procházka, Pavel, Skoumalová, Hana, Šindlerová, Jana, and Škrabal, Michal
Publisher:: Charles University, Faculty of Arts, Institute of the Czech National Corpus
Type:: text and corpus
Subject:: corpus and written language
Language:: Czech
Description:: Corpus of contemporary written (printed) Czech sized 4.7 GW (i.e. 5.7 billion tokens). It covers mostly the 1990-2019 period and features rich metadata including detailed bibliographical information, text-type classification etc. SYN v9 contains a wide variety of text types (fiction, non-fiction, newspapers), but the newspapers prevail noticeably. The corpus is lemmatized and morphologically tagged by the new CNC tagset first utilized for the annotation of the SYN2020 corpus. SYN v9 is provided in a CoNLL-U-like vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query interface to the registered users of CNC at http://www.korpus.cz with one important exception: the corpus is shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) with ordering randomized within the given document.
Rights:: Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA

50. Syntaktická adverbializace typu jaktěživo neměl názor (zrušení shody se subjektem v rodě a čísle)

Creator:: Štěpán, Josef
Format:: bez média and svazek
Type:: model:article and TEXT
Subject:: agreement and disagreement with the subject in gender and number, adverbialisation, frequency, noun, pronoun, adverbial (frozen) expression, negative clause, corpus, shoda a neshoda se subjektem v rodě a čísle, adverbializace, frekvence, substantivum, zájmeno, adverbiální (ustrnulý) výraz, záporná věta, and korpus
Language:: Czech
Description:: On the basis of the material of the corpus SYN, the article deals, at first, with the description of morphologically frozen expressions jakživ, jaktěživ with an adverbial meaning ''never'' in negative clauses, while these expressions are, due to their ending, in syntactic agreement in gender and number with the grammatical subject. Also this agreement in positive clauses, where the frozen expressions mean ''ever (in one’s life)'', is briefly mentioned. However, the principal aim of the article is to show that the syntactic adverbialisation of these expressions in negative clauses causes the disturbance of this agreement, cf. jaktěživo neměl názor ''never in his life had he an opinion'', while there are two possible results of this adverbialisation: the forms of neuter jaktěživo, jakživo are more common in Bohemia, while the forms of masculine jaktěživ, jakživ are used rather in Moravia. The author interprets the frequency of both concordant and non-concordant (frozen) expressions, ordered according to their descending frequency in SYN.
Rights:: http://creativecommons.org/publicdomain/mark/1.0/ and policy:public

41. Prague Dependency Treebank 2.0 (PDT 2.0)

42. Prague Dependency Treebank of Spoken Language (PDTSL) 0.5

43. Preamble 1.0

44. Proměny prózy v letech 1992 až 2018

45. Regionenkorpus (C4-Korpus)

46. Some current problems of corpus and computational linguistics, or Fifteen commandments and general truths

47. Srovnání žánrů v korpusu na základě syntaktických funkcí substantiv

48. SYN v4: large corpus of written Czech

49. SYN v9: large corpus of written Czech

50. Syntaktická adverbializace typu jaktěživo neměl názor (zrušení shody se subjektem v rodě a čísle)

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Coverage

Show values starting with

Creator

Show values starting with

Format

Language

Show values starting with

Publisher

Show values starting with

Rights

Show values starting with

Subject

Show values starting with

Type

Show values starting with

Original context has metadata only

Harvested from