Contributor: Ministerstvo školství, mládeže a tělovýchovy České republiky@@LM2018101@@LINDAT/CLARIAH-CZ: Digitální výzkumná infrastruktura pro jazykové technologie, umění a humanitní vědy@@nationalFunds@@ / Original context has metadata only: false / Type: corpus

Start Over Contributor Ministerstvo školství, mládeže a tělovýchovy České republiky@@LM2018101@@LINDAT/CLARIAH-CZ: Digitální výzkumná infrastruktura pro jazykové technologie, umění a humanitní vědy@@nationalFunds@@ Type corpus Original context has metadata only false

31. LongEval Train Collection

Creator:: Galuščáková, Petra, Devaud, Romain, Gonzalez-Saez, Gabriela, Mulhem, Philippe, Goeuriot, Lorraine, Piroi, Florina, and Popel, Martin
Publisher:: Université Grenoble Alpes, Qwant, Research Studios Austria, and Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: information retrieval, parallel corpus, search, and automatic evaluation
Language:: French and English
Description:: The collection consists of queries and documents provided by the Qwant search Engine (https://www.qwant.com). The queries, which were issued by the users of Qwant, are based on the selected trending topics. The documents in the collection were selected with respect to these queries using the Qwant click model. Apart from the documents selected using this model, the collection also contains randomly selected documents from the Qwant index. All the data were collected over June 2022. In total, the collection contains 672 train queries, with corresponding 9656 assessments coming from the Qwant click model, and 98 heldout queries. The set of documents consist of 1,570,734 downloaded, cleaned and filtered Web Pages. Apart from their original French versions, the collection also contains translations of the webpages and queries into English. The collection serves as the official training collection for the 2023 LongEval Information Retrieval Lab (https://clef-longeval.github.io/) organised at CLEF.
Rights:: Qwant LongEval Attribution-NonCommercial-ShareAlike License, PUB, and https://lindat.mff.cuni.cz/repository/xmlui/page/Qwant_LongEval_BY-NC-SA_License

32. ParCzech 3.0

Creator:: Kopp, Matyáš, Stankov, Vladislav, Bojar, Ondřej, Hladká, Barbora, and Straňák, Pavel
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: audio and corpus
Subject:: Parliament of the Czech Republic, Chamber of Deputies, stenographic protocols, TEI encoding, and speech corpus
Language:: Czech
Description:: The ParCzech 3.0 corpus is the third version of ParCzech consisting of stenographic protocols that record the Chamber of Deputies’ meetings held in the 7th term (2013-2017) and the current 8th term (2017-Mar 2021). The protocols are provided in their original HTML format, Parla-CLARIN TEI format, and the format suitable for Automatic Speech Recognition. The corpus is automatically enriched with the morphological, syntactic, and named-entity annotations using the procedures UDPipe 2 and NameTag 2. The audio files are aligned with the texts in the annotated TEI files.
Rights:: Public Domain Dedication (CC Zero), http://creativecommons.org/publicdomain/zero/1.0/, and PUB

33. ParCzech PS7 1.0

Creator:: Hladká, Barbora, Kopp, Matyáš, and Straňák, Pavel
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: Parliament of the Czech Republic, Chamber of Deputies, stenographic protocols, TEI encoding, and TEITOK
Language:: Czech
Description:: The ParCzech PS7 1.0 corpus is the very first member of the corpus family of data coming from the Parliament of the Czech Republic. ParCzech PS7 1.0 consists of stenographic protocols that record the Chamber of Deputies' meetings held in the 7th term between 2013-2017. The audio recordings are available as well. Transcripts are provided in the original HTML as harvested, and also converted into TEI-derived XML format for use in TEITOK corpus manager. The corpus is automatically enriched with the morphological and named-entity annotations using the procedures MorphoDita and NameTag.
Rights:: Public Domain Dedication (CC Zero), http://creativecommons.org/publicdomain/zero/1.0/, and PUB

34. ParCzech PS7 2.0

Creator:: Hladká, Barbora, Kopp, Matyáš, and Straňák, Pavel
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: Parliament of the Czech Republic, Chamber of Deputies, stenographic protocols, TEI encoding, and TEITOK
Language:: Czech
Description:: The ParCzech PS7 2.0 corpus is the second version of ParCzech PS7 consisting of stenographic protocols that record the Chamber of Deputies' meetings held in the 7th term between 2013-2017. The protocols are provided in their original HTML format, TEI format and TEI-derived format to make them searchable in the TEITOK corpus manager. Their audio recordings are available as well. The corpus is automatically enriched with the morphological, syntactic, and named-entity annotations using the procedures UDPipe 2 and NameTag 2.
Rights:: Public Domain Dedication (CC Zero), http://creativecommons.org/publicdomain/zero/1.0/, and PUB

35. Prague Dependency Treebank - Consolidated 1.0 (PDT-C 1.0)

Creator:: Hajič, Jan, Bejček, Eduard, Bémová, Alevtina, Buráňová, Eva, Fučíková, Eva, Hajičová, Eva, Havelka, Jiří, Hlaváčová, Jaroslava, Homola, Petr, Ircing, Pavel, Kárník, Jiří, Kettnerová, Václava, Klyueva, Natalia, Kolářová, Veronika, Kučová, Lucie, Lopatková, Markéta, Mareček, David, Mikulová, Marie, Mírovský, Jiří, Nedoluzhko, Anna, Novák, Michal, Pajas, Petr, Panevová, Jarmila, Peterek, Nino, Poláková, Lucie, Popel, Martin, Popelka, Jan, Romportl, Jan, Rysová, Magdaléna, Semecký, Jiří, Sgall, Petr, Spoustová, Johanka, Straka, Milan, Straňák, Pavel, Synková, Pavlína, Ševčíková, Magda, Šindlerová, Jana, Štěpánek, Jan, Štěpánková, Barbora, Toman, Josef, Urešová, Zdeňka, Vidová Hladká, Barbora, Zeman, Daniel, Zikánová, Šárka, and Žabokrtský, Zdeněk
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: treebank, dependency, tectogrammatics, topic-focus articulation, multiword expressions, coreference, bridging relations, discourse, morphology, syntax, tokenization, lemmatization, semantic relations, lexical semantics, lexicon, valency, speech reconstruction, clauses, speech recognition, and spoken corpus
Language:: Czech
Description:: A richly annotated and genre-diversified language resource, The Prague Dependency Treebank – Consolidated 1.0 (PDT-C 1.0, or PDT-C in short in the sequel) is a consolidated release of the existing PDT-corpora of Czech data, uniformly annotated using the standard PDT scheme. PDT-corpora included in PDT-C: Prague Dependency Treebank (the original PDT contents, written newspaper and journal texts from three genres); Czech part of Prague Czech-English Dependency Treebank (translated financial texts, from English), Prague Dependency Treebank of Spoken Czech (spoken data, including audio and transcripts and multiple speech reconstruction annotation); PDT-Faust (user-generated texts). The difference from the separately published original treebanks can be briefly described as follows: it is published in one package, to allow easier data handling for all the datasets; the data is enhanced with a manual linguistic annotation at the morphological layer and new version of morphological dictionary is enclosed; a common valency lexicon for all four original parts is enclosed. Documentation provides two browsing and editing desktop tools (TrEd and MEd) and the corpus is also available online for searching using PML-TQ.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

36. Quality and Efficiency of Manual Annotation: Data from the Pre-annotation Bias Experiment (part of the PDT-C 2.0 project)

Creator:: Mikulová, Marie, Straka, Milan, Štěpánek, Jan, Štěpánková, Barbora, and Hajič, Jan
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: annotation, syntax, inter-annotator agreement, pre-annotation bias, annotation efficiency, and annotation quality
Language:: Czech
Description:: Input data, individual experimental annotations, and a complete and detailed overview of the measured results related to the experiment described in the referenced paper.
Rights:: Creative Commons - Attribution 4.0 International (CC BY 4.0), http://creativecommons.org/licenses/by/4.0/, and PUB

37. SnakeCLEF 2021

Creator:: Picek, Lukáš, Bolon, Isabelle, Durso, Andrew M., and Castañeda, Rafael Ruiz de
Publisher:: CEUR Workshop Proceedings (CEUR-WS.org)
Type:: IMAGE and corpus
Subject:: LifeCLEF, SnakeCLEF, global health, epidemiology, snake bite, snake, reptile, benchmark, biodiversity, machine learning, computer vision, and Classification
Language:: No linguistic content
Description:: The dataset with 409,679 images belonging to 772 snake species from 188 countries and all continents (386,006 images with labels targeted for development and 23,673 images without labels for testing). In addition, we provide a simple train/val (90% / 10%) split to validate preliminary results while ensuring the same species distributions. Furthermore, we prepared a compact subset (70,208 images) for fast prototyping. The test set data consists of 23,673 images submitted to the iNaturalist platform within the "first four months of 2021. All data were gathered from online biodiversity platforms (i.e., iNaturalist, HerpMapper) and further extended by data scraped from Flickr. The provided dataset has a heavy long-tailed class distribution, where the most frequent species (Thamnophis sirtalis) is represented by 22,163 images and the least frequent by just 10 (Achalinus formosanus).
Rights:: BSD 3-Clause "New" or "Revised" license, http://opensource.org/licenses/BSD-3-Clause, and PUB

38. SQAD 3.2

Creator:: Medveď, Marek
Publisher:: Masaryk University, NLP Centre
Type:: text and corpus
Subject:: QA, Question Answering, SQAD, and Czech QA
Language:: Czech
Description:: Simple question answering database version 3.2 (SQAD v3.2) created from Czech Wikipedia. The new version consists of more than 16000 records. Each record of SQAD consists of multiple files - question, answer extraction, answer selection, URL, question metadata, and in some cases, answer context.
Rights:: Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0), http://creativecommons.org/licenses/by-sa/3.0/, and PUB

39. Synthetic part of CzEng 2.0

Creator:: Popel, Martin
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: parallel corpus
Language:: Czech and English
Description:: CzEng is a sentence-parallel Czech-English corpus compiled at the Institute of Formal and Applied Linguistics (ÚFAL). While the full CzEng 2.0 is freely available for non-commercial research purposes from the project website (https://ufal.mff.cuni.cz/czeng), this release contains only the original monolingual parts of news text (csmono 53M and enmono 79M sentences) with automatic (synthetic) translations by CUBBITT. See the attached README for additional details such as the file format.
Rights:: Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), http://creativecommons.org/licenses/by-sa/4.0/, and PUB

40. Universal Dependencies 2.10

Creator:: Zeman, Daniel, Nivre, Joakim, Abrams, Mitchell, Ackermann, Elia, Aepli, Noëmi, Aghaei, Hamid, Agić, Željko, Ahmadi, Amir, Ahrenberg, Lars, Ajede, Chika Kennedy, Aleksandravičiūtė, Gabrielė, Alfina, Ika, Algom, Avner, Andersen, Erik, Antonsen, Lene, Aplonova, Katya, Aquino, Angelina, Aragon, Carolina, Aranes, Glyd, Aranzabe, Maria Jesus, Arıcan, Bilge Nas, Arnardóttir, Þórunn, Arutie, Gashaw, Arwidarasti, Jessica Naraiswari, Asahara, Masayuki, Aslan, Deniz Baran, Asmazoğlu, Cengiz, Ateyah, Luma, Atmaca, Furkan, Attia, Mohammed, Atutxa, Aitziber, Augustinus, Liesbeth, Badmaeva, Elena, Balasubramani, Keerthana, Ballesteros, Miguel, Banerjee, Esha, Bank, Sebastian, Barbu Mititelu, Verginica, Barkarson, Starkaður, Basile, Rodolfo, Basmov, Victoria, Batchelor, Colin, Bauer, John, Bedir, Seyyit Talha, Bengoetxea, Kepa, Ben Moshe, Yifat, Berk, Gözde, Berzak, Yevgeni, Bhat, Irshad Ahmad, Bhat, Riyaz Ahmad, Biagetti, Erica, Bick, Eckhard, Bielinskienė, Agnė, Bjarnadóttir, Kristín, Blokland, Rogier, Bobicev, Victoria, Boizou, Loïc, Borges Völker, Emanuel, Börstell, Carl, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Boyd, Adriane, Braggaar, Anouck, Brokaitė, Kristina, Burchardt, Aljoscha, Candito, Marie, Caron, Bernard, Caron, Gauthier, Cassidy, Lauren, Cavalcanti, Tatiana, Cebiroğlu Eryiğit, Gülşen, Cecchini, Flavio Massimiliano, Celano, Giuseppe G. A., Čéplö, Slavomír, Cesur, Neslihan, Cetin, Savas, Çetinoğlu, Özlem, Chalub, Fabricio, Chauhan, Shweta, Chi, Ethan, Chika, Taishi, Cho, Yongseok, Choi, Jinho, Chun, Jayeol, Chung, Juyeon, Cignarella, Alessandra T., Cinková, Silvie, Collomb, Aurélie, Çöltekin, Çağrı, Connor, Miriam, Corbetta, Daniela, Courtin, Marine, Cristescu, Mihaela, Daniel, Philemon, Davidson, Elizabeth, Dehouck, Mathieu, de Laurentiis, Martina, de Marneffe, Marie-Catherine, de Paiva, Valeria, Derin, Mehmet Oguz, de Souza, Elvis, Diaz de Ilarraza, Arantza, Dickerson, Carly, Dinakaramani, Arawinda, Di Nuovo, Elisa, Dione, Bamba, Dirix, Peter, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Dwivedi, Puneet, Eckhoff, Hanne, Eiche, Sandra, Eli, Marhaba, Elkahky, Ali, Ephrem, Binyam, Erina, Olga, Erjavec, Tomaž, Etienne, Aline, Evelyn, Wograine, Facundes, Sidney, Farkas, Richárd, Favero, Federica, Ferdaousi, Jannatul, Fernanda, Marília, Fernandez Alcalde, Hector, Foster, Jennifer, Freitas, Cláudia, Fujita, Kazunori, Gajdošová, Katarína, Galbraith, Daniel, Gamba, Federica, Garcia, Marcos, Gärdenfors, Moa, Garza, Sebastian, Gerardi, Fabrício Ferraz, Gerdes, Kim, Ginter, Filip, Godoy, Gustavo, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, González Saavedra, Berta, Griciūtė, Bernadeta, Grioni, Matias, Grobol, Loïc, Grūzītis, Normunds, Guillaume, Bruno, Guillot-Barbance, Céline, Güngör, Tunga, Habash, Nizar, Hafsteinsson, Hinrik, Hajič, Jan, Hajič jr., Jan, Hämäläinen, Mika, Hà Mỹ, Linh, Han, Na-Rae, Hanifmuti, Muhammad Yudistira, Harada, Takahiro, Hardwick, Sam, Harris, Kim, Haug, Dag, Heinecke, Johannes, Hellwig, Oliver, Hennig, Felix, Hladká, Barbora, Hlaváčová, Jaroslava, Hociung, Florinel, Hohle, Petter, Hwang, Jena, Ikeda, Takumi, Ingason, Anton Karl, Ion, Radu, Irimia, Elena, Ishola, Ọlájídé, Ito, Kaoru, Jannat, Siratun, Jelínek, Tomáš, Jha, Apoorva, Johannsen, Anders, Jónsdóttir, Hildur, Jørgensen, Fredrik, Juutinen, Markus, K, Sarveswaran, Kaşıkara, Hüner, Kaasen, Andre, Kabaeva, Nadezhda, Kahane, Sylvain, Kanayama, Hiroshi, Kanerva, Jenna, Kara, Neslihan, Karahóǧa, Ritván, Katz, Boris, Kayadelen, Tolga, Kenney, Jessica, Kettnerová, Václava, Kirchner, Jesse, Klementieva, Elena, Klyachko, Elena, Köhn, Arne, Köksal, Abdullatif, Kopacewicz, Kamil, Korkiakangas, Timo, Köse, Mehmet, Kotsyba, Natalia, Kovalevskaitė, Jolanta, Krek, Simon, Krishnamurthy, Parameswari, Kübler, Sandra, Kuyrukçu, Oğuzhan, Kuzgun, Aslı, Kwak, Sookyoung, Laippala, Veronika, Lam, Lucia, Lambertino, Lorenzo, Lando, Tatiana, Larasati, Septina Dian, Lavrentiev, Alexei, Lee, John, Lê Hồng, Phương, Lenci, Alessandro, Lertpradit, Saran, Leung, Herman, Levina, Maria, Li, Cheuk Ying, Li, Josie, Li, Keying, Li, Yuan, Lim, KyungTae, Lima Padovani, Bruna, Lindén, Krister, Ljubešić, Nikola, Loginova, Olga, Lusito, Stefano, Luthfi, Andry, Luukko, Mikko, Lyashevskaya, Olga, Lynn, Teresa, Macketanz, Vivien, Mahamdi, Menel, Maillard, Jean, Makazhanov, Aibek, Mandl, Michael, Manning, Christopher, Manurung, Ruli, Marşan, Büşra, Mărănduc, Cătălina, Mareček, David, Marheinecke, Katrin, Markantonatou, Stella, Martínez Alonso, Héctor, Martín Rodríguez, Lorena, Martins, André, Mašek, Jan, Matsuda, Hiroshi, Matsumoto, Yuji, Mazzei, Alessandro, McDonald, Ryan, McGuinness, Sarah, Mendonça, Gustavo, Merzhevich, Tatiana, Miekka, Niko, Mischenkova, Karina, Misirpashayeva, Margarita, Missilä, Anna, Mititelu, Cătălin, Mitrofan, Maria, Miyao, Yusuke, Mojiri Foroushani, AmirHossein, Molnár, Judit, Moloodi, Amirsaeid, Montemagni, Simonetta, More, Amir, Moreno Romero, Laura, Moretti, Giovanni, Mori, Keiko Sophie, Mori, Shinsuke, Morioka, Tomohiko, Moro, Shigeki, Mortensen, Bjartur, Moskalevskyi, Bohdan, Muischnek, Kadri, Munro, Robert, Murawaki, Yugo, Müürisep, Kaili, Nainwani, Pinkey, Nakhlé, Mariam, Navarro Horñiacek, Juan Ignacio, Nedoluzhko, Anna, Nešpore-Bērzkalne, Gunta, Nevaci, Manuela, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikaido, Yoshihiro, Nikolaev, Vitaly, Nitisaroj, Rattima, Nourian, Alireza, Nurmi, Hanna, Ojala, Stina, Ojha, Atul Kr., Olúòkun, Adédayọ̀, Omura, Mai, Onwuegbuzia, Emeka, Ordan, Noam, Osenova, Petya, Östling, Robert, Øvrelid, Lilja, Özateş, Şaziye Betül, Özçelik, Merve, Özgür, Arzucan, Öztürk Başaran, Balkız, Paccosi, Teresa, Palmero Aprosio, Alessio, Park, Hyunji Hayley, Partanen, Niko, Pascual, Elena, Passarotti, Marco, Patejuk, Agnieszka, Paulino-Passos, Guilherme, Pedonese, Giulia, Peljak-Łapińska, Angelika, Peng, Siyao, Perez, Cenel-Augusto, Perkova, Natalia, Perrier, Guy, Petrov, Slav, Petrova, Daria, Peverelli, Andrea, Phelan, Jason, Piitulainen, Jussi, Pirinen, Tommi A, Pitler, Emily, Plank, Barbara, Poibeau, Thierry, Ponomareva, Larisa, Popel, Martin, Pretkalniņa, Lauma, Prévost, Sophie, Prokopidis, Prokopis, Przepiórkowski, Adam, Puolakainen, Tiina, Pyysalo, Sampo, Qi, Peng, Rääbis, Andriela, Rademaker, Alexandre, Rahoman, Mizanur, Rama, Taraka, Ramasamy, Loganathan, Ramisch, Carlos, Rashel, Fam, Rasooli, Mohammad Sadegh, Ravishankar, Vinit, Real, Livy, Rebeja, Petru, Reddy, Siva, Regnault, Mathilde, Rehm, Georg, Riabov, Ivan, Rießler, Michael, Rimkutė, Erika, Rinaldi, Larissa, Rituma, Laura, Rizqiyah, Putri, Rocha, Luisa, Rögnvaldsson, Eiríkur, Romanenko, Mykhailo, Rosa, Rudolf, Roșca, Valentin, Rovati, Davide, Rozonoyer, Ben, Rudina, Olga, Rueter, Jack, Rúnarsson, Kristján, Sadde, Shoval, Safari, Pegah, Sagot, Benoît, Sahala, Aleksi, Saleh, Shadi, Salomoni, Alessio, Samardžić, Tanja, Samson, Stephanie, Sanguinetti, Manuela, Sanıyar, Ezgi, Särg, Dage, Saulīte, Baiba, Sawanakunanon, Yanin, Saxena, Shefali, Scannell, Kevin, Scarlata, Salvatore, Schneider, Nathan, Schuster, Sebastian, Schwartz, Lane, Seddah, Djamé, Seeker, Wolfgang, Seraji, Mojgan, Shahzadi, Syeda, Shen, Mo, Shimada, Atsuko, Shirasu, Hiroyuki, Shishkina, Yana, Shohibussirri, Muh, Sichinava, Dmitry, Siewert, Janine, Sigurðsson, Einar Freyr, Silveira, Aline, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Skachedubova, Maria, Smith, Aaron, Soares-Bastos, Isabela, Sourov, Shafi, Spadine, Carolyn, Sprugnoli, Rachele, Stamou, Vivian, Steingrímsson, Steinþór, Stella, Antonio, Straka, Milan, Strickland, Emmett, Strnadová, Jana, Suhr, Alane, Sulestio, Yogi Lesmana, Sulubacak, Umut, Suzuki, Shingo, Swanson, Daniel, Szántó, Zsolt, Taguchi, Chihiro, Taji, Dima, Takahashi, Yuta, Tamburini, Fabio, Tan, Mary Ann C., Tanaka, Takaaki, Tanaya, Dipta, Tavoni, Mirko, Tella, Samson, Tellier, Isabelle, Testori, Marinella, Thomas, Guillaume, Tonelli, Sara, Torga, Liisi, Toska, Marsida, Trosterud, Trond, Trukhina, Anna, Tsarfaty, Reut, Türk, Utku, Tyers, Francis, Uematsu, Sumire, Untilov, Roman, Urešová, Zdeňka, Uria, Larraitz, Uszkoreit, Hans, Utka, Andrius, Vagnoni, Elena, Vajjala, Sowmya, van der Goot, Rob, Vanhove, Martine, van Niekerk, Daniel, van Noord, Gertjan, Varga, Viktor, Vedenina, Uliana, Villemonte de la Clergerie, Eric, Vincze, Veronika, Vlasova, Natalia, Wakasa, Aya, Wallenberg, Joel C., Wallin, Lars, Walsh, Abigail, Wang, Jing Xian, Washington, Jonathan North, Wendt, Maximilan, Widmer, Paul, Wigderson, Shira, Wijono, Sri Hartati, Williams, Seyi, Wirén, Mats, Wittern, Christian, Woldemariam, Tsegay, Wong, Tak-sum, Wróblewska, Alina, Yako, Mary, Yamashita, Kayo, Yamazaki, Naoki, Yan, Chunxiao, Yasuoka, Koichi, Yavrumyan, Marat M., Yenice, Arife Betül, Yıldız, Olcay Taner, Yu, Zhuoran, Yuliawati, Arlisa, Žabokrtský, Zdeněk, Zahra, Shorouq, Zeldes, Amir, Zhou, He, Zhu, Hanzhi, Zhuravleva, Anna, and Ziane, Rayan
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, Urdu, Russia Buriat, Northern Kurdish, Northern Sami, Upper Sorbian, Afrikaans, Yue Chinese, Marathi, Serbian, Swedish Sign Language, Telugu, Amharic, Armenian, Breton, Faroese, Komi-Zyrian, Nigerian Pidgin, Old French (842-ca. 1400), Tagalog, Thai, Warlpiri, Yoruba, Akkadian, Bambara, Erzya, Maltese, Welsh, Wolof, Assyrian Neo-Aramaic, Literary Chinese, Old Russian, Karelian, Mbyá Guaraní, Bhojpuri, Komi-Permyak, Livvi, Moksha, Scottish Gaelic, Skolt Sami, Swiss German, Albanian, Icelandic, Akuntsu, Apurinã, Chukot, Khunsari, Manx, Mundurukú, Nayini, Old Turkish, Soi, South Levantine Arabic, Tupinambá, Beja, Western Frisian, Guajajára, Urubú-Kaapor, Kangri, K'iche', Low German, Makuráp, Central Siberian Yupik, Western Armenian, Bengali, Javanese, Karo (Brazil), Ligurian, Neapolitan, Tatar, Xibe, Yakut, Ancient Hebrew, Cebuano, Guarani, Hittite, Madi, Emerillon, and Umbrian
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:: Licence Universal Dependencies v2.10, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.10, and PUB

31. LongEval Train Collection

32. ParCzech 3.0

33. ParCzech PS7 1.0

34. ParCzech PS7 2.0

35. Prague Dependency Treebank - Consolidated 1.0 (PDT-C 1.0)

36. Quality and Efficiency of Manual Annotation: Data from the Pre-annotation Bias Experiment (part of the PDT-C 2.0 project)

37. SnakeCLEF 2021

38. SQAD 3.2

39. Synthetic part of CzEng 2.0

40. Universal Dependencies 2.10

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Creator

Show values starting with

Language

Show values starting with

Publisher

Show values starting with

Rights

Show values starting with

Subject

Show values starting with

Type

Date

Original context has metadata only

Harvested from