Language: Chinese / Rights: PUB - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Language Chinese Rights PUB

11. DaMuEL 1.0: A Large Multilingual Dataset for Entity Linking

Creator:: Kubeša, David and Straka, Milan
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: entity linking, NEL, NER, dataset, and knowledge base
Language:: Afrikaans, Arabic, Armenian, Basque, Belarusian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Korean, Latin, Latvian, Lithuanian, Maltese, Marathi, Modern Greek (1453-), Northern Sami, Norwegian Nynorsk, Persian, Polish, Portuguese, Romanian, Russian, Scottish Gaelic, Serbian, Slovak, Slovenian, Spanish, Swedish, Tamil, Telugu, Uighur, Ukrainian, Urdu, Vietnamese, and Wolof
Description:: We present DaMuEL, a large Multilingual Dataset for Entity Linking containing data in 53 languages. DaMuEL consists of two components: a knowledge base that contains language-agnostic information about entities, including their claims from Wikidata and named entity types (PER, ORG, LOC, EVENT, BRAND, WORK_OF_ART, MANUFACTURED); and Wikipedia texts with entity mentions linked to the knowledge base, along with language-specific text from Wikidata such as labels, aliases, and descriptions, stored separately for each language. The Wikidata QID is used as a persistent, language-agnostic identifier, enabling the combination of the knowledge base with language-specific texts and information for each entity. Wikipedia documents deliberately annotate only a single mention for every entity present; we further automatically detect all mentions of named entities linked from each document. The dataset contains 27.9M named entities in the knowledge base and 12.3G tokens from Wikipedia texts. The dataset is published under the CC BY-SA licence.
Rights:: Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), http://creativecommons.org/licenses/by-sa/4.0/, and PUB

12. Morpho-syntactically annotated corpora provided for the PARSEME Shared Task on Semi-Supervised Identification of Verbal Multiword Expressions (edition 1.2)

Creator:: Guillaume, Bruno, Ramisch, Carlos, Waszczuk, Jakub, Monti, Johanna, Di Buono, Maria Pia, Sangati, Federico, Speranza, Giulia, Carlino, Carola, Güngör, Tunga, Yirmibeşoğlu, Zeynep, Sak, Haşim, Saraçlar, Murat, Giouli, Voula, Foufi, Vassiliki, Ramisch, Renata, Rademaker, Alexandre, Vale, Oto, Wilkens, Rodrigo, Candito, Marie, Crabbé, Benoît, Segonne, Vincent, Liebeskind, Chaya, Stymne, Sara, Hajič, Jan, Ginter, Filip, Luotolahti, Juhani, Straka, Milan, Zeman, Daniel, Barbu Mititelu, Verginica, Cristescu, Mihaela, Vaidya, Ashwini, Bhatia, Archna, Lichte, Timm, Ehren, Rafael, Jiang, Menghan, Xu, Hongzhi, Walsh, Abigail, Irimia, Elena, and Dowling, Meghan
Publisher:: PARSEME
Type:: text and corpus
Subject:: morphosyntactic annotation, dependency trees, and morphological analysis
Language:: German, Modern Greek (1453-), Basque, French, Irish, Hebrew, Hindi, Italian, Polish, Portuguese, Romanian, Swedish, Turkish, and Chinese
Description:: This multilingual resource contains corpora for 14 languages, gathered at the occasion of the 1.2 edition of the PARSEME Shared Task on semi-supervised Identification of Verbal MWEs (2020). These corpora were meant to serve as additional "raw" corpora, to help discovering unseen verbal MWEs. The corpora are provided in CONLL-U (https://universaldependencies.org/format.html) format. They contain morphosyntactic annotations (parts of speech, lemmas, morphological features, and syntactic dependencies). Depending on the language, the information comes from treebanks (mostly Universal Dependencies v2.x) or from automatic parsers trained on UD v2.x treebanks (e.g., UDPipe). VMWEs include idioms (let the cat out of the bag), light-verb constructions (make a decision), verb-particle constructions (give up), inherently reflexive verbs (help oneself), and multi-verb constructions (make do). For the 1.2 shared task edition, the data covers 14 languages, for which VMWEs were annotated according to the universal guidelines. The corpora are provided in the cupt format, inspired by the CONLL-U format. Morphological and syntactic information – not necessarily using UD tagsets – including parts of speech, lemmas, morphological features and/or syntactic dependencies are also provided. Depending on the language, the information comes from treebanks (e.g., Universal Dependencies) or from automatic parsers trained on treebanks (e.g., UDPipe). This item contains training, development and test data, as well as the evaluation tools used in the PARSEME Shared Task 1.2 (2020). The annotation guidelines are available online: http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.2
Rights:: PARSEME Shared Task Raw Corpus Data (v. 1.2) Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-mwe-1.2-raw, and PUB

13. Multilingual static embeddings for Verbal Multiword Expressions trained on PARSEME raw corpora

Creator:: Estève, Louis Clément, Savary, Agata, and Lavergne, Thomas
Publisher:: Université Paris-Saclay, CNRS, Laboratoire Interdisciplinaire des Sciences du Numérique
Type:: text, computationalLexicon, and lexicalConceptualResource
Subject:: verbal multiword expressions, word embeddings, and word2vec
Language:: German, Modern Greek (1453-), Basque, French, Irish, Hebrew, Hindi, Italian, Polish, Portuguese, Romanian, Swedish, Turkish, and Chinese
Description:: This resource is a set of 14 vector spaces for single words and Verbal Multiword Expressions (VMWEs) in different languages (German, Greek, Basque, French, Irish, Hebrew, Hindi, Italian, Polish, Brazilian Portuguese, Romanian, Swedish, Turkish, Chinese). They were trained with the Word2Vec algorithm, in its skip-gram version, on PARSEME raw corpora automatically annotated for morpho-syntax (http://hdl.handle.net/11234/1-3367). These corpora were annotated by Seen2Seen, a rule-based VMWE identifier, one of the leading tools of the PARSEME shared task version 1.2. VMWE tokens were merged into single tokens. The format of the vector space files is that of the original Word2Vec implementation by Mikolov et al. (2013), i.e. a binary format. For compression, bzip2 was used.
Rights:: PARSEME Shared Task Raw Corpus Data (v. 1.2) Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-mwe-1.2-raw, and PUB

17. Uniform Meaning Representation

Creator:: Bonn, Julia, Ching-wen, Chen, Cowell, James Andrew, Croft, William, Denk, Lukas, Hajič, Jan, Lai, Kenneth, Palmer, Martha, Palmer, Alexis, Pustejovsky, James, Sun, Haibo, Vallejos Yopán, Rosa, Van Gysel, Jens, Vigus, Meagan, Xue, Nianwen, and Zhao, Jin
Publisher:: UMR Consortium
Type:: text and corpus
Subject:: uniform meaning representation and semantics
Language:: Chinese, Arapaho, English, Cocama-Cocamilla, Navajo, and Sanapaná
Description:: The goal of the Uniform Meaning Representation (UMR) project is to design a meaning representation that can be used to annotate the semantic content of a text. UMR is primarily based on Abstract Meaning Representation (AMR), an annotation framework initially designed for English, but also draws from other meaning representations. UMR extends AMR to other languages, particularly morphologically complex, low-resource languages. UMR also adds features to AMR that are critical to semantic interpretation and enhances AMR by proposing a companion document-level representation that captures linguistic phenomena such as coreference as well as temporal and modal dependencies that potentially go beyond sentence boundaries. UMR is intended to be scalable, learnable, and cross-linguistically plausible. It is designed to support both lexical and logical inference.
Rights:: Uniform Meaning Representation v1.0 License Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/license-umr-1.0, and PUB

18. Universal Dependencies 1.3

Creator:: Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Berzak, Yevgeni, Bhat, Riyaz Ahmad, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Çöltekin, Çağrı, Connor, Miriam, de Marneffe, Marie-Catherine, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Galbraith, Daniel, Garza, Sebastian, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gokirmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grūzītis, Normunds, Guillaume, Bruno, Hajič, Jan, Haug, Dag, Hladká, Barbora, Ion, Radu, Irimia, Elena, Johannsen, Anders, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Katz, Boris, Kenney, Jessica, Krek, Simon, Laippala, Veronika, Lam, Lucia, Lenci, Alessandro, Ljubešić, Nikola, Lyashevskaya, Olga, Lynn, Teresa, Makazhanov, Aibek, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, Mori, Keiko Sophie, Mori, Shunsuke, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nikolaev, Vitaly, Nurmi, Hanna, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Ramasamy, Loganathan, Rituma, Laura, Rosa, Rudolf, Saleh, Shadi, Saulīte, Baiba, Schuster, Sebastian, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Simov, Kiril, Smith, Aaron, Spadine, Carolyn, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Tanaka, Takaaki, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Uria, Larraitz, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Wang, Jing Xian, Washington, Jonathan North, Žabokrtský, Zdeněk, Zeman, Daniel, and Zhu, Hanzhi
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, and Turkish
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:: Licence Universal Dependencies v1.3, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.3, and PUB

19. Universal Dependencies 1.4

Creator:: Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Berzak, Yevgeni, Bhat, Riyaz Ahmad, Bick, Eckhard, Börstell, Carl, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Chalub, Fabricio, Çöltekin, Çağrı, Connor, Miriam, Davidson, Elizabeth, de Marneffe, Marie-Catherine, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Dwivedi, Puneet, Eli, Marhaba, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Freitas, Claudia, Gajdošová, Katarína, Galbraith, Daniel, Garcia, Marcos, Gärdenfors, Moa, Garza, Sebastian, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grioni, Matias, Grūzītis, Normunds, Guillaume, Bruno, Hajič, Jan, Hà Mỹ, Linh, Haug, Dag, Hladká, Barbora, Ion, Radu, Irimia, Elena, Johannsen, Anders, Jørgensen, Fredrik, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Katz, Boris, Kenney, Jessica, Kotsyba, Natalia, Krek, Simon, Laippala, Veronika, Lam, Lucia, Lê Hồng, Phương, Lenci, Alessandro, Ljubešić, Nikola, Lyashevskaya, Olga, Lynn, Teresa, Makazhanov, Aibek, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Martins, André, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, Mori, Keiko Sophie, Mori, Shunsuke, Moskalevskyi, Bohdan, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikolaev, Vitaly, Nurmi, Hanna, Osenova, Petya, Östling, Robert, Øvrelid, Lilja, Paiva, Valeria, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Rademaker, Alexandre, Ramasamy, Loganathan, Real, Livy, Rituma, Laura, Rosa, Rudolf, Saleh, Shadi, Saulīte, Baiba, Schuster, Sebastian, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Smith, Aaron, Spadine, Carolyn, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Tanaka, Takaaki, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Uria, Larraitz, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Wallin, Lars, Wang, Jing Xian, Washington, Jonathan North, Wirén, Mats, Žabokrtský, Zdeněk, Zeldes, Amir, Zeman, Daniel, and Zhu, Hanzhi
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Swedish Sign Language, Ukrainian, Uighur, and Vietnamese
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:: Licence Universal Dependencies v1.4, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.4, and PUB

20. Universal Dependencies 2.0

Creator:: Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Bhat, Riyaz Ahmad, Bick, Eckhard, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Candito, Marie, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Chalub, Fabricio, Choi, Jinho, Çöltekin, Çağrı, Connor, Miriam, Davidson, Elizabeth, de Marneffe, Marie-Catherine, de Paiva, Valeria, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Dwivedi, Puneet, Eli, Marhaba, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Freitas, Cláudia, Gajdošová, Katarína, Galbraith, Daniel, Garcia, Marcos, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grioni, Matias, Grūzītis, Normunds, Guillaume, Bruno, Habash, Nizar, Hajič, Jan, Hà Mỹ, Linh, Haug, Dag, Hladká, Barbora, Hohle, Petter, Ion, Radu, Irimia, Elena, Johannsen, Anders, Jørgensen, Fredrik, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Kotsyba, Natalia, Krek, Simon, Laippala, Veronika, Lê Hồng, Phương, Lenci, Alessandro, Ljubešić, Nikola, Lyashevskaya, Olga, Lynn, Teresa, Makazhanov, Aibek, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Martins, André, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, More, Amir, Mori, Shunsuke, Moskalevskyi, Bohdan, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikolaev, Vitaly, Nurmi, Hanna, Ojala, Stina, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Perrier, Guy, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Rademaker, Alexandre, Ramasamy, Loganathan, Real, Livy, Rituma, Laura, Rosa, Rudolf, Saleh, Shadi, Sanguinetti, Manuela, Saulīte, Baiba, Schuster, Sebastian, Seddah, Djamé, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Sichinava, Dmitry, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Smith, Aaron, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Taji, Dima, Tanaka, Takaaki, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Uria, Larraitz, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Washington, Jonathan North, Žabokrtský, Zdeněk, Zeldes, Amir, Zeman, Daniel, and Zhu, Hanzhi
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, and Urdu
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This release is special in that the treebanks will be used as training/development data in the CoNLL 2017 shared task (http://universaldependencies.org/conll17/). Test data are not released, except for the few treebanks that do not take part in the shared task. 64 treebanks will be in the shared task, and they correspond to the following 45 languages: Ancient Greek, Arabic, Basque, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Gothic, Greek, Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Kazakh, Korean, Latin, Latvian, Norwegian, Old Church Slavonic, Persian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Turkish, Ukrainian, Urdu, Uyghur and Vietnamese. This release fixes a bug in http://hdl.handle.net/11234/1-1976. Changed files: ud-tools-v2.0.tgz (conllu_to_text.pl, conllu_to_conllx.pl; added text_without_spaces.pl), ud-treebanks-conll2017.tgz (fi_ftb-ud-train.txt, he-ud-train.txt, it-ud-train.txt, pt_br-ud-train.txt, es-ud-train.txt) and ud-treebanks-v2.0.tgz (fi_ftb-ud-train.txt, he-ud-train.txt, it-ud-train.txt, pt_br-ud-train.txt, es-ud-train.txt, ar_nyuad-ud-dev.txt, ar_nyuad-ud-test.txt, ar_nyuad-ud-train.txt, cop-ud-dev.txt, cop-ud-test.txt, cop-ud-train.txt, sa-ud-dev.txt, sa-ud-test.txt, sa-ud-train.txt).
Rights:: Licence Universal Dependencies v2.0, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.0, and PUB

11. DaMuEL 1.0: A Large Multilingual Dataset for Entity Linking

12. Morpho-syntactically annotated corpora provided for the PARSEME Shared Task on Semi-Supervised Identification of Verbal Multiword Expressions (edition 1.2)

13. Multilingual static embeddings for Verbal Multiword Expressions trained on PARSEME raw corpora

14. PARSEME corpora annotated for verbal multiword expressions (version 1.3)

15. Plaintext Wikipedia dump 2018

16. UDify Pretrained Model

17. Uniform Meaning Representation

18. Universal Dependencies 1.3

19. Universal Dependencies 1.4

20. Universal Dependencies 2.0

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Creator

Show values starting with

Language

Show values starting with

Publisher

Rights

Show values starting with

Subject

Show values starting with

Type

Original context has metadata only

Harvested from