Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Droganova, Kira , Zeman, Daniel , Kanerva, Jenna , and Ginter, Filip
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
universal dependencies , ellipsis , and gapping
Language:
English , Czech , Finnish , Russian , and Slovak
Description:
Artificially created treebank of elliptical constructions (gapping), in the annotation style of Universal Dependencies. Data taken from UD 2.1 release, and from large web corpora parsed by two parsers. Input data are filtered, sentences are identified where gapping could be applied, then those sentences are transformed, one or more words are omitted, resulting in a sentence with gapping. Details in Droganova et al.: Parse Me if You Can: Artificial Treebanks for Parsing Experiments on Elliptical Constructions, LREC 2018, Miyazaki, Japan.
Rights:
Licence Universal Dependencies v2.1 , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.1 , and PUB
Creator:
Ginter, Filip , Hajič, Jan , Luotolahti, Juhani , Straka, Milan , and Zeman, Daniel
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text , mlmodel , and languageDescription
Subject:
CoNLL 2017 , word embeddings , and automatic annotation
Language:
Multiple languages
Description:
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by UDPipe (http://ufal.mff.cuni.cz/udpipe), together with word embeddings of dimension 100 computed from lowercased texts by word2vec (https://code.google.com/archive/p/word2vec/).
For each language, automatic annotations in CoNLL-U format are provided in a separate archive. The word embeddings for all languages are distributed in one archive.
Note that the CC BY-SA-NC 4.0 license applies to the automatically generated annotations and word embeddings, not to the underlying data, which may have different license and impose additional restrictions.
Update 2018-09-03
===============
Added data in the 4 “surprise languages” from the 2017 ST: Buryat, Kurmanji, North Sami and Upper Sorbian. This has been promised before, during CoNLL-ST 2018 we gave the participants a link to this record saying the data was here. It wasn't, sorry. But now it is.
Rights:
Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) , http://creativecommons.org/licenses/by-nc-sa/4.0/ , and PUB
Creator:
Zeman, Daniel , Potthast, Martin , Straka, Milan , Popel, Martin , Dozat, Timothy , Qi, Peng , Manning, Christopher , Shi, Tianze , Wu, Felix G. , Chen, Xilun , Cheng, Yao , Björkelund, Anders , Falenska, Agnieszka , Yu, Xiang , Kuhn, Jonas , Che, Wanxiang , Guo, Jiang , Wang, Yuxuan , Zheng, Bo , Zhao, Huaipeng , Liu, Yang , Teng, Dechuan , Liu, Ting , Lim, Kyungtae , Poibeau, Thierry , Sato, Motoki , Manabe, Hitoshi , Noji, Hiroshi , Matsumoto, Yuji , Kırnap, Ömer , Önder, Berkay Furkan , Yuret, Deniz , Straková, Jana , Vania, Clara , Zhang, Xingxing , Lopez, Adam , Heinecke, Johannes , Asadullah, Munshi , Kanerva, Jenna , Luotolahti, Juhani , Ginter, Filip , Kuan, Yu , Sofroniev, Pavel , Schill, Erik , Hinrichs, Erhard , Nguyen, Dat Quoc , Dras, Mark , Johnson, Mark , Qian, Xian , Vilares, David , Gómez-Rodríguez, Carlos , Aufrant, Lauriane , Wisniewski, Guillaume , Yvon, François , Dumitrescu, Stefan Daniel , Boroş, Tiberiu , Tufiş, Dan , Das, Ayan , Zaffar, Affan , Sarkar, Sudeshna , Wang, Hao , Zhao, Hai , Zhang, Zhisong , Hornby, Ryan , Taylor, Clark , Park, Jungyeul , de Lhoneux, Miryam , Shao, Yan , Basirat, Ali , Kiperwasser, Eliyahu , Stymne, Sara , Goldberg, Yoav , Nivre, Joakim , Akkuş, Burak Kerim , Azizoglu, Heval , Cakici, Ruket , Moor, Christophe , Merlo, Paola , Henderson, James , Wang, Haozhou , Ji, Tao , Wu, Yuanbin , Lan, Man , de la Clergerie, Eric , Sagot, Benoît , Seddah, Djamé , More, Amir , Tsarfaty, Reut , Kanayama, Hiroshi , Muraoka, Masayasu , Yoshikawa, Katsumasa , Garcia, Marcos , and Gamallo, Pablo
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
dependency parser and parsebank
Language:
Arabic , Bulgarian , Russia Buriat , Czech , Catalan , Church Slavic , Danish , German , Modern Greek (1453-) , English , Spanish , Estonian , Basque , Persian , Finnish , French , Irish , Galician , Gothic , Ancient Greek (to 1453) , Hebrew , Hindi , Croatian , Upper Sorbian , Hungarian , Indonesian , Italian , Japanese , Kazakh , Northern Kurdish , Korean , Latin , Latvian , Dutch , Norwegian , Polish , Portuguese , Romanian , Russian , Slovak , Slovenian , Northern Sami , Swedish , Turkish , Uighur , Ukrainian , Urdu , Vietnamese , and Chinese
Description:
This package contains the system outputs from the CoNLL 2017 Shared Task in Multilingual Parsing from Raw Text to Universal Dependencies.
Rights:
Licence Universal Dependencies v2.0 , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.0 , and PUB
Creator:
Zeman, Daniel , Potthast, Martin , Duthoo, Elie , Mesnard, Olivier , Rybak, Piotr , Wróblewska, Alina , Che, Wanxiang , Liu, Yijia , Wang, Yuxuan , Zheng, Bo , Liu, Ting , Li, Zuchao , He, Shexia , Zhang, Zhuosheng , Zhao, Hai , Wu, Yingting , Tong, Jia-Jun , Nguyen, Dat Quoc , Verspoor, Karin , Wan, Hui , Naseem, Tahira , Lee, Young-Suk , Castelli, Vittorio , Ballesteros, Miguel , Hershcovich, Daniel , Abend, Omri , Rappoport, Ari , Smith, Aaron , Bohnet, Bernd , de Lhoneux, Miryam , Nivre, Joakim , Shao, Yan , Stymne, Sara , Kırnap, Ömer , Dayanık, Erenay , Yuret, Deniz , Kanerva, Jenna , Ginter, Filip , Miekka, Niko , Leino, Akseli , Salakoski, Tapio , Lim, KyungTae , Park, Cheoneum , Lee, Changki , Poibeau, Thierry , Bhat, Riyaz Ahmad , Bhat, Irshad , Bangalore, Srinivas , Qi, Peng , Dozat, Timothy , Zhang, Yuhao , Manning, Christopher , Boroș, Tiberiu , Dumitrescu, Stefan Daniel , Burtica, Ruxandra , Arakelyan, Gor , Hambardzumyan, Karen , Khachatrian, Hrant , Rosa, Rudolf , Mareček, David , Straka, Milan , Seker, Amit , More, Amir , Tsarfaty, Reut , Önder, Berkay Furkan , Gümeli, Can , Jawahar, Ganesh , Muller, Benjamin , Fethi, Amal , Martin, Louis , Villemonte de la Clergerie, Eric , Sagot, Benoît , Seddah, Djamé , Özateş, Şaziye Betül , Özgür, Arzucan , Gungor, Tunga , Öztürk, Balkız , Ji, Tao , Liu, Yufang , Wang, Yijun , Wu, Yuanbin , Lan, Man , Chen, Danlu , Lin, Mengxiao , Hu, Zhifeng , and Qiu, Xipeng
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
parsed data , conllu , and universal dependencies
Language:
Afrikaans , Arabic , Breton , Bulgarian , Russia Buriat , Catalan , Czech , Church Slavic , Danish , German , Modern Greek (1453-) , English , Estonian , Basque , Faroese , Persian , Finnish , French , Old French (842-ca. 1400) , Irish , Galician , Gothic , Ancient Greek (to 1453) , Hebrew , Hindi , Croatian , Upper Sorbian , Hungarian , Armenian , Indonesian , Italian , Japanese , Kazakh , Northern Kurdish , Korean , Latin , Latvian , Dutch , Norwegian , Nigerian Pidgin , Polish , Portuguese , Romanian , Russian , Slovak , Slovenian , Northern Sami , Spanish , Serbian , Swedish , Thai , Turkish , Uighur , Ukrainian , Urdu , Vietnamese , and Chinese
Description:
Test data parsed by systems submitted to the CoNLL 2018 UD parsing shared task.
Rights:
Licence Universal Dependencies v2.2 , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.2 , and PUB
Creator:
Guillaume, Bruno , Ramisch, Carlos , Waszczuk, Jakub , Monti, Johanna , Di Buono, Maria Pia , Sangati, Federico , Speranza, Giulia , Carlino, Carola , Güngör, Tunga , Yirmibeşoğlu, Zeynep , Sak, Haşim , Saraçlar, Murat , Giouli, Voula , Foufi, Vassiliki , Ramisch, Renata , Rademaker, Alexandre , Vale, Oto , Wilkens, Rodrigo , Candito, Marie , Crabbé, Benoît , Segonne, Vincent , Liebeskind, Chaya , Stymne, Sara , Hajič, Jan , Ginter, Filip , Luotolahti, Juhani , Straka, Milan , Zeman, Daniel , Barbu Mititelu, Verginica , Cristescu, Mihaela , Vaidya, Ashwini , Bhatia, Archna , Lichte, Timm , Ehren, Rafael , Jiang, Menghan , Xu, Hongzhi , Walsh, Abigail , Irimia, Elena , and Dowling, Meghan
Publisher:
PARSEME
Type:
text and corpus
Subject:
morphosyntactic annotation , dependency trees , and morphological analysis
Language:
German , Modern Greek (1453-) , Basque , French , Irish , Hebrew , Hindi , Italian , Polish , Portuguese , Romanian , Swedish , Turkish , and Chinese
Description:
This multilingual resource contains corpora for 14 languages, gathered at the occasion of the 1.2 edition of the PARSEME Shared Task on semi-supervised Identification of Verbal MWEs (2020). These corpora were meant to serve as additional "raw" corpora, to help discovering unseen verbal MWEs.
The corpora are provided in CONLL-U (https://universaldependencies.org/format.html) format. They contain morphosyntactic annotations (parts of speech, lemmas, morphological features, and syntactic dependencies). Depending on the language, the information comes from treebanks (mostly Universal Dependencies v2.x) or from automatic parsers trained on UD v2.x treebanks (e.g., UDPipe).
VMWEs include idioms (let the cat out of the bag), light-verb constructions (make a decision), verb-particle constructions (give up), inherently reflexive verbs (help oneself), and multi-verb constructions (make do).
For the 1.2 shared task edition, the data covers 14 languages, for which VMWEs were annotated according to the universal guidelines. The corpora are provided in the cupt format, inspired by the CONLL-U format.
Morphological and syntactic information – not necessarily using UD tagsets – including parts of speech, lemmas, morphological features and/or syntactic dependencies are also provided. Depending on the language, the information comes from treebanks (e.g., Universal Dependencies) or from automatic parsers trained on treebanks (e.g., UDPipe).
This item contains training, development and test data, as well as the evaluation tools used in the PARSEME Shared Task 1.2 (2020). The annotation guidelines are available online: http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.2
Rights:
PARSEME Shared Task Raw Corpus Data (v. 1.2) Agreement , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-mwe-1.2-raw , and PUB
Creator:
Nivre, Joakim , Bosco, Cristina , Choi, Jinho , de Marneffe, Marie-Catherine , Dozat, Timothy , Farkas, Richárd , Foster, Jennifer , Ginter, Filip , Goldberg, Yoav , Hajič, Jan , Kanerva, Jenna , Laippala, Veronika , Lenci, Alessandro , Lynn, Teresa , Manning, Christopher , McDonald, Ryan , Missilä, Anna , Montemagni, Simonetta , Petrov, Slav , Pyysalo, Sampo , Silveira, Natalia , Simi, Maria , Smith, Aaron , Tsarfaty, Reut , Vincze, Veronika , and Zeman, Daniel
Publisher:
Universal Dependencies Consortium
Type:
text and corpus
Subject:
treebank , dependency , syntax , morphology , harmonized annotation , interset , universal tagset , and stanford dependencies
Language:
Czech , German , English , Spanish , Finnish , French , Irish , Italian , Swedish , and Hungarian
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:
Universal Dependencies 1.0 License Set , https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-1.0 , and PUB
Creator:
Agić, Željko , Aranzabe, Maria Jesus , Atutxa, Aitziber , Bosco, Cristina , Choi, Jinho , de Marneffe, Marie-Catherine , Dozat, Timothy , Farkas, Richárd , Foster, Jennifer , Ginter, Filip , Goenaga, Iakes , Gojenola, Koldo , Goldberg, Yoav , Hajič, Jan , Johannsen, Anders Trærup , Kanerva, Jenna , Kuokkala, Juha , Laippala, Veronika , Lenci, Alessandro , Lindén, Krister , Ljubešić, Nikola , Lynn, Teresa , Manning, Christopher , Martínez, Héctor Alonso , McDonald, Ryan , Missilä, Anna , Montemagni, Simonetta , Nivre, Joakim , Nurmi, Hanna , Osenova, Petya , Petrov, Slav , Piitulainen, Jussi , Plank, Barbara , Prokopidis, Prokopis , Pyysalo, Sampo , Seeker, Wolfgang , Seraji, Mojgan , Silveira, Natalia , Simi, Maria , Simov, Kiril , Smith, Aaron , Tsarfaty, Reut , Vincze, Veronika , and Zeman, Daniel
Publisher:
Universal Dependencies Consortium
Type:
text and corpus
Subject:
treebank , dependency syntax , morphology , harmonized annotation , interset , universal tagset , stanford dependencies , and universal dependencies
Language:
Basque , Bulgarian , Croatian , Czech , Danish , English , Finnish , French , German , Modern Greek (1453-) , Hebrew , Hungarian , Indonesian , Irish , Italian , Persian , Spanish , and Swedish
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This is the second release of UD Treebanks, Version 1.1.
Rights:
Licence Universal Dependencies v1.1 , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.1 , and PUB
Creator:
Nivre, Joakim , Agić, Željko , Aranzabe, Maria Jesus , Asahara, Masayuki , Atutxa, Aitziber , Ballesteros, Miguel , Bauer, John , Bengoetxea, Kepa , Bhat, Riyaz Ahmad , Bosco, Cristina , Bowman, Sam , Celano, Giuseppe G. A. , Connor, Miriam , de Marneffe, Marie-Catherine , Diaz de Ilarraza, Arantza , Dobrovoljc, Kaja , Dozat, Timothy , Erjavec, Tomaž , Farkas, Richárd , Foster, Jennifer , Galbraith, Daniel , Ginter, Filip , Goenaga, Iakes , Gojenola, Koldo , Goldberg, Yoav , Gonzales, Berta , Guillaume, Bruno , Hajič, Jan , Haug, Dag , Ion, Radu , Irimia, Elena , Johannsen, Anders , Kanayama, Hiroshi , Kanerva, Jenna , Krek, Simon , Laippala, Veronika , Lenci, Alessandro , Ljubešić, Nikola , Lynn, Teresa , Manning, Christopher , Mărănduc, Cătălina , Mareček, David , Martínez Alonso, Héctor , Mašek, Jan , Matsumoto, Yuji , McDonald, Ryan , Missilä, Anna , Mititelu, Verginica , Miyao, Yusuke , Montemagni, Simonetta , Mori, Shunsuke , Nurmi, Hanna , Osenova, Petya , Øvrelid, Lilja , Pascual, Elena , Passarotti, Marco , Perez, Cenel-Augusto , Petrov, Slav , Piitulainen, Jussi , Plank, Barbara , Popel, Martin , Prokopidis, Prokopis , Pyysalo, Sampo , Ramasamy, Loganathan , Rosa, Rudolf , Saleh, Shadi , Schuster, Sebastian , Seeker, Wolfgang , Seraji, Mojgan , Silveira, Natalia , Simi, Maria , Simionescu, Radu , Simkó, Katalin , Simov, Kiril , Smith, Aaron , Štěpánek, Jan , Suhr, Alane , Szántó, Zsolt , Tanaka, Takaaki , Tsarfaty, Reut , Uematsu, Sumire , Uria, Larraitz , Varga, Viktor , Vincze, Veronika , Žabokrtský, Zdeněk , Zeman, Daniel , and Zhu, Hanzhi
Publisher:
Universal Dependencies Consortium
Type:
text and corpus
Subject:
treebank , dependency , syntax , morphology , harmonized annotation , interset , universal tagset , and stanford dependencies
Language:
Ancient Greek (to 1453) , Arabic , Basque , Bulgarian , Croatian , Czech , Danish , Dutch , English , Estonian , Finnish , French , German , Gothic , Modern Greek (1453-) , Hebrew , Hindi , Hungarian , Indonesian , Irish , Italian , Japanese , Latin , Norwegian , Church Slavic , Persian , Polish , Portuguese , Romanian , Slovenian , Spanish , Swedish , and Tamil
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:
Licence Universal Dependencies v1.2 , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.2 , and PUB
Creator:
Nivre, Joakim , Agić, Željko , Ahrenberg, Lars , Aranzabe, Maria Jesus , Asahara, Masayuki , Atutxa, Aitziber , Ballesteros, Miguel , Bauer, John , Bengoetxea, Kepa , Berzak, Yevgeni , Bhat, Riyaz Ahmad , Bosco, Cristina , Bouma, Gosse , Bowman, Sam , Cebiroğlu Eryiğit, Gülşen , Celano, Giuseppe G. A. , Çöltekin, Çağrı , Connor, Miriam , de Marneffe, Marie-Catherine , Diaz de Ilarraza, Arantza , Dobrovoljc, Kaja , Dozat, Timothy , Droganova, Kira , Erjavec, Tomaž , Farkas, Richárd , Foster, Jennifer , Galbraith, Daniel , Garza, Sebastian , Ginter, Filip , Goenaga, Iakes , Gojenola, Koldo , Gokirmak, Memduh , Goldberg, Yoav , Gómez Guinovart, Xavier , Gonzáles Saavedra, Berta , Grūzītis, Normunds , Guillaume, Bruno , Hajič, Jan , Haug, Dag , Hladká, Barbora , Ion, Radu , Irimia, Elena , Johannsen, Anders , Kaşıkara, Hüner , Kanayama, Hiroshi , Kanerva, Jenna , Katz, Boris , Kenney, Jessica , Krek, Simon , Laippala, Veronika , Lam, Lucia , Lenci, Alessandro , Ljubešić, Nikola , Lyashevskaya, Olga , Lynn, Teresa , Makazhanov, Aibek , Manning, Christopher , Mărănduc, Cătălina , Mareček, David , Martínez Alonso, Héctor , Mašek, Jan , Matsumoto, Yuji , McDonald, Ryan , Missilä, Anna , Mititelu, Verginica , Miyao, Yusuke , Montemagni, Simonetta , Mori, Keiko Sophie , Mori, Shunsuke , Muischnek, Kadri , Mustafina, Nina , Müürisep, Kaili , Nikolaev, Vitaly , Nurmi, Hanna , Osenova, Petya , Øvrelid, Lilja , Pascual, Elena , Passarotti, Marco , Perez, Cenel-Augusto , Petrov, Slav , Piitulainen, Jussi , Plank, Barbara , Popel, Martin , Pretkalniņa, Lauma , Prokopidis, Prokopis , Puolakainen, Tiina , Pyysalo, Sampo , Ramasamy, Loganathan , Rituma, Laura , Rosa, Rudolf , Saleh, Shadi , Saulīte, Baiba , Schuster, Sebastian , Seeker, Wolfgang , Seraji, Mojgan , Shakurova, Lena , Shen, Mo , Silveira, Natalia , Simi, Maria , Simionescu, Radu , Simkó, Katalin , Simov, Kiril , Smith, Aaron , Spadine, Carolyn , Suhr, Alane , Sulubacak, Umut , Szántó, Zsolt , Tanaka, Takaaki , Tsarfaty, Reut , Tyers, Francis , Uematsu, Sumire , Uria, Larraitz , van Noord, Gertjan , Varga, Viktor , Vincze, Veronika , Wang, Jing Xian , Washington, Jonathan North , Žabokrtský, Zdeněk , Zeman, Daniel , and Zhu, Hanzhi
Publisher:
Universal Dependencies Consortium
Type:
text and corpus
Subject:
treebank , dependency , syntax , morphology , harmonized annotation , interset , universal tagset , and stanford dependencies
Language:
Ancient Greek (to 1453) , Arabic , Basque , Bulgarian , Croatian , Czech , Danish , Dutch , English , Estonian , Finnish , French , German , Gothic , Modern Greek (1453-) , Hebrew , Hindi , Hungarian , Indonesian , Irish , Italian , Japanese , Latin , Norwegian , Church Slavic , Persian , Polish , Portuguese , Romanian , Slovenian , Spanish , Swedish , Tamil , Catalan , Chinese , Galician , Kazakh , Latvian , Russian , and Turkish
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:
Licence Universal Dependencies v1.3 , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.3 , and PUB
Creator:
Nivre, Joakim , Agić, Željko , Ahrenberg, Lars , Aranzabe, Maria Jesus , Asahara, Masayuki , Atutxa, Aitziber , Ballesteros, Miguel , Bauer, John , Bengoetxea, Kepa , Berzak, Yevgeni , Bhat, Riyaz Ahmad , Bick, Eckhard , Börstell, Carl , Bosco, Cristina , Bouma, Gosse , Bowman, Sam , Cebiroğlu Eryiğit, Gülşen , Celano, Giuseppe G. A. , Chalub, Fabricio , Çöltekin, Çağrı , Connor, Miriam , Davidson, Elizabeth , de Marneffe, Marie-Catherine , Diaz de Ilarraza, Arantza , Dobrovoljc, Kaja , Dozat, Timothy , Droganova, Kira , Dwivedi, Puneet , Eli, Marhaba , Erjavec, Tomaž , Farkas, Richárd , Foster, Jennifer , Freitas, Claudia , Gajdošová, Katarína , Galbraith, Daniel , Garcia, Marcos , Gärdenfors, Moa , Garza, Sebastian , Ginter, Filip , Goenaga, Iakes , Gojenola, Koldo , Gökırmak, Memduh , Goldberg, Yoav , Gómez Guinovart, Xavier , Gonzáles Saavedra, Berta , Grioni, Matias , Grūzītis, Normunds , Guillaume, Bruno , Hajič, Jan , Hà Mỹ, Linh , Haug, Dag , Hladká, Barbora , Ion, Radu , Irimia, Elena , Johannsen, Anders , Jørgensen, Fredrik , Kaşıkara, Hüner , Kanayama, Hiroshi , Kanerva, Jenna , Katz, Boris , Kenney, Jessica , Kotsyba, Natalia , Krek, Simon , Laippala, Veronika , Lam, Lucia , Lê Hồng, Phương , Lenci, Alessandro , Ljubešić, Nikola , Lyashevskaya, Olga , Lynn, Teresa , Makazhanov, Aibek , Manning, Christopher , Mărănduc, Cătălina , Mareček, David , Martínez Alonso, Héctor , Martins, André , Mašek, Jan , Matsumoto, Yuji , McDonald, Ryan , Missilä, Anna , Mititelu, Verginica , Miyao, Yusuke , Montemagni, Simonetta , Mori, Keiko Sophie , Mori, Shunsuke , Moskalevskyi, Bohdan , Muischnek, Kadri , Mustafina, Nina , Müürisep, Kaili , Nguyễn Thị, Lương , Nguyễn Thị Minh, Huyền , Nikolaev, Vitaly , Nurmi, Hanna , Osenova, Petya , Östling, Robert , Øvrelid, Lilja , Paiva, Valeria , Pascual, Elena , Passarotti, Marco , Perez, Cenel-Augusto , Petrov, Slav , Piitulainen, Jussi , Plank, Barbara , Popel, Martin , Pretkalniņa, Lauma , Prokopidis, Prokopis , Puolakainen, Tiina , Pyysalo, Sampo , Rademaker, Alexandre , Ramasamy, Loganathan , Real, Livy , Rituma, Laura , Rosa, Rudolf , Saleh, Shadi , Saulīte, Baiba , Schuster, Sebastian , Seeker, Wolfgang , Seraji, Mojgan , Shakurova, Lena , Shen, Mo , Silveira, Natalia , Simi, Maria , Simionescu, Radu , Simkó, Katalin , Šimková, Mária , Simov, Kiril , Smith, Aaron , Spadine, Carolyn , Suhr, Alane , Sulubacak, Umut , Szántó, Zsolt , Tanaka, Takaaki , Tsarfaty, Reut , Tyers, Francis , Uematsu, Sumire , Uria, Larraitz , van Noord, Gertjan , Varga, Viktor , Vincze, Veronika , Wallin, Lars , Wang, Jing Xian , Washington, Jonathan North , Wirén, Mats , Žabokrtský, Zdeněk , Zeldes, Amir , Zeman, Daniel , and Zhu, Hanzhi
Publisher:
Universal Dependencies Consortium
Type:
text and corpus
Subject:
treebank , dependency , syntax , morphology , harmonized annotation , interset , universal tagset , and stanford dependencies
Language:
Ancient Greek (to 1453) , Arabic , Basque , Bulgarian , Croatian , Czech , Danish , Dutch , English , Estonian , Finnish , French , German , Gothic , Modern Greek (1453-) , Hebrew , Hindi , Hungarian , Indonesian , Irish , Italian , Japanese , Latin , Norwegian , Church Slavic , Persian , Polish , Portuguese , Romanian , Slovenian , Spanish , Swedish , Tamil , Catalan , Chinese , Galician , Kazakh , Latvian , Russian , Turkish , Coptic , Sanskrit , Slovak , Swedish Sign Language , Ukrainian , Uighur , and Vietnamese
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:
Licence Universal Dependencies v1.4 , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.4 , and PUB