Number of results to display per page
Search Results
52. The Karjalainen Corpus
- Publisher:
- University of Joensuu
- Type:
- corpus
- Language:
- Finnish
- Description:
- computer corpus of Finnish newspaper texts of the 1990s (newspaper Karjalainen, Joensuu)
- Rights:
- Not specified
53. The National Certificates corpus
- Publisher:
- Centre for Applied Language Studies, University of Jyväskylä
- Type:
- corpus
- Language:
- English, Finnish, French, German, Italian, Russian, Spanish, and Swedish
- Description:
- The NC test results, background information, speaking and writing performances in 9 foreign / second languages. A web-based data base (html files).
- Rights:
- Not specified
54. UDify Pretrained Model
- Creator:
- Kondratyuk, Dan and Straka, Milan
- Publisher:
- Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
- Type:
- tool and toolService
- Subject:
- syntax, dependency parser, and universal dependencies
- Language:
- Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, Urdu, Russia Buriat, Northern Kurdish, Northern Sami, Upper Sorbian, Afrikaans, Yue Chinese, Marathi, Serbian, Swedish Sign Language, Telugu, Amharic, Armenian, Breton, Faroese, Komi-Zyrian, Nigerian Pidgin, Old French (842-ca. 1400), Tagalog, Thai, Warlpiri, Yoruba, Akkadian, Bambara, Erzya, and Maltese
- Description:
- Pretrained model weights for the UDify model, and extracted BERT weights in pytorch-transformers format. Note that these weights slightly differ from those used in the paper.
- Rights:
- Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), http://creativecommons.org/licenses/by-sa/4.0/, and PUB
55. Universal Dependencies 1.0
- Creator:
- Nivre, Joakim, Bosco, Cristina, Choi, Jinho, de Marneffe, Marie-Catherine, Dozat, Timothy, Farkas, Richárd, Foster, Jennifer, Ginter, Filip, Goldberg, Yoav, Hajič, Jan, Kanerva, Jenna, Laippala, Veronika, Lenci, Alessandro, Lynn, Teresa, Manning, Christopher, McDonald, Ryan, Missilä, Anna, Montemagni, Simonetta, Petrov, Slav, Pyysalo, Sampo, Silveira, Natalia, Simi, Maria, Smith, Aaron, Tsarfaty, Reut, Vincze, Veronika, and Zeman, Daniel
- Publisher:
- Universal Dependencies Consortium
- Type:
- text and corpus
- Subject:
- treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
- Language:
- Czech, German, English, Spanish, Finnish, French, Irish, Italian, Swedish, and Hungarian
- Description:
- Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
- Rights:
- Universal Dependencies 1.0 License Set, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-1.0, and PUB
56. Universal Dependencies 1.1
- Creator:
- Agić, Željko, Aranzabe, Maria Jesus, Atutxa, Aitziber, Bosco, Cristina, Choi, Jinho, de Marneffe, Marie-Catherine, Dozat, Timothy, Farkas, Richárd, Foster, Jennifer, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Goldberg, Yoav, Hajič, Jan, Johannsen, Anders Trærup, Kanerva, Jenna, Kuokkala, Juha, Laippala, Veronika, Lenci, Alessandro, Lindén, Krister, Ljubešić, Nikola, Lynn, Teresa, Manning, Christopher, Martínez, Héctor Alonso, McDonald, Ryan, Missilä, Anna, Montemagni, Simonetta, Nivre, Joakim, Nurmi, Hanna, Osenova, Petya, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Prokopidis, Prokopis, Pyysalo, Sampo, Seeker, Wolfgang, Seraji, Mojgan, Silveira, Natalia, Simi, Maria, Simov, Kiril, Smith, Aaron, Tsarfaty, Reut, Vincze, Veronika, and Zeman, Daniel
- Publisher:
- Universal Dependencies Consortium
- Type:
- text and corpus
- Subject:
- treebank, dependency syntax, morphology, harmonized annotation, interset, universal tagset, stanford dependencies, and universal dependencies
- Language:
- Basque, Bulgarian, Croatian, Czech, Danish, English, Finnish, French, German, Modern Greek (1453-), Hebrew, Hungarian, Indonesian, Irish, Italian, Persian, Spanish, and Swedish
- Description:
- Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This is the second release of UD Treebanks, Version 1.1.
- Rights:
- Licence Universal Dependencies v1.1, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.1, and PUB
57. Universal Dependencies 1.2
- Creator:
- Nivre, Joakim, Agić, Željko, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Bhat, Riyaz Ahmad, Bosco, Cristina, Bowman, Sam, Celano, Giuseppe G. A., Connor, Miriam, de Marneffe, Marie-Catherine, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Galbraith, Daniel, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Goldberg, Yoav, Gonzales, Berta, Guillaume, Bruno, Hajič, Jan, Haug, Dag, Ion, Radu, Irimia, Elena, Johannsen, Anders, Kanayama, Hiroshi, Kanerva, Jenna, Krek, Simon, Laippala, Veronika, Lenci, Alessandro, Ljubešić, Nikola, Lynn, Teresa, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, Mori, Shunsuke, Nurmi, Hanna, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Prokopidis, Prokopis, Pyysalo, Sampo, Ramasamy, Loganathan, Rosa, Rudolf, Saleh, Shadi, Schuster, Sebastian, Seeker, Wolfgang, Seraji, Mojgan, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Simov, Kiril, Smith, Aaron, Štěpánek, Jan, Suhr, Alane, Szántó, Zsolt, Tanaka, Takaaki, Tsarfaty, Reut, Uematsu, Sumire, Uria, Larraitz, Varga, Viktor, Vincze, Veronika, Žabokrtský, Zdeněk, Zeman, Daniel, and Zhu, Hanzhi
- Publisher:
- Universal Dependencies Consortium
- Type:
- text and corpus
- Subject:
- treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
- Language:
- Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, and Tamil
- Description:
- Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
- Rights:
- Licence Universal Dependencies v1.2, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.2, and PUB
58. Universal Dependencies 1.3
- Creator:
- Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Berzak, Yevgeni, Bhat, Riyaz Ahmad, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Çöltekin, Çağrı, Connor, Miriam, de Marneffe, Marie-Catherine, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Galbraith, Daniel, Garza, Sebastian, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gokirmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grūzītis, Normunds, Guillaume, Bruno, Hajič, Jan, Haug, Dag, Hladká, Barbora, Ion, Radu, Irimia, Elena, Johannsen, Anders, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Katz, Boris, Kenney, Jessica, Krek, Simon, Laippala, Veronika, Lam, Lucia, Lenci, Alessandro, Ljubešić, Nikola, Lyashevskaya, Olga, Lynn, Teresa, Makazhanov, Aibek, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, Mori, Keiko Sophie, Mori, Shunsuke, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nikolaev, Vitaly, Nurmi, Hanna, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Ramasamy, Loganathan, Rituma, Laura, Rosa, Rudolf, Saleh, Shadi, Saulīte, Baiba, Schuster, Sebastian, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Simov, Kiril, Smith, Aaron, Spadine, Carolyn, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Tanaka, Takaaki, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Uria, Larraitz, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Wang, Jing Xian, Washington, Jonathan North, Žabokrtský, Zdeněk, Zeman, Daniel, and Zhu, Hanzhi
- Publisher:
- Universal Dependencies Consortium
- Type:
- text and corpus
- Subject:
- treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
- Language:
- Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, and Turkish
- Description:
- Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
- Rights:
- Licence Universal Dependencies v1.3, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.3, and PUB
59. Universal Dependencies 1.4
- Creator:
- Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Berzak, Yevgeni, Bhat, Riyaz Ahmad, Bick, Eckhard, Börstell, Carl, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Chalub, Fabricio, Çöltekin, Çağrı, Connor, Miriam, Davidson, Elizabeth, de Marneffe, Marie-Catherine, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Dwivedi, Puneet, Eli, Marhaba, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Freitas, Claudia, Gajdošová, Katarína, Galbraith, Daniel, Garcia, Marcos, Gärdenfors, Moa, Garza, Sebastian, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grioni, Matias, Grūzītis, Normunds, Guillaume, Bruno, Hajič, Jan, Hà Mỹ, Linh, Haug, Dag, Hladká, Barbora, Ion, Radu, Irimia, Elena, Johannsen, Anders, Jørgensen, Fredrik, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Katz, Boris, Kenney, Jessica, Kotsyba, Natalia, Krek, Simon, Laippala, Veronika, Lam, Lucia, Lê Hồng, Phương, Lenci, Alessandro, Ljubešić, Nikola, Lyashevskaya, Olga, Lynn, Teresa, Makazhanov, Aibek, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Martins, André, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, Mori, Keiko Sophie, Mori, Shunsuke, Moskalevskyi, Bohdan, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikolaev, Vitaly, Nurmi, Hanna, Osenova, Petya, Östling, Robert, Øvrelid, Lilja, Paiva, Valeria, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Rademaker, Alexandre, Ramasamy, Loganathan, Real, Livy, Rituma, Laura, Rosa, Rudolf, Saleh, Shadi, Saulīte, Baiba, Schuster, Sebastian, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Smith, Aaron, Spadine, Carolyn, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Tanaka, Takaaki, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Uria, Larraitz, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Wallin, Lars, Wang, Jing Xian, Washington, Jonathan North, Wirén, Mats, Žabokrtský, Zdeněk, Zeldes, Amir, Zeman, Daniel, and Zhu, Hanzhi
- Publisher:
- Universal Dependencies Consortium
- Type:
- text and corpus
- Subject:
- treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
- Language:
- Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Swedish Sign Language, Ukrainian, Uighur, and Vietnamese
- Description:
- Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
- Rights:
- Licence Universal Dependencies v1.4, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.4, and PUB
60. Universal Dependencies 2.0
- Creator:
- Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Bhat, Riyaz Ahmad, Bick, Eckhard, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Candito, Marie, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Chalub, Fabricio, Choi, Jinho, Çöltekin, Çağrı, Connor, Miriam, Davidson, Elizabeth, de Marneffe, Marie-Catherine, de Paiva, Valeria, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Dwivedi, Puneet, Eli, Marhaba, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Freitas, Cláudia, Gajdošová, Katarína, Galbraith, Daniel, Garcia, Marcos, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grioni, Matias, Grūzītis, Normunds, Guillaume, Bruno, Habash, Nizar, Hajič, Jan, Hà Mỹ, Linh, Haug, Dag, Hladká, Barbora, Hohle, Petter, Ion, Radu, Irimia, Elena, Johannsen, Anders, Jørgensen, Fredrik, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Kotsyba, Natalia, Krek, Simon, Laippala, Veronika, Lê Hồng, Phương, Lenci, Alessandro, Ljubešić, Nikola, Lyashevskaya, Olga, Lynn, Teresa, Makazhanov, Aibek, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Martins, André, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, More, Amir, Mori, Shunsuke, Moskalevskyi, Bohdan, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikolaev, Vitaly, Nurmi, Hanna, Ojala, Stina, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Perrier, Guy, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Rademaker, Alexandre, Ramasamy, Loganathan, Real, Livy, Rituma, Laura, Rosa, Rudolf, Saleh, Shadi, Sanguinetti, Manuela, Saulīte, Baiba, Schuster, Sebastian, Seddah, Djamé, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Sichinava, Dmitry, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Smith, Aaron, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Taji, Dima, Tanaka, Takaaki, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Uria, Larraitz, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Washington, Jonathan North, Žabokrtský, Zdeněk, Zeldes, Amir, Zeman, Daniel, and Zhu, Hanzhi
- Publisher:
- Universal Dependencies Consortium
- Type:
- text and corpus
- Subject:
- treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
- Language:
- Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, and Urdu
- Description:
- Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This release is special in that the treebanks will be used as training/development data in the CoNLL 2017 shared task (http://universaldependencies.org/conll17/). Test data are not released, except for the few treebanks that do not take part in the shared task. 64 treebanks will be in the shared task, and they correspond to the following 45 languages: Ancient Greek, Arabic, Basque, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Gothic, Greek, Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Kazakh, Korean, Latin, Latvian, Norwegian, Old Church Slavonic, Persian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Turkish, Ukrainian, Urdu, Uyghur and Vietnamese. This release fixes a bug in http://hdl.handle.net/11234/1-1976. Changed files: ud-tools-v2.0.tgz (conllu_to_text.pl, conllu_to_conllx.pl; added text_without_spaces.pl), ud-treebanks-conll2017.tgz (fi_ftb-ud-train.txt, he-ud-train.txt, it-ud-train.txt, pt_br-ud-train.txt, es-ud-train.txt) and ud-treebanks-v2.0.tgz (fi_ftb-ud-train.txt, he-ud-train.txt, it-ud-train.txt, pt_br-ud-train.txt, es-ud-train.txt, ar_nyuad-ud-dev.txt, ar_nyuad-ud-test.txt, ar_nyuad-ud-train.txt, cop-ud-dev.txt, cop-ud-test.txt, cop-ud-train.txt, sa-ud-dev.txt, sa-ud-test.txt, sa-ud-train.txt).
- Rights:
- Licence Universal Dependencies v2.0, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.0, and PUB