« Previous |
1 - 10 of 22
|
Next »
Number of results to display per page
Search Results
2. A Human-Annotated Dataset for Language Modeling and Named Entity Recognition in Medieval Documents (2023-01-05)
- Creator:
- Novotný, Vít, Luger, Kristýna, Štefánik, Michal, Vrabcová, Tereza, and Horák, Aleš
- Publisher:
- Masaryk University, Brno
- Type:
- text and corpus
- Subject:
- NER, named entity recognition, and Medieval
- Language:
- Czech, English, German, and Latin
- Description:
- This is an open dataset of sentences from 19th and 20th century letterpress reprints of documents from the Hussite era. The dataset contains a corpus for language modeling and human annotations for named entity recognition (NER).
- Rights:
- Public Domain Dedication (CC Zero), http://creativecommons.org/publicdomain/zero/1.0/, and PUB
3. A Human-Annotated Dataset of Scanned Images and OCR Texts from Medieval Documents
- Creator:
- Novotný, Vít, Seidlová, Kristýna, Vrabcová, Tereza, and Horák, Aleš
- Publisher:
- Masaryk University, Brno
- Type:
- image and corpus
- Subject:
- ocr, optical character recognition, language identification, image super-resolution, sr, and Medieval
- Language:
- German, Czech, Latin, and English
- Description:
- This is an open dataset of scanned images and OCR texts from 19th and 20th century letterpress reprints of documents from the Hussite era. The dataset contains human annotations for layout analysis, OCR evaluation, and language identification.
- Rights:
- Public Domain Dedication (CC Zero), http://creativecommons.org/publicdomain/zero/1.0/, and PUB
4. A Human-Annotated Dataset of Scanned Images and OCR Texts from Medieval Documents: Supplementary Materials
- Creator:
- Novotný, Vít and Horák, Aleš
- Publisher:
- Masaryk University, Brno
- Type:
- text and corpus
- Subject:
- ocr, optical character recognition, language identification, image super-resolution, sr, and Medieval
- Language:
- Czech, English, German, and Latin
- Description:
- These are supplementary materials for an open dataset of scanned images and OCR texts from 19th and 20th century letterpress reprints of documents from the Hussite era. The dataset contains human annotations for layout analysis, OCR evaluation, and language identification and is available at http://hdl.handle.net/11234/1-4615. These supplementary materials contain OCR texts from different OCR engines for book pages for which we have both high-resolution scanned images and annotations for OCR evaluation.
- Rights:
- Public Domain Dedication (CC Zero), http://creativecommons.org/publicdomain/zero/1.0/, and PUB
5. Deep Universal Dependencies 2.6
- Creator:
- Zeman, Daniel and Droganova, Kira
- Publisher:
- Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
- Type:
- text and corpus
- Subject:
- semantic dependency and universal dependencies
- Language:
- Afrikaans, Assyrian Neo-Aramaic, Akkadian, Amharic, Arabic, Belarusian, Breton, Bulgarian, Russia Buriat, Catalan, Czech, Church Slavic, Mandarin Chinese, Coptic, Welsh, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Faroese, Finnish, French, Irish, Gothic, Ancient Greek (to 1453), Mbyá Guaraní, Hebrew, Hindi, Croatian, Upper Sorbian, Hungarian, Armenian, Indonesian, Italian, Japanese, Kazakh, Northern Kurdish, Korean, Komi-Zyrian, Karelian, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Erzya, Dutch, Norwegian, Old Russian, Nigerian Pidgin, Polish, Portuguese, Romanian, Russian, Sanskrit, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Warlpiri, Wolof, Yoruba, Galician, Bhojpuri, Komi-Permyak, Livvi, Moksha, Scottish Gaelic, Skolt Sami, Icelandic, Albanian, and Persian
- Description:
- Deep Universal Dependencies is a collection of treebanks derived semi-automatically from Universal Dependencies (http://hdl.handle.net/11234/1-3226). It contains additional deep-syntactic and semantic annotations. Version of Deep UD corresponds to the version of UD it is based on. Note however that some UD treebanks have been omitted from Deep UD.
- Rights:
- Licence Universal Dependencies v2.6, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.6, and PUB
6. Deep Universal Dependencies 2.7
- Creator:
- Zeman, Daniel and Droganova, Kira
- Publisher:
- Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
- Type:
- text and corpus
- Subject:
- semantic dependency and universal dependencies
- Language:
- Afrikaans, Assyrian Neo-Aramaic, Akkadian, Amharic, Arabic, Belarusian, Breton, Bulgarian, Russia Buriat, Catalan, Czech, Church Slavic, Mandarin Chinese, Coptic, Welsh, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Faroese, Finnish, French, Irish, Gothic, Ancient Greek (to 1453), Mbyá Guaraní, Hebrew, Hindi, Croatian, Upper Sorbian, Hungarian, Armenian, Indonesian, Italian, Japanese, Kazakh, Northern Kurdish, Korean, Komi-Zyrian, Karelian, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Erzya, Dutch, Norwegian, Old Russian, Nigerian Pidgin, Polish, Portuguese, Romanian, Russian, Sanskrit, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Warlpiri, Wolof, Yoruba, Galician, Bhojpuri, Komi-Permyak, Livvi, Moksha, Scottish Gaelic, Skolt Sami, Icelandic, Albanian, Persian, Akuntsu, Apurinã, Khunsari, Manx, Mundurukú, Nayini, Soi, South Levantine Arabic, and Tupinambá
- Description:
- Deep Universal Dependencies is a collection of treebanks derived semi-automatically from Universal Dependencies (http://hdl.handle.net/11234/1-3424). It contains additional deep-syntactic and semantic annotations. Version of Deep UD corresponds to the version of UD it is based on. Note however that some UD treebanks have been omitted from Deep UD.
- Rights:
- Licence Universal Dependencies v2.7, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.7, and PUB
7. Deep Universal Dependencies 2.8
- Creator:
- Zeman, Daniel and Droganova, Kira
- Publisher:
- Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
- Type:
- text and corpus
- Subject:
- semantic dependency and universal dependencies
- Language:
- Afrikaans, Assyrian Neo-Aramaic, Akkadian, Amharic, Arabic, Belarusian, Breton, Bulgarian, Russia Buriat, Catalan, Czech, Church Slavic, Mandarin Chinese, Coptic, Welsh, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Faroese, Finnish, French, Irish, Gothic, Ancient Greek (to 1453), Mbyá Guaraní, Hebrew, Hindi, Croatian, Upper Sorbian, Hungarian, Armenian, Indonesian, Italian, Japanese, Kazakh, Northern Kurdish, Korean, Komi-Zyrian, Karelian, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Erzya, Dutch, Norwegian, Old Russian, Nigerian Pidgin, Polish, Portuguese, Romanian, Russian, Sanskrit, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Warlpiri, Wolof, Yoruba, Galician, Bhojpuri, Komi-Permyak, Livvi, Moksha, Scottish Gaelic, Skolt Sami, Icelandic, Albanian, Persian, Akuntsu, Apurinã, Khunsari, Manx, Mundurukú, Nayini, Soi, South Levantine Arabic, Tupinambá, Beja, Western Frisian, Urubú-Kaapor, Kangri, K'iche', Low German, Makuráp, Western Armenian, and Central Siberian Yupik
- Description:
- Deep Universal Dependencies is a collection of treebanks derived semi-automatically from Universal Dependencies (http://hdl.handle.net/11234/1-3687). It contains additional deep-syntactic and semantic annotations. Version of Deep UD corresponds to the version of UD it is based on. Note however that some UD treebanks have been omitted from Deep UD.
- Rights:
- Licence Universal Dependencies v2.8, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.8, and PUB
8. On-line Dictionary of medieval latin in the Czech lands
- Creator:
- Ctibor, Jan and Nývlt, Pavel
- Publisher:
- Institute of Philosophy of the Czech Academy of Sciences
- Type:
- text, lexicon, and lexicalConceptualResource
- Subject:
- dictionary, latin, Medieval, digital humanities, lexicography, and Medieval Latin
- Language:
- Latin and Czech
- Description:
- The Dictionary of Medieval Latin in the Czech Lands registers and explains the vocabulary of Medieval Latin as used in the Czech lands since the beginnings of Latin writing in this area (from about 1000 CE) to 1500 CE, so far covering the letters A-M. For more information about the Dictionary, see the webpage of the Department of Medieval Lexicography of the Institute of Philosophy of Czech Academy of Sciences. The data uploaded present the on-line version of the dictionary (API and XML data), making it possible to put the application into operation at a localhost.
- Rights:
- Dictionary of Medieval Latin in the Czech Lands - digital version 2.2 License Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/license-lb, and ACA
9. Universal Dependencies 2.10
- Creator:
- Zeman, Daniel, Nivre, Joakim, Abrams, Mitchell, Ackermann, Elia, Aepli, Noëmi, Aghaei, Hamid, Agić, Željko, Ahmadi, Amir, Ahrenberg, Lars, Ajede, Chika Kennedy, Aleksandravičiūtė, Gabrielė, Alfina, Ika, Algom, Avner, Andersen, Erik, Antonsen, Lene, Aplonova, Katya, Aquino, Angelina, Aragon, Carolina, Aranes, Glyd, Aranzabe, Maria Jesus, Arıcan, Bilge Nas, Arnardóttir, Þórunn, Arutie, Gashaw, Arwidarasti, Jessica Naraiswari, Asahara, Masayuki, Aslan, Deniz Baran, Asmazoğlu, Cengiz, Ateyah, Luma, Atmaca, Furkan, Attia, Mohammed, Atutxa, Aitziber, Augustinus, Liesbeth, Badmaeva, Elena, Balasubramani, Keerthana, Ballesteros, Miguel, Banerjee, Esha, Bank, Sebastian, Barbu Mititelu, Verginica, Barkarson, Starkaður, Basile, Rodolfo, Basmov, Victoria, Batchelor, Colin, Bauer, John, Bedir, Seyyit Talha, Bengoetxea, Kepa, Ben Moshe, Yifat, Berk, Gözde, Berzak, Yevgeni, Bhat, Irshad Ahmad, Bhat, Riyaz Ahmad, Biagetti, Erica, Bick, Eckhard, Bielinskienė, Agnė, Bjarnadóttir, Kristín, Blokland, Rogier, Bobicev, Victoria, Boizou, Loïc, Borges Völker, Emanuel, Börstell, Carl, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Boyd, Adriane, Braggaar, Anouck, Brokaitė, Kristina, Burchardt, Aljoscha, Candito, Marie, Caron, Bernard, Caron, Gauthier, Cassidy, Lauren, Cavalcanti, Tatiana, Cebiroğlu Eryiğit, Gülşen, Cecchini, Flavio Massimiliano, Celano, Giuseppe G. A., Čéplö, Slavomír, Cesur, Neslihan, Cetin, Savas, Çetinoğlu, Özlem, Chalub, Fabricio, Chauhan, Shweta, Chi, Ethan, Chika, Taishi, Cho, Yongseok, Choi, Jinho, Chun, Jayeol, Chung, Juyeon, Cignarella, Alessandra T., Cinková, Silvie, Collomb, Aurélie, Çöltekin, Çağrı, Connor, Miriam, Corbetta, Daniela, Courtin, Marine, Cristescu, Mihaela, Daniel, Philemon, Davidson, Elizabeth, Dehouck, Mathieu, de Laurentiis, Martina, de Marneffe, Marie-Catherine, de Paiva, Valeria, Derin, Mehmet Oguz, de Souza, Elvis, Diaz de Ilarraza, Arantza, Dickerson, Carly, Dinakaramani, Arawinda, Di Nuovo, Elisa, Dione, Bamba, Dirix, Peter, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Dwivedi, Puneet, Eckhoff, Hanne, Eiche, Sandra, Eli, Marhaba, Elkahky, Ali, Ephrem, Binyam, Erina, Olga, Erjavec, Tomaž, Etienne, Aline, Evelyn, Wograine, Facundes, Sidney, Farkas, Richárd, Favero, Federica, Ferdaousi, Jannatul, Fernanda, Marília, Fernandez Alcalde, Hector, Foster, Jennifer, Freitas, Cláudia, Fujita, Kazunori, Gajdošová, Katarína, Galbraith, Daniel, Gamba, Federica, Garcia, Marcos, Gärdenfors, Moa, Garza, Sebastian, Gerardi, Fabrício Ferraz, Gerdes, Kim, Ginter, Filip, Godoy, Gustavo, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, González Saavedra, Berta, Griciūtė, Bernadeta, Grioni, Matias, Grobol, Loïc, Grūzītis, Normunds, Guillaume, Bruno, Guillot-Barbance, Céline, Güngör, Tunga, Habash, Nizar, Hafsteinsson, Hinrik, Hajič, Jan, Hajič jr., Jan, Hämäläinen, Mika, Hà Mỹ, Linh, Han, Na-Rae, Hanifmuti, Muhammad Yudistira, Harada, Takahiro, Hardwick, Sam, Harris, Kim, Haug, Dag, Heinecke, Johannes, Hellwig, Oliver, Hennig, Felix, Hladká, Barbora, Hlaváčová, Jaroslava, Hociung, Florinel, Hohle, Petter, Hwang, Jena, Ikeda, Takumi, Ingason, Anton Karl, Ion, Radu, Irimia, Elena, Ishola, Ọlájídé, Ito, Kaoru, Jannat, Siratun, Jelínek, Tomáš, Jha, Apoorva, Johannsen, Anders, Jónsdóttir, Hildur, Jørgensen, Fredrik, Juutinen, Markus, K, Sarveswaran, Kaşıkara, Hüner, Kaasen, Andre, Kabaeva, Nadezhda, Kahane, Sylvain, Kanayama, Hiroshi, Kanerva, Jenna, Kara, Neslihan, Karahóǧa, Ritván, Katz, Boris, Kayadelen, Tolga, Kenney, Jessica, Kettnerová, Václava, Kirchner, Jesse, Klementieva, Elena, Klyachko, Elena, Köhn, Arne, Köksal, Abdullatif, Kopacewicz, Kamil, Korkiakangas, Timo, Köse, Mehmet, Kotsyba, Natalia, Kovalevskaitė, Jolanta, Krek, Simon, Krishnamurthy, Parameswari, Kübler, Sandra, Kuyrukçu, Oğuzhan, Kuzgun, Aslı, Kwak, Sookyoung, Laippala, Veronika, Lam, Lucia, Lambertino, Lorenzo, Lando, Tatiana, Larasati, Septina Dian, Lavrentiev, Alexei, Lee, John, Lê Hồng, Phương, Lenci, Alessandro, Lertpradit, Saran, Leung, Herman, Levina, Maria, Li, Cheuk Ying, Li, Josie, Li, Keying, Li, Yuan, Lim, KyungTae, Lima Padovani, Bruna, Lindén, Krister, Ljubešić, Nikola, Loginova, Olga, Lusito, Stefano, Luthfi, Andry, Luukko, Mikko, Lyashevskaya, Olga, Lynn, Teresa, Macketanz, Vivien, Mahamdi, Menel, Maillard, Jean, Makazhanov, Aibek, Mandl, Michael, Manning, Christopher, Manurung, Ruli, Marşan, Büşra, Mărănduc, Cătălina, Mareček, David, Marheinecke, Katrin, Markantonatou, Stella, Martínez Alonso, Héctor, Martín Rodríguez, Lorena, Martins, André, Mašek, Jan, Matsuda, Hiroshi, Matsumoto, Yuji, Mazzei, Alessandro, McDonald, Ryan, McGuinness, Sarah, Mendonça, Gustavo, Merzhevich, Tatiana, Miekka, Niko, Mischenkova, Karina, Misirpashayeva, Margarita, Missilä, Anna, Mititelu, Cătălin, Mitrofan, Maria, Miyao, Yusuke, Mojiri Foroushani, AmirHossein, Molnár, Judit, Moloodi, Amirsaeid, Montemagni, Simonetta, More, Amir, Moreno Romero, Laura, Moretti, Giovanni, Mori, Keiko Sophie, Mori, Shinsuke, Morioka, Tomohiko, Moro, Shigeki, Mortensen, Bjartur, Moskalevskyi, Bohdan, Muischnek, Kadri, Munro, Robert, Murawaki, Yugo, Müürisep, Kaili, Nainwani, Pinkey, Nakhlé, Mariam, Navarro Horñiacek, Juan Ignacio, Nedoluzhko, Anna, Nešpore-Bērzkalne, Gunta, Nevaci, Manuela, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikaido, Yoshihiro, Nikolaev, Vitaly, Nitisaroj, Rattima, Nourian, Alireza, Nurmi, Hanna, Ojala, Stina, Ojha, Atul Kr., Olúòkun, Adédayọ̀, Omura, Mai, Onwuegbuzia, Emeka, Ordan, Noam, Osenova, Petya, Östling, Robert, Øvrelid, Lilja, Özateş, Şaziye Betül, Özçelik, Merve, Özgür, Arzucan, Öztürk Başaran, Balkız, Paccosi, Teresa, Palmero Aprosio, Alessio, Park, Hyunji Hayley, Partanen, Niko, Pascual, Elena, Passarotti, Marco, Patejuk, Agnieszka, Paulino-Passos, Guilherme, Pedonese, Giulia, Peljak-Łapińska, Angelika, Peng, Siyao, Perez, Cenel-Augusto, Perkova, Natalia, Perrier, Guy, Petrov, Slav, Petrova, Daria, Peverelli, Andrea, Phelan, Jason, Piitulainen, Jussi, Pirinen, Tommi A, Pitler, Emily, Plank, Barbara, Poibeau, Thierry, Ponomareva, Larisa, Popel, Martin, Pretkalniņa, Lauma, Prévost, Sophie, Prokopidis, Prokopis, Przepiórkowski, Adam, Puolakainen, Tiina, Pyysalo, Sampo, Qi, Peng, Rääbis, Andriela, Rademaker, Alexandre, Rahoman, Mizanur, Rama, Taraka, Ramasamy, Loganathan, Ramisch, Carlos, Rashel, Fam, Rasooli, Mohammad Sadegh, Ravishankar, Vinit, Real, Livy, Rebeja, Petru, Reddy, Siva, Regnault, Mathilde, Rehm, Georg, Riabov, Ivan, Rießler, Michael, Rimkutė, Erika, Rinaldi, Larissa, Rituma, Laura, Rizqiyah, Putri, Rocha, Luisa, Rögnvaldsson, Eiríkur, Romanenko, Mykhailo, Rosa, Rudolf, Roșca, Valentin, Rovati, Davide, Rozonoyer, Ben, Rudina, Olga, Rueter, Jack, Rúnarsson, Kristján, Sadde, Shoval, Safari, Pegah, Sagot, Benoît, Sahala, Aleksi, Saleh, Shadi, Salomoni, Alessio, Samardžić, Tanja, Samson, Stephanie, Sanguinetti, Manuela, Sanıyar, Ezgi, Särg, Dage, Saulīte, Baiba, Sawanakunanon, Yanin, Saxena, Shefali, Scannell, Kevin, Scarlata, Salvatore, Schneider, Nathan, Schuster, Sebastian, Schwartz, Lane, Seddah, Djamé, Seeker, Wolfgang, Seraji, Mojgan, Shahzadi, Syeda, Shen, Mo, Shimada, Atsuko, Shirasu, Hiroyuki, Shishkina, Yana, Shohibussirri, Muh, Sichinava, Dmitry, Siewert, Janine, Sigurðsson, Einar Freyr, Silveira, Aline, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Skachedubova, Maria, Smith, Aaron, Soares-Bastos, Isabela, Sourov, Shafi, Spadine, Carolyn, Sprugnoli, Rachele, Stamou, Vivian, Steingrímsson, Steinþór, Stella, Antonio, Straka, Milan, Strickland, Emmett, Strnadová, Jana, Suhr, Alane, Sulestio, Yogi Lesmana, Sulubacak, Umut, Suzuki, Shingo, Swanson, Daniel, Szántó, Zsolt, Taguchi, Chihiro, Taji, Dima, Takahashi, Yuta, Tamburini, Fabio, Tan, Mary Ann C., Tanaka, Takaaki, Tanaya, Dipta, Tavoni, Mirko, Tella, Samson, Tellier, Isabelle, Testori, Marinella, Thomas, Guillaume, Tonelli, Sara, Torga, Liisi, Toska, Marsida, Trosterud, Trond, Trukhina, Anna, Tsarfaty, Reut, Türk, Utku, Tyers, Francis, Uematsu, Sumire, Untilov, Roman, Urešová, Zdeňka, Uria, Larraitz, Uszkoreit, Hans, Utka, Andrius, Vagnoni, Elena, Vajjala, Sowmya, van der Goot, Rob, Vanhove, Martine, van Niekerk, Daniel, van Noord, Gertjan, Varga, Viktor, Vedenina, Uliana, Villemonte de la Clergerie, Eric, Vincze, Veronika, Vlasova, Natalia, Wakasa, Aya, Wallenberg, Joel C., Wallin, Lars, Walsh, Abigail, Wang, Jing Xian, Washington, Jonathan North, Wendt, Maximilan, Widmer, Paul, Wigderson, Shira, Wijono, Sri Hartati, Williams, Seyi, Wirén, Mats, Wittern, Christian, Woldemariam, Tsegay, Wong, Tak-sum, Wróblewska, Alina, Yako, Mary, Yamashita, Kayo, Yamazaki, Naoki, Yan, Chunxiao, Yasuoka, Koichi, Yavrumyan, Marat M., Yenice, Arife Betül, Yıldız, Olcay Taner, Yu, Zhuoran, Yuliawati, Arlisa, Žabokrtský, Zdeněk, Zahra, Shorouq, Zeldes, Amir, Zhou, He, Zhu, Hanzhi, Zhuravleva, Anna, and Ziane, Rayan
- Publisher:
- Universal Dependencies Consortium
- Type:
- text and corpus
- Subject:
- treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
- Language:
- Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, Urdu, Russia Buriat, Northern Kurdish, Northern Sami, Upper Sorbian, Afrikaans, Yue Chinese, Marathi, Serbian, Swedish Sign Language, Telugu, Amharic, Armenian, Breton, Faroese, Komi-Zyrian, Nigerian Pidgin, Old French (842-ca. 1400), Tagalog, Thai, Warlpiri, Yoruba, Akkadian, Bambara, Erzya, Maltese, Welsh, Wolof, Assyrian Neo-Aramaic, Literary Chinese, Old Russian, Karelian, Mbyá Guaraní, Bhojpuri, Komi-Permyak, Livvi, Moksha, Scottish Gaelic, Skolt Sami, Swiss German, Albanian, Icelandic, Akuntsu, Apurinã, Chukot, Khunsari, Manx, Mundurukú, Nayini, Old Turkish, Soi, South Levantine Arabic, Tupinambá, Beja, Western Frisian, Guajajára, Urubú-Kaapor, Kangri, K'iche', Low German, Makuráp, Central Siberian Yupik, Western Armenian, Bengali, Javanese, Karo (Brazil), Ligurian, Neapolitan, Tatar, Xibe, Yakut, Ancient Hebrew, Cebuano, Guarani, Hittite, Madi, Emerillon, and Umbrian
- Description:
- Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
- Rights:
- Licence Universal Dependencies v2.10, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.10, and PUB
10. Universal Dependencies 2.10 models for UDPipe 2 (2022-07-11)
- Creator:
- Straka, Milan
- Publisher:
- Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
- Type:
- tool and toolService
- Subject:
- tokenizer, POS tagger, lemmatization, tagger, parser, and dependency parser
- Language:
- Afrikaans, Arabic, Belarusian, Bulgarian, Catalan, Czech, Church Slavic, Coptic, Welsh, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Faroese, Persian, Finnish, French, Old French (842-ca. 1400), Scottish Gaelic, Irish, Galician, Gothic, Ancient Greek (to 1453), Ancient Hebrew, Hebrew, Hindi, Croatian, Hungarian, Armenian, Western Armenian, Indonesian, Icelandic, Italian, Japanese, Korean, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Maltese, Dutch, Norwegian Nynorsk, Norwegian Bokmål, Old Russian, Nigerian Pidgin, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Telugu, Turkish, Uighur, Ukrainian, Urdu, Vietnamese, Gambian Wolof, Wolof, and Chinese
- Description:
- Tokenizer, POS Tagger, Lemmatizer and Parser models for 123 treebanks of 69 languages of Universal Depenencies 2.10 Treebanks, created solely using UD 2.10 data (https://hdl.handle.net/11234/1-4758). The model documentation including performance can be found at https://ufal.mff.cuni.cz/udpipe/2/models#universal_dependencies_210_models . To use these models, you need UDPipe version 2.0, which you can download from https://ufal.mff.cuni.cz/udpipe/2 .
- Rights:
- Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB
- « Previous
- Next »
- 1
- 2
- 3