Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Abdulmumin, Idris , Das, Satya Ranja , Dawud, Musa Abdullahi , Parida, Shantipriya , Muhammad, Shamsuddeen Hassan , Ahmad, Ibrahim Sa'id , Panda, Subhadarshi , Bojar, Ondřej , Galadanci, Bashir Shehu , and Bello, Bello Shehu
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
image and corpus
Subject:
multi-modal , machine translation , image captioning , image annotation , and neural machine translation
Language:
Hausa and English
Description:
Data
-------
Hausa Visual Genome 1.0, a multimodal dataset consisting of text and images suitable for English-to-Hausa multimodal machine translation tasks and multimodal research. We follow the same selection of short English segments (captions) and the associated images from Visual Genome as the dataset Hindi Visual Genome 1.1 has. We automatically translated the English captions to Hausa and manually post-edited, taking the associated images into account.
The training set contains 29K segments. Further 1K and 1.6K segments are provided in development and test sets, respectively, which follow the same (random) sampling from the original Hindi Visual Genome.
Additionally, a challenge test set of 1400 segments is available for the multi-modal task. This challenge test set was created in Hindi Visual Genome by searching for (particularly) ambiguous English words based on the embedding similarity and manually selecting those where the image helps to resolve the ambiguity.
Dataset Formats
-----------------------
The multimodal dataset contains both text and images.
The text parts of the dataset (train and test sets) are in simple tab-delimited plain text files.
All the text files have seven columns as follows:
Column1 - image_id
Column2 - X
Column3 - Y
Column4 - Width
Column5 - Height
Column6 - English Text
Column7 - Hausa Text
The image part contains the full images with the corresponding image_id as the file name. The X, Y, Width, and Height columns indicate the rectangular region in the image described by the caption.
Data Statistics
--------------------
The statistics of the current release are given below.
Parallel Corpus Statistics
-----------------------------------
Dataset Segments English Words Hausa Words
---------- -------- ------------- -----------
Train 28930 143106 140981
Dev 998 4922 4857
Test 1595 7853 7736
Challenge Test 1400 8186 8752
---------- -------- ------------- -----------
Total 32923 164067 162326
The word counts are approximate, prior to tokenization.
Citation
-----------
If you use this corpus, please cite the following paper:
@InProceedings{abdulmumin-EtAl:2022:LREC,
author = {Abdulmumin, Idris
and Dash, Satya Ranjan
and Dawud, Musa Abdullahi
and Parida, Shantipriya
and Muhammad, Shamsuddeen
and Ahmad, Ibrahim Sa'id
and Panda, Subhadarshi
and Bojar, Ond{\v{r}}ej
and Galadanci, Bashir Shehu
and Bello, Bello Shehu},
title = "{Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation}",
booktitle = {Proceedings of the Language Resources and Evaluation Conference},
month = {June},
year = {2022},
address = {Marseille, France},
publisher = {European Language Resources Association},
pages = {6471--6479},
url = {https://aclanthology.org/2022.lrec-1.694}
}
Rights:
Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) , http://creativecommons.org/licenses/by-nc-sa/4.0/ , and PUB
Creator:
Rosa, Rudolf
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
Wikipedia , text corpora , and monolingual corpus
Language:
Abkhazian , Achinese , Adyghe , Afrikaans , Akan , Tosk Albanian , Amharic , Old English (ca. 450-1100) , Arabic , Official Aramaic (700-300 BCE) , Aragonese , Egyptian Arabic , Assamese , Asturian , Atikamekw , Avaric , Aymara , South Azerbaijani , Azerbaijani , Bashkir , Bambara , Bavarian , Central Bikol , Belarusian , Bengali , Bislama , Banjar , Tibetan , Bosnian , Bishnupriya , Breton , Buginese , Bulgarian , Russia Buriat , Catalan , Min Dong Chinese , Cebuano , Czech , Chamorro , Chechen , Cherokee , Church Slavic , Chuvash , Cheyenne , Central Kurdish , Cornish , Corsican , Cree , Crimean Tatar , Kashubian , Welsh , Danish , German , Dinka , Dimli (individual language) , Dhivehi , Lower Sorbian , Dzongkha , Modern Greek (1453-) , English , Esperanto , Estonian , Basque , Ewe , Extremaduran , Faroese , Persian , Fijian , Finnish , French , Arpitan , Northern Frisian , Western Frisian , Fulah , Friulian , Gagauz , Gan Chinese , Scottish Gaelic , Irish , Galician , Gilaki , Manx , Goan Konkani , Gothic , Guarani , Gujarati , Hakka Chinese , Haitian , Hausa , Hawaiian , Serbo-Croatian , Hebrew , Herero , Fiji Hindi , Hindi , Hiri Motu , Croatian , Upper Sorbian , Hungarian , Armenian , Igbo , Ido , Inuktitut , Interlingue , Iloko , Interlingua (International Auxiliary Language Association) , Indonesian , Inupiaq , Icelandic , Italian , Jamaican Creole English , Javanese , Lojban , Japanese , Kara-Kalpak , Kabyle , Kalaallisut , Kannada , Kashmiri , Georgian , Kanuri , Kazakh , Kabardian , Kabiyè , Khmer , Kikuyu , Kinyarwanda , Kirghiz , Komi-Permyak , Komi , Kongo , Korean , Karachay-Balkar , Kölsch , Kurdish , Ladino , Lao , Latin , Latvian , Lak , Lezghian , Ligurian , Limburgan , Lingala , Lithuanian , Lombard , Northern Luri , Latgalian , Luxembourgish , Ganda , Literary Chinese , Marshallese , Maithili , Malayalam , Marathi , Moksha , Eastern Mari , Minangkabau , Macedonian , Malagasy , Maltese , Mongolian , Maori , Western Mari , Malay (macrolanguage) , Creek , Mirandese , Burmese , Erzya , Mazanderani , Min Nan Chinese , Neapolitan , Nauru , Navajo , Ndonga , Low German , Nepali (macrolanguage) , Newari , Dutch , Norwegian Nynorsk , Norwegian , Novial , Pedi , Nyanja , Occitan (post 1500) , Livvi , Oriya (macrolanguage) , Oromo , Ossetian , Pangasinan , Pampanga , Panjabi , Papiamento , Picard , Pennsylvania German , Pfaelzisch , Pitcairn-Norfolk , Pali , Piemontese , Western Panjabi , Pontic , Polish , Portuguese , Pushto , Quechua , Vlax Romani , Romansh , Romanian , Rusyn , Rundi , Macedo-Romanian , Russian , Sango , Yakut , Sanskrit , Sicilian , Scots , Samogitian , Sinhala , Slovak , Slovenian , Northern Sami , Samoan , Shona , Sindhi , Somali , Southern Sotho , Spanish , Albanian , Sardinian , Sranan Tongo , Serbian , Swati , Saterfriesisch , Sundanese , Swahili (macrolanguage) , Swedish , Silesian , Tahitian , Tamil , Tatar , Tulu , Telugu , Tama (Colombia) , Tetum , Tajik , Tagalog , Thai , Tigrinya , Tonga (Tonga Islands) , Tok Pisin , Tswana , Tsonga , Turkmen , Tumbuka , Turkish , Twi , Tuvinian , Udmurt , Uighur , Ukrainian , Urdu , Uzbek , Venetian , Venda , Veps , Vietnamese , Vlaams , Volapük , Võro , Waray (Philippines) , Walloon , Wolof , Wu Chinese , Kalmyk , Xhosa , Mingrelian , Yiddish , Yoruba , Yue Chinese , Zeeuws , Zhuang , Chinese , Zulu , and Dotyali
Description:
Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018.
The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages).
For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias].
The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast).
Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day.
The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
Rights:
Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) , http://creativecommons.org/licenses/by-sa/3.0/ , and PUB
Creator:
Zeman, Daniel , Nivre, Joakim , Abrams, Mitchell , Ackermann, Elia , Aepli, Noëmi , Aghaei, Hamid , Agić, Željko , Ahmadi, Amir , Ahrenberg, Lars , Ajede, Chika Kennedy , Akkurt, Salih Furkan , Aleksandravičiūtė, Gabrielė , Alfina, Ika , Algom, Avner , Alnajjar, Khalid , Alzetta, Chiara , Andersen, Erik , Antonsen, Lene , Aoyama, Tatsuya , Aplonova, Katya , Aquino, Angelina , Aragon, Carolina , Aranes, Glyd , Aranzabe, Maria Jesus , Arıcan, Bilge Nas , Arnardóttir, Þórunn , Arutie, Gashaw , Arwidarasti, Jessica Naraiswari , Asahara, Masayuki , Ásgeirsdóttir, Katla , Aslan, Deniz Baran , Asmazoğlu, Cengiz , Ateyah, Luma , Atmaca, Furkan , Attia, Mohammed , Atutxa, Aitziber , Augustinus, Liesbeth , Avelãs, Mariana , Badmaeva, Elena , Balasubramani, Keerthana , Ballesteros, Miguel , Banerjee, Esha , Bank, Sebastian , Barbu Mititelu, Verginica , Barkarson, Starkaður , Basile, Rodolfo , Basmov, Victoria , Batchelor, Colin , Bauer, John , Bedir, Seyyit Talha , Behzad, Shabnam , Belieni, Juan , Bengoetxea, Kepa , Benli, İbrahim , Ben Moshe, Yifat , Berg, Ansu , Berk, Gözde , Bhat, Riyaz Ahmad , Biagetti, Erica , Bick, Eckhard , Bielinskienė, Agnė , Bilgin Taşdemir, Esma Fatıma , Bjarnadóttir, Kristín , Blaschke, Verena , Blokland, Rogier , Bobicev, Victoria , Boizou, Loïc , Bonilla, Johnatan , Borges Völker, Emanuel , Börstell, Carl , Bosco, Cristina , Bouma, Gosse , Bowman, Sam , Boyd, Adriane , Braggaar, Anouck , Branco, António , Brokaitė, Kristina , Burchardt, Aljoscha , Campos, Marisa , Candito, Marie , Caron, Bernard , Caron, Gauthier , Carvalheiro, Catarina , Carvalho, Rita , Cassidy, Lauren , Castro, Maria Clara , Castro, Sérgio , Cavalcanti, Tatiana , Cebiroğlu Eryiğit, Gülşen , Cecchini, Flavio Massimiliano , Celano, Giuseppe G. A. , Čéplö, Slavomír , Cesur, Neslihan , Cetin, Savas , Çetinoğlu, Özlem , Chalub, Fabricio , Chamila, Liyanage , Chauhan, Shweta , Chen, Yifei , Chi, Ethan , Chika, Taishi , Cho, Yongseok , Choi, Jinho , Chontaeva, Bermet , Chun, Jayeol , Chung, Juyeon , Cignarella, Alessandra T. , Cinková, Silvie , Collomb, Aurélie , Çöltekin, Çağrı , Connor, Miriam , Corbetta, Claudia , Corbetta, Daniela , Costa, Francisco , Courtin, Marine , Crabbé, Benoît , Cristescu, Mihaela , Cvetkoski, Vladimir , Dale, Ingerid Løyning , Daniel, Philemon , Davidson, Elizabeth , de Alencar, Leonel Figueiredo , Dehouck, Mathieu , de Laurentiis, Martina , de Marneffe, Marie-Catherine , de Paiva, Valeria , Derin, Mehmet Oguz , de Souza, Elvis , Diaz de Ilarraza, Arantza , Díaz Hernández, Roberto Antonio , Dickerson, Carly , Dinakaramani, Arawinda , Di Nuovo, Elisa , Dione, Bamba , Dirix, Peter , Do, Hoa , Dobrovoljc, Kaja , Döhmer, Caroline , Doyle, Adrian , Dozat, Timothy , Droganova, Kira , Duran, Magali Sanches , Dwivedi, Puneet , Ebert, Christian , Eckhoff, Hanne , Eguchi, Masaki , Eiche, Sandra , Eiselen, Roald , Eli, Marhaba , Elkahky, Ali , Ephrem, Binyam , Erina, Olga , Erjavec, Tomaž , Eslami, Soudabeh , Essaidi, Farah , Etienne, Aline , Evelyn, Wograine , Facundes, Sidney , Farkas, Richárd , Favero, Federica , Ferdaousi, Jannatul , Fernanda, Marília , Fernandez Alcalde, Hector , Fethi, Amal , Foster, Jennifer , Fransen, Theodorus , Freitas, Cláudia , Fujita, Kazunori , Gajdošová, Katarína , Galbraith, Daniel , Galy, Edith , Gamba, Federica , Garcia, Marcos , Gärdenfors, Moa , Gaustad, Tanja , Genç, Efe Eren , Gerardi, Fabrício Ferraz , Gerdes, Kim , Gessler, Luke , Ginter, Filip , Godoy, Gustavo , Goenaga, Iakes , Gojenola, Koldo , Gökırmak, Memduh , Goldberg, Yoav , Gómez Guinovart, Xavier , González Saavedra, Berta , Griciūtė, Bernadeta , Grioni, Matias , Grobol, Loïc , Grūzītis, Normunds , Guillaume, Bruno , Guiller, Kirian , Guillot-Barbance, Céline , Güngör, Tunga , Habash, Nizar , Hafsteinsson, Hinrik , Hajič, Jan , Hajič jr., Jan , Hämäläinen, Mika , Hà Mỹ, Linh , Han, Na-Rae , Hanifmuti, Muhammad Yudistira , Harada, Takahiro , Hardwick, Sam , Harris, Kim , Hassert, Naïma , Haug, Dag , Heinecke, Johannes , Hellwig, Oliver , Hennig, Felix , Hladká, Barbora , Hlaváčová, Jaroslava , Hociung, Florinel , Hoefels, Diana , Hohle, Petter , Huang, Yidi , Huerta Mendez, Marivel , Hwang, Jena , Ikeda, Takumi , Iliadou, Inessa , Ingason, Anton Karl , Ion, Radu , Irimia, Elena , Ishola, Ọlájídé , Islamaj, Artan , Ito, Kaoru , Iurescia, Federica , Jagodzińska, Sandra , Jannat, Siratun , Jelínek, Tomáš , Jha, Apoorva , Jiang, Katharine , Jobanputra, Mayank , Johannsen, Anders , Jónsdóttir, Hildur , Jørgensen, Fredrik , Juutinen, Markus , Kaşıkara, Hüner , Kabaeva, Nadezhda , Kahane, Sylvain , Kanayama, Hiroshi , Kanerva, Jenna , Kara, Neslihan , Karahóǧa, Ritván , Kåsen, Andre , Kayadelen, Tolga , Kengatharaiyer, Sarveswaran , Kettnerová, Václava , Kharatyan, Lilit , Kirchner, Jesse , Klementieva, Elena , Klyachko, Elena , Kocharov, Petr , Köhn, Arne , Köksal, Abdullatif , Kopacewicz, Kamil , Korkiakangas, Timo , Köse, Mehmet , Koshevoy, Alexey , Kotsyba, Natalia , Kovačić, Barbara , Kovalevskaitė, Jolanta , Krek, Simon , Krishnamurthy, Parameswari , Kübler, Sandra , Kuqi, Adrian , Kuyrukçu, Oğuzhan , Kuzgun, Aslı , Kwak, Sookyoung , Kyle, Kris , Laan, Käbi , Laippala, Veronika , Lambertino, Lorenzo , Lando, Tatiana , Larasati, Septina Dian , Lavrentiev, Alexei , Lee, John , Lê Hồng, Phương , Lenci, Alessandro , Lertpradit, Saran , Leung, Herman , Levina, Maria , Levine, Lauren , Li, Cheuk Ying , Li, Josie , Li, Keying , Li, Yixuan , Li, Yuan , Lim, KyungTae , Lima Padovani, Bruna , Lin, Yi-Ju Jessica , Lindén, Krister , Liu, Yang Janet , Ljubešić, Nikola , Lobzhanidze, Irina , Loginova, Olga , Lopes, Lucelene , Lusito, Stefano , Lutgen, Anne-Marie , Luthfi, Andry , Luukko, Mikko , Lyashevskaya, Olga , Lynn, Teresa , Macketanz, Vivien , Mahamdi, Menel , Maillard, Jean , Makarchuk, Ilya , Makazhanov, Aibek , Mambrini, Francesco , Mandl, Michael , Manning, Christopher , Manurung, Ruli , Marşan, Büşra , Mărănduc, Cătălina , Mareček, David , Marheinecke, Katrin , Markantonatou, Stella , Martínez Alonso, Héctor , Martín Rodríguez, Lorena , Martins, André , Martins, Cláudia , Mašek, Jan , Matsuda, Hiroshi , Matsumoto, Yuji , Mazzei, Alessandro , McDonald, Ryan , McGuinness, Sarah , Mehta, Maitrey , Ménard, Pierre André , Mendonça, Gustavo , Merzhevich, Tatiana , Meurer, Paul , Miekka, Niko , Milano, Emilia , Miller, Aaron , Mischenkova, Karina , Missilä, Anna , Mititelu, Cătălin , Mitrofan, Maria , Miyao, Yusuke , Mojiri Foroushani, AmirHossein , Molnár, Judit , Moloodi, Amirsaeid , Montemagni, Simonetta , More, Amir , Moreno Romero, Laura , Moretti, Giovanni , Mori, Shinsuke , Morioka, Tomohiko , Moro, Shigeki , Mortensen, Bjartur , Moskalevskyi, Bohdan , Muischnek, Kadri , Munro, Robert , Murawaki, Yugo , Müürisep, Kaili , Nainwani, Pinkey , Nakhlé, Mariam , Navarro Horñiacek, Juan Ignacio , Nedoluzhko, Anna , Nešpore-Bērzkalne, Gunta , Nevaci, Manuela , Nguyễn Thị, Lương , Nguyễn Thị Minh, Huyền , Nikaido, Yoshihiro , Nikolaev, Vitaly , Nitisaroj, Rattima , Norrman, Victor , Nourian, Alireza , Nunes, Maria das Graças Volpe , Nurmi, Hanna , Ojala, Stina , Ojha, Atul Kr. , Óladóttir, Hulda , Olúòkun, Adédayọ̀ , Omura, Mai , Onwuegbuzia, Emeka , Ordan, Noam , Osenova, Petya , Östling, Robert , Ott, Annika , Øvrelid, Lilja , Özateş, Şaziye Betül , Özçelik, Merve , Özgür, Arzucan , Öztürk Başaran, Balkız , Paccosi, Teresa , Palmero Aprosio, Alessio , Panova, Anastasia , Pardo, Thiago Alexandre Salgueiro , Park, Hyunji Hayley , Partanen, Niko , Pascual, Elena , Passarotti, Marco , Patejuk, Agnieszka , Paulino-Passos, Guilherme , Pedonese, Giulia , Peljak-Łapińska, Angelika , Peng, Siyao , Peng, Siyao Logan , Pereira, Rita , Pereira, Sílvia , Perez, Cenel-Augusto , Perkova, Natalia , Perrier, Guy , Petrov, Slav , Petrova, Daria , Peverelli, Andrea , Phelan, Jason , Pierre-Louis, Claudel , Piitulainen, Jussi , Pinter, Yuval , Pinto, Clara , Pintucci, Rodrigo , Pirinen, Tommi A , Pitler, Emily , Plamada, Magdalena , Plank, Barbara , Plum, Alistair , Poibeau, Thierry , Ponomareva, Larisa , Popel, Martin , Pretkalniņa, Lauma , Pretorius, Rigardt , Prévost, Sophie , Prokopidis, Prokopis , Przepiórkowski, Adam , Pugh, Robert , Puolakainen, Tiina , Purschke, Christoph , Pyysalo, Sampo , Qi, Peng , Querido, Andreia , Rääbis, Andriela , Rademaker, Alexandre , Rahoman, Mizanur , Rama, Taraka , Ramasamy, Loganathan , Ramisch, Carlos , Ramos, Joana , Rashel, Fam , Rasooli, Mohammad Sadegh , Ravishankar, Vinit , Real, Livy , Rebeja, Petru , Reddy, Siva , Regnault, Mathilde , Rehm, Georg , Riabi, Arij , Riabov, Ivan , Rießler, Michael , Rimkutė, Erika , Rinaldi, Larissa , Rituma, Laura , Rizqiyah, Putri , Rocha, Luisa , Rögnvaldsson, Eiríkur , Roksandic, Ivan , Romanenko, Mykhailo , Rosa, Rudolf , Roșca, Valentin , Rovati, Davide , Rozonoyer, Ben , Rudina, Olga , Rueter, Jack , Ruffolo, Paolo , Rúnarsson, Kristján , Sadde, Shoval , Safari, Pegah , Sahala, Aleksi , Saleh, Shadi , Salomoni, Alessio , Samardžić, Tanja , Samson, Stephanie , Sánchez-Rodríguez, Xulia , Sanguinetti, Manuela , Sanıyar, Ezgi , Särg, Dage , Sartor, Marta , Sarymsakova, Albina , Sasaki, Mitsuya , Saulīte, Baiba , Savary, Agata , Sawanakunanon, Yanin , Saxena, Shefali , Scannell, Kevin , Scarlata, Salvatore , Schang, Emmanuel , Schneider, Nathan , Schuster, Sebastian , Schwartz, Lane , Seddah, Djamé , Seeker, Wolfgang , Sellmer, Sven , Seraji, Mojgan , Shahzadi, Syeda , Shen, Mo , Shimada, Atsuko , Shirasu, Hiroyuki , Shishkina, Yana , Shohibussirri, Muh , Shvedova, Maria , Siewert, Janine , Sigurðsson, Einar Freyr , Silva, João , Silveira, Aline , Silveira, Natalia , Silveira, Sara , Simi, Maria , Simionescu, Radu , Simkó, Katalin , Šimková, Mária , Símonarson, Haukur Barri , Simov, Kiril , Sitchinava, Dmitri , Sither, Ted , Smith, Aaron , Soares-Bastos, Isabela , Solberg, Per Erik , Sonnenhauser, Barbara , Sourov, Shafi , Sprugnoli, Rachele , Stamou, Vivian , Steingrímsson, Steinþór , Stella, Antonio , Stephen, Abishek , Straka, Milan , Strickland, Emmett , Strnadová, Jana , Suhr, Alane , Sulestio, Yogi Lesmana , Sulubacak, Umut , Suzuki, Shingo , Swanson, Daniel , Szántó, Zsolt , Taguchi, Chihiro , Taji, Dima , Tamburini, Fabio , Tan, Mary Ann C. , Tanaka, Takaaki , Tanaya, Dipta , Tavoni, Mirko , Tella, Samson , Tellier, Isabelle , Testori, Marinella , Thomas, Guillaume , Tıraş, Tarık Emre , Tonelli, Sara , Torga, Liisi , Toska, Marsida , Trosterud, Trond , Trukhina, Anna , Tsarfaty, Reut , Türk, Utku , Tyers, Francis , Þórðarson, Sveinbjörn , Þorsteinsson, Vilhjálmur , Uematsu, Sumire , Untilov, Roman , Urešová, Zdeňka , Uria, Larraitz , Uszkoreit, Hans , Utka, Andrius , Vagnoni, Elena , Vajjala, Sowmya , Vak, Socrates , van der Goot, Rob , Vanhove, Martine , van Niekerk, Daniel , van Noord, Gertjan , Varga, Viktor , Vedenina, Uliana , Venturi, Giulia , Villemonte de la Clergerie, Eric , Vincze, Veronika , Vissamsetty, Anishka , Vlasova, Natalia , Vligouridou, Eleni , Wakasa, Aya , Wallenberg, Joel C. , Wallin, Lars , Walsh, Abigail , Wang, John , Washington, Jonathan North , Wendt, Maximilan , Widmer, Paul , Wigderson, Shira , Wijono, Sri Hartati , Wille, Vanessa Berwanger , Williams, Seyi , Wirén, Mats , Wittern, Christian , Woldemariam, Tsegay , Wong, Tak-sum , Wróblewska, Alina , Wu, Qishen , Yako, Mary , Yamashita, Kayo , Yamazaki, Naoki , Yan, Chunxiao , Yasuoka, Koichi , Yavrumyan, Marat M. , Yenice, Arife Betül , Yılandiloğlu, Enes , Yıldız, Olcay Taner , Yu, Zhuoran , Yuliawati, Arlisa , Žabokrtský, Zdeněk , Zahra, Shorouq , Zeldes, Amir , Zhou, He , Zhu, Hanzhi , Zhu, Yilun , Zhuravleva, Anna , and Ziane, Rayan
Publisher:
Universal Dependencies Consortium
Type:
text and corpus
Subject:
treebank , dependency , syntax , morphology , harmonized annotation , interset , universal tagset , and stanford dependencies
Language:
Ancient Greek (to 1453) , Arabic , Basque , Bulgarian , Croatian , Czech , Danish , Dutch , English , Estonian , Finnish , French , German , Gothic , Modern Greek (1453-) , Hebrew , Hindi , Hungarian , Indonesian , Irish , Italian , Japanese , Latin , Norwegian , Church Slavic , Persian , Polish , Portuguese , Romanian , Slovenian , Spanish , Swedish , Tamil , Catalan , Chinese , Galician , Kazakh , Latvian , Russian , Turkish , Coptic , Sanskrit , Slovak , Ukrainian , Uighur , Vietnamese , Belarusian , Korean , Lithuanian , Urdu , Russia Buriat , Northern Kurdish , Northern Sami , Upper Sorbian , Afrikaans , Yue Chinese , Marathi , Serbian , Swedish Sign Language , Telugu , Amharic , Armenian , Breton , Faroese , Komi-Zyrian , Nigerian Pidgin , Old French (842-ca. 1400) , Tagalog , Thai , Warlpiri , Yoruba , Akkadian , Bambara , Erzya , Maltese , Welsh , Wolof , Assyrian Neo-Aramaic , Literary Chinese , Old Russian , Karelian , Mbyá Guaraní , Bhojpuri , Komi-Permyak , Livvi , Moksha , Scottish Gaelic , Skolt Sami , Swiss German , Albanian , Icelandic , Akuntsu , Apurinã , Chukot , Khunsari , Manx , Mundurukú , Nayini , Old Turkish , Soi , South Levantine Arabic , Tupinambá , Beja , Western Frisian , Guajajára , Urubú-Kaapor , Kangri , K'iche' , Low German , Makuráp , Central Siberian Yupik , Western Armenian , Bengali , Javanese , Karo (Brazil) , Ligurian , Neapolitan , Tatar , Xibe , Yakut , Ancient Hebrew , Cebuano , Guarani , Hittite , Madi , Emerillon , Umbrian , Abaza , Gheg Albanian , Malayalam , Nhengatu , Sinhala , Zacatlán-Ahuacatlán-Tepetzintla Nahuatl , Xavánte , Saya , Borôro , Kirghiz , Algerian Arabic , Old Irish (to 900) , Classical Armenian , Georgian , Haitian , Highland Puebla Nahuatl , Macedonian , Middle French (ca. 1400-1600) , Veps , Abkhazian , Azerbaijani , Bavarian , Cappadocian Greek , Egyptian (Ancient) , Gujarati , Hausa , Latgalian , Luxembourgish , Ottoman Turkish (1500-1928) , Paumarí , and Tswana
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:
Licence Universal Dependencies v2.14 , https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.14 , and PUB