Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Zeman, Daniel , Bouma, Gosse , and Seddah, Djamé
Publisher:
Universal Dependencies Consortium
Type:
text and corpus
Subject:
treebank , dependency , syntax , enhanced universal dependencies , shared task , and parsing
Language:
Arabic , Bulgarian , Czech , Dutch , English , Estonian , Finnish , French , Italian , Latvian , Lithuanian , Polish , Russian , Slovak , Swedish , Tamil , and Ukrainian
Description:
This package contains data used in the IWPT 2021 shared task. It contains training, development and test (evaluation) datasets. The data is based on a subset of Universal Dependencies release 2.7 (http://hdl.handle.net/11234/1-3424) but some treebanks contain additional enhanced annotations. Moreover, not all of these additions became part of Universal Dependencies release 2.8 (http://hdl.handle.net/11234/1-3687), which makes the shared task data unique and worth a separate release to enable later comparison with new parsing algorithms. The package also contains a number of Perl and Python scripts that have been used to process the data during preparation and during the shared task. Finally, the package includes the official primary submission of each team participating in the shared task.
Rights:
Licence Universal Dependencies v2.7 , https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.7 , and PUB
Creator:
Jan Patočka
Publisher:
Str. 160–203. Stať.
Type:
Text
Subject:
1975 , 1979/25 , 1980/27 , 1981/6 , 1981/7 , 1988/28 , 1988/31 , 1988/32 , 1988/34 , 1994/7 , 1996/4 , 1996/7 , 1997/7 , 1998/3 , 1999/8 , 2001/9 , 2002/21 , 2006/1 , 2007/1 , 2008/3 , be , bg , cs , de , en , es , fr , fulltext , hu , it , lt , no , pl , ru , sl , sr , SS-3/PD-III , sv , and uk
Language:
Czech , English , Bulgarian , French , Italian , Lithuanian , Hungarian , German , Norwegian , Polish , Russian , Belarusian , Slovenian , Serbian , Spanish , Swedish , and Ukrainian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Masaryk, Tomáš Garrigue,
Type:
text , korespondence , and edice
Subject:
Politika , Biografie , Masaryk, Tomáš Garrigue, , politici čeští , prezidenti českoslovenští , vztahy mezinárodní , Slované , Rusové , and Ukrajinci
Language:
Czech , English , French , German , Polish , Russian , and Ukrainian
Description:
Název v tiráži: Korespondence TGM :, Nad názvem: TGM - MZ, Obsahuje doplňky k 1. svazku, seznam korespondence, anotovaných dokumentů a jmenný rejstřík, Nad názvem: TGM - AVA, and T. G. Masaryk anf The Slavs
Rights:
unknown
Creator:
Masaryk, Tomáš Garrigue,
Type:
text , korespondence , and edice
Subject:
Politika , Biografie , Masaryk, Tomáš Garrigue, , politici čeští , prezidenti českoslovenští , vztahy mezinárodní , Slované , Poláci , Rusové , and Ukrajinci
Language:
Czech , English , French , German , Polish , Russian , and Ukrainian
Description:
Název v tiráži: Korespondence TGM :, Nad názvem: TGM - MZ, and T. G. Masaryk anf The Slavs
Rights:
unknown
Creator:
Jan Patočka
Publisher:
Str. 89–131. Stať. [Součástí eseje i text To platí též..., v. 1988/25H.]
Type:
Text
Subject:
1975 , 1979/25 , 1981/6 , 1981/7 , 1988/25H , 1988/28 , 1988/31 , 1988/32 , 1988/34 , 1994/7 , 1996/4 , 1996/7 , 1998/3 , 1999/8 , 2 , 2001/9 , 2002/21 , 2002/7 , 2006/1 , 2007/1 , 2008/3 , bg , cs , de , en , es , fr , fulltext , hu , it , jp , lt , no , pl , ru , SS-3/PD-III , sv , uk , and v
Language:
Czech , English , Bulgarian , French , Italian , Lithuanian , Hungarian , German , Norwegian , Polish , Russian , Spanish , Swedish , and Ukrainian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Bojar, Ondřej , Cífka, Ondřej , Pecina, Pavel , and Tamchyna, Aleš
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
toolService and tool
Subject:
machine translation , web service , and demo
Language:
Czech , English , Russian , Ukrainian , French , and German
Description:
An interactive web demo of selected ÚFAL MT systems. and FP7-ICT-2011-7-288487 MosesCore
Rights:
Not specified
Creator:
Jan Patočka
Publisher:
Str. 9–44. Stať. [Psáno r. 1951–1953, původně snad zamýšleno autorem jako jeho příspěvek do oslavného sborníku k 70. narozeninám F. Novotného (1951), protože se práce rozrostla a autor text dokončil včas, kolovala jako samostatná strojopisná kopie.] — 2. otisk in: Proměny 24 (New York 1987), č. 1, str. 108–135. — 3. otisk in: Negativní platonismus, 1. knižní vyd., Praha 1990, str. 9–58 (v. 1990/2). — 4. otisk (2. knižní, opr. vyd.) in: Péče o duši I (SS-1/PD-I), Praha 1996, str. 303–336 (v. 1996/2). — 5. otisk: Negativní platonismus, 3. knižní, opr. vyd., Praha 2007, 71 s. (v. 2007/9). — Částečný otisk úryvku ze začátku V. kapitoly (v. 4. otisk, str. 327–330) pod názvem IDEA a CHÓRISMOS, in: Idea, hypotéza a otázka, ed. P. Rezek, Praha (OIKOYMENH) 1991, str. 51–53, Edice PomFil, sv. 1.
Type:
Text
Subject:
1987 , 1988/28 , 1989/16 , 1990/2 , 1990/6 , 1996/2 , 1996/7 , 1996/8 , 2001/9 , 2007/7 , 2007/9 , ca , cs , de , en , es , fr , fulltext , hu , ru , SS-1/PD-I , stať , and uk
Language:
English , French , Catalan , Hungarian , German , Spanish , Russian , Ukrainian , and Czech
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Publisher:
Universität Bamberg, World Language Documentation Centre
Format:
application/octet-stream
Type:
lexicalConceptualResource
Language:
Afrikaans , Arabic , Basque , Bulgarian , Catalan , Chinese , Czech , Danish , Dutch , English , Esperanto , Estonian , Finnish , French , Galician , Georgian , Modern Greek (1453-) , Hebrew , Hungarian , Icelandic , Indonesian , Interlingua (International Auxiliary Language Association) , Irish , Italian , Japanese , Khmer , Norwegian , Polish , Portuguese , Romanian , Russian , Serbian , Slovak , Spanish , Swedish , Turkish , Ukrainian , and Welsh
Rights:
GFDL or CC and http://www.omegawiki.org/Licensing
Creator:
Rosa, Rudolf
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
Wikipedia , text corpora , and monolingual corpus
Language:
Abkhazian , Achinese , Adyghe , Afrikaans , Akan , Tosk Albanian , Amharic , Old English (ca. 450-1100) , Arabic , Official Aramaic (700-300 BCE) , Aragonese , Egyptian Arabic , Assamese , Asturian , Atikamekw , Avaric , Aymara , South Azerbaijani , Azerbaijani , Bashkir , Bambara , Bavarian , Central Bikol , Belarusian , Bengali , Bislama , Banjar , Tibetan , Bosnian , Bishnupriya , Breton , Buginese , Bulgarian , Russia Buriat , Catalan , Min Dong Chinese , Cebuano , Czech , Chamorro , Chechen , Cherokee , Church Slavic , Chuvash , Cheyenne , Central Kurdish , Cornish , Corsican , Cree , Crimean Tatar , Kashubian , Welsh , Danish , German , Dinka , Dimli (individual language) , Dhivehi , Lower Sorbian , Dzongkha , Modern Greek (1453-) , English , Esperanto , Estonian , Basque , Ewe , Extremaduran , Faroese , Persian , Fijian , Finnish , French , Arpitan , Northern Frisian , Western Frisian , Fulah , Friulian , Gagauz , Gan Chinese , Scottish Gaelic , Irish , Galician , Gilaki , Manx , Goan Konkani , Gothic , Guarani , Gujarati , Hakka Chinese , Haitian , Hausa , Hawaiian , Serbo-Croatian , Hebrew , Herero , Fiji Hindi , Hindi , Hiri Motu , Croatian , Upper Sorbian , Hungarian , Armenian , Igbo , Ido , Inuktitut , Interlingue , Iloko , Interlingua (International Auxiliary Language Association) , Indonesian , Inupiaq , Icelandic , Italian , Jamaican Creole English , Javanese , Lojban , Japanese , Kara-Kalpak , Kabyle , Kalaallisut , Kannada , Kashmiri , Georgian , Kanuri , Kazakh , Kabardian , Kabiyè , Khmer , Kikuyu , Kinyarwanda , Kirghiz , Komi-Permyak , Komi , Kongo , Korean , Karachay-Balkar , Kölsch , Kurdish , Ladino , Lao , Latin , Latvian , Lak , Lezghian , Ligurian , Limburgan , Lingala , Lithuanian , Lombard , Northern Luri , Latgalian , Luxembourgish , Ganda , Literary Chinese , Marshallese , Maithili , Malayalam , Marathi , Moksha , Eastern Mari , Minangkabau , Macedonian , Malagasy , Maltese , Mongolian , Maori , Western Mari , Malay (macrolanguage) , Creek , Mirandese , Burmese , Erzya , Mazanderani , Min Nan Chinese , Neapolitan , Nauru , Navajo , Ndonga , Low German , Nepali (macrolanguage) , Newari , Dutch , Norwegian Nynorsk , Norwegian , Novial , Pedi , Nyanja , Occitan (post 1500) , Livvi , Oriya (macrolanguage) , Oromo , Ossetian , Pangasinan , Pampanga , Panjabi , Papiamento , Picard , Pennsylvania German , Pfaelzisch , Pitcairn-Norfolk , Pali , Piemontese , Western Panjabi , Pontic , Polish , Portuguese , Pushto , Quechua , Vlax Romani , Romansh , Romanian , Rusyn , Rundi , Macedo-Romanian , Russian , Sango , Yakut , Sanskrit , Sicilian , Scots , Samogitian , Sinhala , Slovak , Slovenian , Northern Sami , Samoan , Shona , Sindhi , Somali , Southern Sotho , Spanish , Albanian , Sardinian , Sranan Tongo , Serbian , Swati , Saterfriesisch , Sundanese , Swahili (macrolanguage) , Swedish , Silesian , Tahitian , Tamil , Tatar , Tulu , Telugu , Tama (Colombia) , Tetum , Tajik , Tagalog , Thai , Tigrinya , Tonga (Tonga Islands) , Tok Pisin , Tswana , Tsonga , Turkmen , Tumbuka , Turkish , Twi , Tuvinian , Udmurt , Uighur , Ukrainian , Urdu , Uzbek , Venetian , Venda , Veps , Vietnamese , Vlaams , Volapük , Võro , Waray (Philippines) , Walloon , Wolof , Wu Chinese , Kalmyk , Xhosa , Mingrelian , Yiddish , Yoruba , Yue Chinese , Zeeuws , Zhuang , Chinese , Zulu , and Dotyali
Description:
Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018.
The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages).
For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias].
The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast).
Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day.
The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
Rights:
Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) , http://creativecommons.org/licenses/by-sa/3.0/ , and PUB
Creator:
Jan Patočka
Publisher:
Str. 46–88. Stať.
Type:
Text
Subject:
1975 , 1979/25 , 1981/6 , 1981/7 , 1988/28 , 1988/31 , 1988/32 , 1988/34 , 1994/7 , 1996/4 , 1996/7 , 1998/3 , 1999/8 , 2 , 2001/9 , 2002/21 , 2002/6 , 2006/1 , 2007/1 , 2008/3 , bg , cs , de , en , es , fr , fulltext , hu , it , lt , no , pl , ru , SS-3/PD-III , sv , uk , and v
Language:
Czech , English , Bulgarian , French , Italian , Lithuanian , Hungarian , German , Norwegian , Polish , Russian , Spanish , Swedish , and Ukrainian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.