Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Rosa, Rudolf
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
Wikipedia , text corpora , and monolingual corpus
Language:
Abkhazian , Achinese , Adyghe , Afrikaans , Akan , Tosk Albanian , Amharic , Old English (ca. 450-1100) , Arabic , Official Aramaic (700-300 BCE) , Aragonese , Egyptian Arabic , Assamese , Asturian , Atikamekw , Avaric , Aymara , South Azerbaijani , Azerbaijani , Bashkir , Bambara , Bavarian , Central Bikol , Belarusian , Bengali , Bislama , Banjar , Tibetan , Bosnian , Bishnupriya , Breton , Buginese , Bulgarian , Russia Buriat , Catalan , Min Dong Chinese , Cebuano , Czech , Chamorro , Chechen , Cherokee , Church Slavic , Chuvash , Cheyenne , Central Kurdish , Cornish , Corsican , Cree , Crimean Tatar , Kashubian , Welsh , Danish , German , Dinka , Dimli (individual language) , Dhivehi , Lower Sorbian , Dzongkha , Modern Greek (1453-) , English , Esperanto , Estonian , Basque , Ewe , Extremaduran , Faroese , Persian , Fijian , Finnish , French , Arpitan , Northern Frisian , Western Frisian , Fulah , Friulian , Gagauz , Gan Chinese , Scottish Gaelic , Irish , Galician , Gilaki , Manx , Goan Konkani , Gothic , Guarani , Gujarati , Hakka Chinese , Haitian , Hausa , Hawaiian , Serbo-Croatian , Hebrew , Herero , Fiji Hindi , Hindi , Hiri Motu , Croatian , Upper Sorbian , Hungarian , Armenian , Igbo , Ido , Inuktitut , Interlingue , Iloko , Interlingua (International Auxiliary Language Association) , Indonesian , Inupiaq , Icelandic , Italian , Jamaican Creole English , Javanese , Lojban , Japanese , Kara-Kalpak , Kabyle , Kalaallisut , Kannada , Kashmiri , Georgian , Kanuri , Kazakh , Kabardian , Kabiyè , Khmer , Kikuyu , Kinyarwanda , Kirghiz , Komi-Permyak , Komi , Kongo , Korean , Karachay-Balkar , Kölsch , Kurdish , Ladino , Lao , Latin , Latvian , Lak , Lezghian , Ligurian , Limburgan , Lingala , Lithuanian , Lombard , Northern Luri , Latgalian , Luxembourgish , Ganda , Literary Chinese , Marshallese , Maithili , Malayalam , Marathi , Moksha , Eastern Mari , Minangkabau , Macedonian , Malagasy , Maltese , Mongolian , Maori , Western Mari , Malay (macrolanguage) , Creek , Mirandese , Burmese , Erzya , Mazanderani , Min Nan Chinese , Neapolitan , Nauru , Navajo , Ndonga , Low German , Nepali (macrolanguage) , Newari , Dutch , Norwegian Nynorsk , Norwegian , Novial , Pedi , Nyanja , Occitan (post 1500) , Livvi , Oriya (macrolanguage) , Oromo , Ossetian , Pangasinan , Pampanga , Panjabi , Papiamento , Picard , Pennsylvania German , Pfaelzisch , Pitcairn-Norfolk , Pali , Piemontese , Western Panjabi , Pontic , Polish , Portuguese , Pushto , Quechua , Vlax Romani , Romansh , Romanian , Rusyn , Rundi , Macedo-Romanian , Russian , Sango , Yakut , Sanskrit , Sicilian , Scots , Samogitian , Sinhala , Slovak , Slovenian , Northern Sami , Samoan , Shona , Sindhi , Somali , Southern Sotho , Spanish , Albanian , Sardinian , Sranan Tongo , Serbian , Swati , Saterfriesisch , Sundanese , Swahili (macrolanguage) , Swedish , Silesian , Tahitian , Tamil , Tatar , Tulu , Telugu , Tama (Colombia) , Tetum , Tajik , Tagalog , Thai , Tigrinya , Tonga (Tonga Islands) , Tok Pisin , Tswana , Tsonga , Turkmen , Tumbuka , Turkish , Twi , Tuvinian , Udmurt , Uighur , Ukrainian , Urdu , Uzbek , Venetian , Venda , Veps , Vietnamese , Vlaams , Volapük , Võro , Waray (Philippines) , Walloon , Wolof , Wu Chinese , Kalmyk , Xhosa , Mingrelian , Yiddish , Yoruba , Yue Chinese , Zeeuws , Zhuang , Chinese , Zulu , and Dotyali
Description:
Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018.
The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages).
For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias].
The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast).
Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day.
The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
Rights:
Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) , http://creativecommons.org/licenses/by-sa/3.0/ , and PUB
Creator:
Jan Patočka
Publisher:
Ed. I. Chvatík, str. 1–276. Předn. cykl. [Přepis mgf. záznamu soukromých přednášek z roku 1973.] — 2. otisk in: Péče o duši II (SS-2/PD-II), Praha 1999, str. 149–355 (v. 1999/6). — Části z 1. a 5. přednášky (str. 8–16 a 84–108) otištěny in: Filosofický časopis 40 (1992), č. 6, str. 921–943. S něm. a angl. shrnutím. — Části diskuse k 8. a 9. přednášce (str. 168–217) otištěny in: M. Bednář, České myšlení, Praha (Filosofia) 1996. str. 357–390. — Srv. 1979/29, 1979/30, 1988/12, 1991/3.
Type:
Text
Subject:
1979 , 1979/29 , 1979/30 , 1983/34 , 1988/12 , 1991/3 , 1991/6 , 1997/1 , 1997/5 , 1998/8 , 1999/6 , 2002/20 , cs , en , es , fr , it , lt , pl , SS-2/PD-II , str–16–108 , str–217 , and sv
Language:
English , French , Italian , Lithuanian , Polish , Spanish , Swedish , and Czech
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Jan Patočka
Publisher:
Ed. I. Chvatík, str. 1–276. Předn. cykl. [Přepis mgf. záznamu soukromých přednášek z roku 1973.]
Type:
Text
Subject:
1979 , 1979/29 , 1979/30 , 1983/34 , 1988/12 , 1991/3 , 1991/6 , 1997/1 , 1997/5 , 1998/8 , 1999/6 , 2002/20 , cs , en , es , fr , it , lt , pl , Předn. cykl. , SS-2/PD-II , and sv
Language:
Czech , English , French , Italian , Lithuanian , Polish , Spanish , and Swedish
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Jan Patočka
Publisher:
Str. 46–88. Stať.
Type:
Text
Subject:
1975 , 1979/25 , 1981/6 , 1981/7 , 1988/28 , 1988/31 , 1988/32 , 1988/34 , 1994/7 , 1996/4 , 1996/7 , 1998/3 , 1999/8 , 2 , 2001/9 , 2002/21 , 2002/6 , 2006/1 , 2007/1 , 2008/3 , bg , cs , de , en , es , fr , fulltext , hu , it , lt , no , pl , ru , SS-3/PD-III , sv , uk , and v
Language:
Czech , English , Bulgarian , French , Italian , Lithuanian , Hungarian , German , Norwegian , Polish , Russian , Spanish , Swedish , and Ukrainian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Král, Ivan,
Type:
text and publikace fotografické
Subject:
Geografie Česka a Slovenska, reálie, cestování , Architektura , dějiny měst , památky umělecké , přehledná zpracování dějin českých zemí (chronologicky) , and města, obce
Language:
Czech , English , German , Spanish , French , Italian , and Russian
Description:
Obálkový název: Naše Praha
Rights:
unknown
Creator:
Jan Patočka
Publisher:
Str. 1–45. Stať.
Type:
Text
Subject:
1975 , 1979/25 , 1981/6 , 1981/7 , 1988/28 , 1988/31 , 1988/32 , 1988/34 , 1994/7 , 1996/4 , 1996/7 , 1998/3 , 1999/8 , 2001/9 , 2002/1 , 2002/21 , 2002/5 , 2006/1 , 2007/1 , 2008/3 , bg , cs , de , en , es , fr , fulltext , hu , it , lt , no , pl , ru , SS-3/PD-III , sv , and uk
Language:
Czech , English , Bulgarian , French , Italian , Lithuanian , Hungarian , German , Norwegian , Polish , Russian , Spanish , Swedish , and Ukrainian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Jan Patočka
Publisher:
1.3, 48 s. Stať. [Předloha pro slovenský překlad (v. 1967/2). Začátky textů se mírně liší.] — 2. otisk in: Fenomenologické spisy II (SS-7/Fen-II), Praha 2009, str. 202–237 (v. 2009/1).
Type:
Text
Subject:
1967/2 , 1969/8 , 1970/10 , 1972/1 , 1972/2 , 1976/7 , 1980 , 1988/29 , 1989/16 , 1991/2 , 1996/7 , 2003/23 , 2004/10 , 2009/1 , cs , de , en , es , fr , hu , it , sk , SS-7/Fen-II , and stať
Language:
English , French , Italian , Hungarian , German , Slovak , Spanish , and Czech
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Type:
corpus
Language:
Czech , Danish , Dutch , English , Finnish , French , German , Hungarian , Italian , Polish , Portuguese , Russian , Spanish , Swedish , Turkish , Chinese , Hebrew , Japanese , Korean , and Thai
Description:
28 speech databases containing broadband recordings from 550 adults and 50 children per language. Contains interesting phonetically rich material. All orthographically transcribed. Speaker information included for gender, age, accent. Including pronunciation lexicon.
Rights:
Not specified
Creator:
Jan Patočka
Publisher:
Str. 67–85. Stať. [Přepracovaná česká verze přednášky Die Funktion der Literatur in der Gesellschaft, v. 1968/17.]
Type:
Text
Subject:
1969 , bg , cs , de , es , fr , hu , and it
Language:
Czech , Bulgarian , French , Italian , Hungarian , German , and Spanish
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Kondratyuk, Dan and Straka, Milan
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
tool and toolService
Subject:
syntax , dependency parser , and universal dependencies
Language:
Ancient Greek (to 1453) , Arabic , Basque , Bulgarian , Croatian , Czech , Danish , Dutch , English , Estonian , Finnish , French , German , Gothic , Modern Greek (1453-) , Hebrew , Hindi , Hungarian , Indonesian , Irish , Italian , Japanese , Latin , Norwegian , Church Slavic , Persian , Polish , Portuguese , Romanian , Slovenian , Spanish , Swedish , Tamil , Catalan , Chinese , Galician , Kazakh , Latvian , Russian , Turkish , Coptic , Sanskrit , Slovak , Ukrainian , Uighur , Vietnamese , Belarusian , Korean , Lithuanian , Urdu , Russia Buriat , Northern Kurdish , Northern Sami , Upper Sorbian , Afrikaans , Yue Chinese , Marathi , Serbian , Swedish Sign Language , Telugu , Amharic , Armenian , Breton , Faroese , Komi-Zyrian , Nigerian Pidgin , Old French (842-ca. 1400) , Tagalog , Thai , Warlpiri , Yoruba , Akkadian , Bambara , Erzya , and Maltese
Description:
Pretrained model weights for the UDify model, and extracted BERT weights in pytorch-transformers format. Note that these weights slightly differ from those used in the paper.
Rights:
Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) , http://creativecommons.org/licenses/by-sa/4.0/ , and PUB