Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Rosa, Rudolf
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
Wikipedia , text corpora , and monolingual corpus
Language:
Abkhazian , Achinese , Adyghe , Afrikaans , Akan , Tosk Albanian , Amharic , Old English (ca. 450-1100) , Arabic , Official Aramaic (700-300 BCE) , Aragonese , Egyptian Arabic , Assamese , Asturian , Atikamekw , Avaric , Aymara , South Azerbaijani , Azerbaijani , Bashkir , Bambara , Bavarian , Central Bikol , Belarusian , Bengali , Bislama , Banjar , Tibetan , Bosnian , Bishnupriya , Breton , Buginese , Bulgarian , Russia Buriat , Catalan , Min Dong Chinese , Cebuano , Czech , Chamorro , Chechen , Cherokee , Church Slavic , Chuvash , Cheyenne , Central Kurdish , Cornish , Corsican , Cree , Crimean Tatar , Kashubian , Welsh , Danish , German , Dinka , Dimli (individual language) , Dhivehi , Lower Sorbian , Dzongkha , Modern Greek (1453-) , English , Esperanto , Estonian , Basque , Ewe , Extremaduran , Faroese , Persian , Fijian , Finnish , French , Arpitan , Northern Frisian , Western Frisian , Fulah , Friulian , Gagauz , Gan Chinese , Scottish Gaelic , Irish , Galician , Gilaki , Manx , Goan Konkani , Gothic , Guarani , Gujarati , Hakka Chinese , Haitian , Hausa , Hawaiian , Serbo-Croatian , Hebrew , Herero , Fiji Hindi , Hindi , Hiri Motu , Croatian , Upper Sorbian , Hungarian , Armenian , Igbo , Ido , Inuktitut , Interlingue , Iloko , Interlingua (International Auxiliary Language Association) , Indonesian , Inupiaq , Icelandic , Italian , Jamaican Creole English , Javanese , Lojban , Japanese , Kara-Kalpak , Kabyle , Kalaallisut , Kannada , Kashmiri , Georgian , Kanuri , Kazakh , Kabardian , Kabiyè , Khmer , Kikuyu , Kinyarwanda , Kirghiz , Komi-Permyak , Komi , Kongo , Korean , Karachay-Balkar , Kölsch , Kurdish , Ladino , Lao , Latin , Latvian , Lak , Lezghian , Ligurian , Limburgan , Lingala , Lithuanian , Lombard , Northern Luri , Latgalian , Luxembourgish , Ganda , Literary Chinese , Marshallese , Maithili , Malayalam , Marathi , Moksha , Eastern Mari , Minangkabau , Macedonian , Malagasy , Maltese , Mongolian , Maori , Western Mari , Malay (macrolanguage) , Creek , Mirandese , Burmese , Erzya , Mazanderani , Min Nan Chinese , Neapolitan , Nauru , Navajo , Ndonga , Low German , Nepali (macrolanguage) , Newari , Dutch , Norwegian Nynorsk , Norwegian , Novial , Pedi , Nyanja , Occitan (post 1500) , Livvi , Oriya (macrolanguage) , Oromo , Ossetian , Pangasinan , Pampanga , Panjabi , Papiamento , Picard , Pennsylvania German , Pfaelzisch , Pitcairn-Norfolk , Pali , Piemontese , Western Panjabi , Pontic , Polish , Portuguese , Pushto , Quechua , Vlax Romani , Romansh , Romanian , Rusyn , Rundi , Macedo-Romanian , Russian , Sango , Yakut , Sanskrit , Sicilian , Scots , Samogitian , Sinhala , Slovak , Slovenian , Northern Sami , Samoan , Shona , Sindhi , Somali , Southern Sotho , Spanish , Albanian , Sardinian , Sranan Tongo , Serbian , Swati , Saterfriesisch , Sundanese , Swahili (macrolanguage) , Swedish , Silesian , Tahitian , Tamil , Tatar , Tulu , Telugu , Tama (Colombia) , Tetum , Tajik , Tagalog , Thai , Tigrinya , Tonga (Tonga Islands) , Tok Pisin , Tswana , Tsonga , Turkmen , Tumbuka , Turkish , Twi , Tuvinian , Udmurt , Uighur , Ukrainian , Urdu , Uzbek , Venetian , Venda , Veps , Vietnamese , Vlaams , Volapük , Võro , Waray (Philippines) , Walloon , Wolof , Wu Chinese , Kalmyk , Xhosa , Mingrelian , Yiddish , Yoruba , Yue Chinese , Zeeuws , Zhuang , Chinese , Zulu , and Dotyali
Description:
Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018.
The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages).
For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias].
The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast).
Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day.
The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
Rights:
Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) , http://creativecommons.org/licenses/by-sa/3.0/ , and PUB
Creator:
Križnar, Naško,
Type:
text and studie
Subject:
Etnologie. Etnografie. Folklor , Murko, Matija, , fotografie , národopisci , and národopis, jednotlivci
Language:
Slovenian
Rights:
unknown
Creator:
Blažič, Milena,
Type:
text and studie
Subject:
Slovanské literatury (o nich) , Cankar, Ivan, , literatura slovinská , and spisovatelé slovinští
Language:
Slovenian
Rights:
unknown
Creator:
Mikuž, Jure,
Type:
text and studie
Subject:
Hudba , Gallus-Handl, Jacob, , skladatelé , dějiny umění , Slovinsko , světové dějiny 1492-1648 , and dějiny umění, mecenát
Language:
Slovenian and Czech
Description:
Souběžně uveden text slovinského originálu a českého překladu
Rights:
unknown
Creator:
Mikuž, Jure,
Publisher:
Skupščina občine Ribnica,
Subject:
Gallus-Handl, Jacob, , skladatelé , dějiny umění , skladatelé slovinští , světové dějiny 1492-1648 , Slovinsko , and dějiny umění, mecenát
Language:
Slovenian
Rights:
unknown
Creator:
Mlačnik, Primož,
Type:
text and studie
Subject:
Slovanské literatury (o nich) , Německá literatura, německy psaná (o ní) , Cankar, Ivan, , Kafka, Franz, , spisovatelé slovinští , literatura slovinská , and literatura německojazyčná
Language:
Slovenian
Rights:
unknown
Creator:
Jensterle-Doležal, Alenka,
Subject:
vztahy česko-slovinské , vztahy slovinsko-české , vztahy literární , literatura česká , literatura slovinská , světové dějiny 1789-1918 , Slovinsko , české země 1792-1918 , and literatura, spisovatelé
Language:
Slovenian
Rights:
unknown
Creator:
Domej, Teodor
Subject:
Majar-Ziljski, Matija, , jazyk slovinský , otázka jazyková , politika jazyková , světové dějiny 1789-1918 , Slovinsko , and národnosti, vztahy mezi národnostmi a národní hnutí
Language:
Slovenian
Description:
Matija Majars Ansichten zur Sprachenfrage.
Rights:
unknown
Creator:
Junaht, Janez
Subject:
historiografie , přehledná zpracování (tematicky) , přehledná zpracování světových dějin (chronologicky) , Slovinsko , and historiografie, vědecké projekty
Language:
Slovenian
Description:
Das Verstehen der Geschichte und das Problem der Reinterpretation der geschichtlichen Entwicklung Sloweniens.
Rights:
unknown
Creator:
Granda, Stane
Subject:
Majar-Ziljski, Matija, , vztahy česko-slovinské , vědci slovinští , myšlení politické , činnost politická , politické dějiny, politici , světové dějiny 1789-1918 , Habsburská monarchie , Slovinsko , and české země 1792-1918
Language:
Slovenian
Description:
Die Bedeutung von Matija Majar für die slowenische Geschichte.
Rights:
unknown