Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Calligaris, Luigi
Type:
model:monograph and TEXT
Language:
French , Arabic , Latin , Italian , Spanish , Portuguese , German , English , and Modern Greek (1453-)
Rights:
http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
Creator:
Palacký, František,
Type:
text and zprávy výzkumné
Subject:
Historická věda. Pomocné vědy historické. Archivnictví , historiografie , vztahy česko-italské , české země 1792-1847 , dějepisectví, historické vědy, historici , Itálie , and přehledná zpracování (tematicky)
Language:
German , Italian , and Latin
Description:
Název na doplňkové titulní stránce: Palacky's italienische Reise im Jahre 1837
Rights:
unknown
Type:
text and sborníky jubilejní
Subject:
Historická věda. Pomocné vědy historické. Archivnictví , Ehrle, Franz, , historiografie , and zahraniční periodika a sborníky
Language:
Italian , German , and Latin
Rights:
unknown
Creator:
<<ze >>Žerotína, Karel,
Type:
text , tisky pamětní , korespondence , and edice
Subject:
Historická věda. Pomocné vědy historické. Archivnictví , <<ze >>Žerotína, Karel, , šlechtici , myšlení politické , české země 1526-1620 , and šlechta, buržoazie, měšťanstvo, podnikatelé
Language:
Czech , French , German , Italian , and Latin
Description:
Bibliofilie, Sáňka 5495, 600 výtisků a 10 číslovaných výtisků na měditiskovém papíře Sanders, and Obálkový název: Deset listů Karla st. z Žerotína
Rights:
unknown
Creator:
Rosa, Rudolf
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
Wikipedia , text corpora , and monolingual corpus
Language:
Abkhazian , Achinese , Adyghe , Afrikaans , Akan , Tosk Albanian , Amharic , Old English (ca. 450-1100) , Arabic , Official Aramaic (700-300 BCE) , Aragonese , Egyptian Arabic , Assamese , Asturian , Atikamekw , Avaric , Aymara , South Azerbaijani , Azerbaijani , Bashkir , Bambara , Bavarian , Central Bikol , Belarusian , Bengali , Bislama , Banjar , Tibetan , Bosnian , Bishnupriya , Breton , Buginese , Bulgarian , Russia Buriat , Catalan , Min Dong Chinese , Cebuano , Czech , Chamorro , Chechen , Cherokee , Church Slavic , Chuvash , Cheyenne , Central Kurdish , Cornish , Corsican , Cree , Crimean Tatar , Kashubian , Welsh , Danish , German , Dinka , Dimli (individual language) , Dhivehi , Lower Sorbian , Dzongkha , Modern Greek (1453-) , English , Esperanto , Estonian , Basque , Ewe , Extremaduran , Faroese , Persian , Fijian , Finnish , French , Arpitan , Northern Frisian , Western Frisian , Fulah , Friulian , Gagauz , Gan Chinese , Scottish Gaelic , Irish , Galician , Gilaki , Manx , Goan Konkani , Gothic , Guarani , Gujarati , Hakka Chinese , Haitian , Hausa , Hawaiian , Serbo-Croatian , Hebrew , Herero , Fiji Hindi , Hindi , Hiri Motu , Croatian , Upper Sorbian , Hungarian , Armenian , Igbo , Ido , Inuktitut , Interlingue , Iloko , Interlingua (International Auxiliary Language Association) , Indonesian , Inupiaq , Icelandic , Italian , Jamaican Creole English , Javanese , Lojban , Japanese , Kara-Kalpak , Kabyle , Kalaallisut , Kannada , Kashmiri , Georgian , Kanuri , Kazakh , Kabardian , Kabiyè , Khmer , Kikuyu , Kinyarwanda , Kirghiz , Komi-Permyak , Komi , Kongo , Korean , Karachay-Balkar , Kölsch , Kurdish , Ladino , Lao , Latin , Latvian , Lak , Lezghian , Ligurian , Limburgan , Lingala , Lithuanian , Lombard , Northern Luri , Latgalian , Luxembourgish , Ganda , Literary Chinese , Marshallese , Maithili , Malayalam , Marathi , Moksha , Eastern Mari , Minangkabau , Macedonian , Malagasy , Maltese , Mongolian , Maori , Western Mari , Malay (macrolanguage) , Creek , Mirandese , Burmese , Erzya , Mazanderani , Min Nan Chinese , Neapolitan , Nauru , Navajo , Ndonga , Low German , Nepali (macrolanguage) , Newari , Dutch , Norwegian Nynorsk , Norwegian , Novial , Pedi , Nyanja , Occitan (post 1500) , Livvi , Oriya (macrolanguage) , Oromo , Ossetian , Pangasinan , Pampanga , Panjabi , Papiamento , Picard , Pennsylvania German , Pfaelzisch , Pitcairn-Norfolk , Pali , Piemontese , Western Panjabi , Pontic , Polish , Portuguese , Pushto , Quechua , Vlax Romani , Romansh , Romanian , Rusyn , Rundi , Macedo-Romanian , Russian , Sango , Yakut , Sanskrit , Sicilian , Scots , Samogitian , Sinhala , Slovak , Slovenian , Northern Sami , Samoan , Shona , Sindhi , Somali , Southern Sotho , Spanish , Albanian , Sardinian , Sranan Tongo , Serbian , Swati , Saterfriesisch , Sundanese , Swahili (macrolanguage) , Swedish , Silesian , Tahitian , Tamil , Tatar , Tulu , Telugu , Tama (Colombia) , Tetum , Tajik , Tagalog , Thai , Tigrinya , Tonga (Tonga Islands) , Tok Pisin , Tswana , Tsonga , Turkmen , Tumbuka , Turkish , Twi , Tuvinian , Udmurt , Uighur , Ukrainian , Urdu , Uzbek , Venetian , Venda , Veps , Vietnamese , Vlaams , Volapük , Võro , Waray (Philippines) , Walloon , Wolof , Wu Chinese , Kalmyk , Xhosa , Mingrelian , Yiddish , Yoruba , Yue Chinese , Zeeuws , Zhuang , Chinese , Zulu , and Dotyali
Description:
Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018.
The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages).
For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias].
The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast).
Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day.
The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
Rights:
Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) , http://creativecommons.org/licenses/by-sa/3.0/ , and PUB
Type:
corpus
Language:
Danish , Dutch , English , Finnish , French , German , Italian , Latin , Portuguese , Russian , Spanish , Swedish , and Telugu
Description:
Possibility to download or to browse free electronic books; Angebot: Download von und Online-Zugang zu frei verfügbaren E-Books; deutschsprachige Literatur stellt nur einen Teilbereich der verfügbaren E-Books dar
Rights:
Not specified
Type:
text and sborníky jubilejní
Subject:
Dějiny Česka a Slovenska , Hledíková, Zdeňka, , historici čeští , jubilea životní , and české (československé) sborníky a kolektivní monografie
Language:
Italian , English , German , French , Latin , and Czech
Description:
400 výt.
Rights:
unknown
Creator:
Svobodová, Milada,
Type:
text , katalogy , bibliografie , monografie , and rejstříky
Subject:
Univerzální knihovny. Veřejné knihovny. Soukromé knihovny , Bibliografie. Katalogy , Hannl, Václav Řehoř, , Troilo von Lessoth, Franz Gottfried, , Czerninové z Chudenic (rod) , Lobkowiczové (rod) , rukopisy , knihovny soukromé , and knihovny šlechtické
Language:
Czech , English , French , German , Italian , Latin , and Spanish
Rights:
unknown
Creator:
Svobodová, Milada,
Type:
text and katalogy
Subject:
Rukopisy, prvotisky, staré tisky. Vzácná a pozoruhodná díla , Czerninové z Chudenic (rod) , Lobkowiczové (rod) , rukopisy , knihovny soukromé , and knihovny šlechtické
Language:
Czech , English , French , German , Italian , and Latin
Description:
Název z titulní obrazovky
Rights:
unknown
Creator:
Dudík, Beda,
Type:
text and monografie
Subject:
Dějiny Evropy , válka třicetiletá (1618-1648) , české země 1620-1740 , přehledná zpracování (tematicky) , Švédsko , and světové dějiny novověku (1492-1918)
Language:
German , Italian , and Latin
Rights:
unknown