Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Majliš, Martin
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
multilingual corpora
Language:
Afrikaans , Tosk Albanian , Amharic , Arabic , Aragonese , Egyptian Arabic , Asturian , Azerbaijani , Belarusian , Bengali , Bosnian , Bishnupriya , Breton , Buginese , Bulgarian , Catalan , Cebuano , Czech , Chuvash , Corsican , Welsh , Danish , German , Dimli (individual language) , Modern Greek (1453-) , English , Esperanto , Estonian , Basque , Faroese , Persian , Finnish , French , Western Frisian , Gan Chinese , Scottish Gaelic , Irish , Galician , Gilaki , Gujarati , Haitian , Serbo-Croatian , Hebrew , Fiji Hindi , Hindi , Croatian , Upper Sorbian , Hungarian , Armenian , Ido , Interlingua (International Auxiliary Language Association) , Indonesian , Icelandic , Italian , Javanese , Japanese , Kannada , Georgian , Kazakh , Korean , Kurdish , Latin , Latvian , Limburgan , Lithuanian , Lombard , Luxembourgish , Malayalam , Marathi , Macedonian , Malagasy , Mongolian , Maori , Malay (macrolanguage) , Burmese , Neapolitan , Low German , Nepali (macrolanguage) , Newari , Dutch , Norwegian Nynorsk , Norwegian , Occitan (post 1500) , Ossetian , Pampanga , Piemontese , Polish , Portuguese , Quechua , Romanian , Russian , Yakut , Sicilian , Scots , Slovak , Slovenian , Spanish , Albanian , Serbian , Sundanese , Swahili (macrolanguage) , Swedish , Tamil , Tatar , Telugu , Tajik , Tagalog , Thai , Turkish , Ukrainian , Urdu , Uzbek , Venetian , Vietnamese , Volapük , Waray (Philippines) , Walloon , Yiddish , Yoruba , and Chinese
Description:
A set of corpora for 120 languages automatically collected from wikipedia and the web.
Collected using the W2C toolset: http://hdl.handle.net/11858/00-097C-0000-0022-60D6-1
Rights:
Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) , http://creativecommons.org/licenses/by-sa/3.0/ , and PUB
Creator:
Hoang, Duc Tam and Bojar, Ondřej
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
test data , parallel corpus , and Vietnamese
Language:
Vietnamese , Czech , English , German , French , Spanish , and Russian
Description:
We provide the Vietnamese version of the multi-lingual test set from WMT 2013 [1] competition. The Vietnamese version was manually translated from English. For completeness, this record contains the 3000 sentences in all the WMT 2013 original languages (Czech, English, French, German, Russian and Spanish), extended with our Vietnamese version. Test set is used in [2] to evaluate translation between Czech, English and Vietnamese.
References
1. http://www.statmt.org/wmt13/evaluation-task.html
2. Duc Tam Hoang and Ondřej Bojar, The Prague Bulletin of Mathematical Linguistics. Volume 104, Issue 1, Pages 75--86, ISSN 1804-0462. 9/2015
Rights:
Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) , http://creativecommons.org/licenses/by-nc-sa/4.0/ , and PUB
Publisher:
University of Leipzig
Type:
corpus
Language:
Afrikaans , Albanian , Bulgarian , Catalan , Chinese , Croatian , Czech , Danish , Dutch , English , Esperanto , Estonian , Finnish , French , German , Hungarian , Icelandic , Indonesian , Italian , Japanese , Korean , Latin , Latvian , Lithuanian , Malay (macrolanguage) , Norwegian , Occitan (post 1500) , Romanian , Russian , Slovak , Slovenian , Spanish , Sundanese , Swedish , Tagalog , Turkish , Vietnamese , and Welsh
Description:
Collected from newspaper texts, webcrawling, etc.: words (+frequency), cooccurrences (+graph), left/right neighbours, example sentences
Rights:
Not specified
Creator:
Ailly, Pierre d',
Type:
text and traktáty
Subject:
Křesťanská teologie. Dogmatická teologie , traktáty , edice , duchovenstvo katolické , teologie, ikonografie, zbožnost, hagiografie , and světové dějiny středověku (do r. 1492)
Language:
French and Latin
Rights:
unknown
Type:
text and sborníky
Subject:
Dějiny (obecně) , Plaschka, Richard Georg, , historici rakouští , and zahraniční periodika a sborníky
Language:
German , French , and English
Rights:
unknown
Creator:
Novák, Pavel,
Publisher:
[Ecomusée du Cresot-Montceau],
Type:
katalogy výstav
Subject:
Architektura , města průmyslová , architektura městská , funkcionalismus , Československo 1918-1992 , architektura, architekti , české země 1848-1918 , and zahraniční výstavy
Language:
French
Description:
Katalog k výstavě probíhající v rámci přehlídky české kultury ve Francii květen - prosinec 2002, Text francouzsky a česky, and Z obsahu: Novák, P. : Průmyslová výstavba ve Zlíně/ Horňáková, L. : Výrazné osobnosti zlínské funkcionalistické architektury/ Zemánková, H. : Jak uvést v soulad ochranu kulturního dědictví Zlína s rozvojem města?/ Kabát, M. : Konverze výrobního objektu č. 24 v závodu T. Bati ve Zlíně/ Sedlák, J. : Zlínské paradoxy
Rights:
unknown
Creator:
Sinigaglia, Leone
Format:
print
Type:
supplement , model:supplement , and TEXT
Language:
English , French , and German
Rights:
http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
Creator:
Sinigaglia, Leone
Format:
print
Type:
supplement , model:supplement , and TEXT
Language:
English , French , and German
Rights:
http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
Creator:
Sinigaglia, Leone
Format:
print
Type:
supplement , model:supplement , and TEXT
Language:
English , French , and German
Rights:
http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
Creator:
Beneš, Pavel,
Type:
text and studie
Subject:
Dějiny zemí střední Evropy , Metoděj, , legendy cyrilometodějské , jména místní , sv. Konstantin a Metoděj , and teologie, ikonografie, zbožnost, hagiografie
Language:
French
Rights:
unknown