Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Mareček, David , Yu, Zhiwei , Zeman, Daniel , and Žabokrtský, Zdeněk
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
part of speech , tagging , semi-supervised , and cross-language
Language:
Belarusian , Bosnian , Bulgarian , Czech , Serbo-Croatian , Croatian , Upper Sorbian , Macedonian , Polish , Russian , Slovak , Slovenian , Serbian , Ukrainian , Latvian , Lithuanian , Afrikaans , Danish , German , English , Faroese , Western Frisian , Swiss German , Icelandic , Limburgan , Luxembourgish , Low German , Dutch , Norwegian Nynorsk , Norwegian , Scots , Swedish , Yiddish , Aragonese , Asturian , Catalan , French , Galician , Haitian , Italian , Latin , Lombard , Neapolitan , Piemontese , Portuguese , Romanian , Spanish , Venetian , Walloon , Breton , Welsh , Scottish Gaelic , Irish , Modern Greek (1453-) , Armenian , Albanian , Dimli (individual language) , Persian , Gilaki , Kurdish , Tajik , Bengali , Bishnupriya , Gujarati , Fiji Hindi , Hindi , Marathi , Nepali (macrolanguage) , Urdu , Amharic , Arabic , Egyptian Arabic , Hebrew , Estonian , Finnish , Hungarian , Basque , Georgian , Chuvash , Azerbaijani , Turkish , Uzbek , Kazakh , Tatar , Yakut , Korean , Mongolian , Telugu , Kannada , Malayalam , Tamil , Newari , Vietnamese , Indonesian , Javanese , Malagasy , Maori , Malay (macrolanguage) , Pampanga , Sundanese , Tagalog , Waray (Philippines) , Swahili (macrolanguage) , Esperanto , Ido , Interlingua (International Auxiliary Language Association) , and Volapük
Description:
Texts in 107 languages from the W2C corpus (http://hdl.handle.net/11858/00-097C-0000-0022-6133-9), first 1,000,000 tokens per language, tagged by the delexicalized tagger described in Yu et al. (2016, LREC, Portorož, Slovenia).
Rights:
Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) , http://creativecommons.org/licenses/by-sa/4.0/ , and PUB
Creator:
Mareček, David , Yu, Zhiwei , Zeman, Daniel , and Žabokrtský, Zdeněk
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
part of speech , tagging , semi-supervised , and cross-language
Language:
Belarusian , Bosnian , Bulgarian , Czech , Serbo-Croatian , Croatian , Upper Sorbian , Macedonian , Polish , Russian , Slovak , Slovenian , Serbian , Ukrainian , Latvian , Lithuanian , Afrikaans , Danish , German , English , Faroese , Western Frisian , Swiss German , Icelandic , Limburgan , Luxembourgish , Low German , Dutch , Norwegian Nynorsk , Norwegian , Scots , Swedish , Yiddish , Aragonese , Asturian , Catalan , French , Galician , Haitian , Italian , Latin , Lombard , Neapolitan , Piemontese , Portuguese , Romanian , Spanish , Venetian , Walloon , Breton , Welsh , Scottish Gaelic , Irish , Modern Greek (1453-) , Armenian , Albanian , Dimli (individual language) , Persian , Gilaki , Kurdish , Tajik , Bengali , Bishnupriya , Gujarati , Fiji Hindi , Hindi , Marathi , Nepali (macrolanguage) , Urdu , Amharic , Arabic , Egyptian Arabic , Hebrew , Estonian , Finnish , Hungarian , Basque , Georgian , Chuvash , Azerbaijani , Turkish , Uzbek , Kazakh , Tatar , Yakut , Korean , Mongolian , Telugu , Kannada , Malayalam , Tamil , Newari , Vietnamese , Indonesian , Javanese , Malagasy , Maori , Malay (macrolanguage) , Pampanga , Sundanese , Tagalog , Waray (Philippines) , Swahili (macrolanguage) , Esperanto , Ido , Interlingua (International Auxiliary Language Association) , and Volapük
Description:
Texts in 107 languages from the W2C corpus (http://hdl.handle.net/11858/00-097C-0000-0022-6133-9), first 1,000,000 tokens per language, tagged by the delexicalized tagger described in Yu et al. (2016, LREC, Portorož, Slovenia).
Changes in version 1.1:
1. Universal Dependencies tagset instead of the older and smaller Google Universal POS tagset.
2. SVM classifier trained on Universal Dependencies 1.2 instead of HamleDT 2.0.
3. Balto-Slavic languages, Germanic languages and Romance languages were tagged by classifier trained only on the respective group of languages. Other languages were tagged by a classifier trained on all available languages. The "c7" combination from version 1.0 is no longer used.
Rights:
Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) , http://creativecommons.org/licenses/by-sa/4.0/ , and PUB
Creator:
Mulder, Willem Johannes Maria,
Type:
text
Subject:
Křesťanská teologie. Dogmatická teologie , Theodoricus de Nieheim, , reformátoři , koncil kostnický (1414-1418) , jednotlivci (církevní dějiny) , and světové dějiny středověku (do r. 1492)
Language:
Dutch
Rights:
unknown
Publisher:
Katholieke Universiteit Leuven Campus Kortrijk, Hogeschool Gent
Type:
corpus
Language:
Dutch , English , and French
Description:
Parallel corpus, with Dutch as first language, 10 M words (under construction). DPC is a STEVIN-project.
Rights:
Not specified
Publisher:
Radboud University Nijmegen , Max Planck Institute for Psycholinguistics , Meertens Institute KNAW The Netherlands , and Babylon Centre for Studies of Multilingualism in the Multicultural Society
Type:
corpus
Language:
Arabic , Dutch , and Turkish
Description:
Audio recordings, transcripts,
Rights:
Not specified
Publisher:
Meertens Institute KNAW The Netherlands
Type:
corpus
Language:
Dutch
Description:
The Dynamic Syntactic Atlas of the Dutch dialects (DynaSAND) is an on-line tool for dialect syntax research. DynaSAND consists of a database, a search engine, a cartographic component and a bibliography.
Rights:
Not specified
Creator:
Goeje, Michael Jan de,
Type:
text and studie
Subject:
Dějiny Evropy , Ibrahim ibn Jákúb, , Al-Bekri, , kronikáři arabští , kodexy , Slované , kroniky arabské , edice , paleografie , filologie , vztahy arabsko-slovanské , dějepisectví, historické vědy, historici , světové dějiny středověku (do r. 1492) , and rukopisy
Language:
Dutch
Description:
Overgedrukt uit de Verslagen en Medeelingen der Koninklijke Akademie van Wetenschappen, Afdeling Letterkunde, 2. de Reeks, Deel IX.
Rights:
unknown
Creator:
Kersseboom, Willem,
Type:
text and prameny
Subject:
Demografie. Populace , demografie historická , demografové nizozemští , and dějiny historické demografie a statistiky, jednotlivci
Language:
French and Dutch
Rights:
unknown
Creator:
Presser, Jacob,
Type:
text and studie
Subject:
Germánské literatury , dějiny evropské , dějiny společnosti , společenská struktura , and přehledná zpracování světových dějin (chronologicky)
Language:
Dutch
Rights:
unknown
Creator:
Salfellner, Harald,
Publisher:
Vitalis,
Subject:
Kafka, Franz, , spisovatelé německojazyční , Židé , biografie , české země 1848-1918 , Československo 1918-1938 , and literatura, spisovatelé
Language:
Dutch
Rights:
unknown