Skip to search
Skip to main content
Skip to first result
Search
Search Results
Type:
corpus
Language:
Arabic , Danish , Dutch , English , German , Modern Greek (1453-) , Italian , Japanese , Korean , Portuguese , Russian , Spanish , and Turkish
Description:
Large set of subtitles available for download in multiple languages. Can be used as parallel corpus.
Rights:
Not specified
Publisher:
Center for Sprogteknologi, University of Copenhagen
Type:
toolService
Language:
Danish , Dutch , English , German , Modern Greek (1453-) , Icelandic , Norwegian , Russian , Slovenian , and Swedish
Description:
1) Fully automatic rule based lemmatization of inflected languages 2) Fully automatic training of lemmatization rules based on full form-lemma list
Rights:
Not specified
Creator:
Bojar, Ondřej , Cífka, Ondřej , Pecina, Pavel , and Tamchyna, Aleš
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
toolService and tool
Subject:
machine translation , web service , and demo
Language:
Czech , English , Russian , Ukrainian , French , and German
Description:
An interactive web demo of selected ÚFAL MT systems. and FP7-ICT-2011-7-288487 MosesCore
Rights:
Not specified
Publisher:
University of Tampere
Format:
application/octet-stream
Type:
corpus
Subject:
parallel corpus and multilingual
Language:
English , German , Russian , and Swedish
Description:
International conventions and treaties arranged as a paralell corpus aligned on paragraph level
Rights:
Not specified
Type:
corpus
Language:
Danish , Dutch , English , Finnish , French , German , Italian , Latin , Portuguese , Russian , Spanish , Swedish , and Telugu
Description:
Possibility to download or to browse free electronic books; Angebot: Download von und Online-Zugang zu frei verfügbaren E-Books; deutschsprachige Literatur stellt nur einen Teilbereich der verfügbaren E-Books dar
Rights:
Not specified
Publisher:
Department of Literature, Area Studies and European Languages, University of Oslo and Department of Linguistics and Nordic Studies, University of Oslo
Type:
corpus
Language:
English , Norwegian , and Russian
Description:
The RuN corpus is a parallel corpus consisting of Norwegian, Russian and English texts. The texts are aligned at the sentence level and have been tagged for grammatical information at the word level.
Rights:
Not specified
Type:
corpus
Language:
Czech , Danish , Dutch , English , Finnish , French , German , Hungarian , Italian , Polish , Portuguese , Russian , Spanish , Swedish , Turkish , Chinese , Hebrew , Japanese , Korean , and Thai
Description:
28 speech databases containing broadband recordings from 550 adults and 50 children per language. Contains interesting phonetically rich material. All orthographically transcribed. Speaker information included for gender, age, accent. Including pronunciation lexicon.
Rights:
Not specified
Publisher:
Centre for Applied Language Studies, University of Jyväskylä
Type:
corpus
Language:
English , Finnish , French , German , Italian , Russian , Spanish , and Swedish
Description:
The NC test results, background information, speaking and writing performances in 9 foreign / second languages. A web-based data base (html files).
Rights:
Not specified
Publisher:
University of Stuttgart
Type:
toolService
Subject:
POS tagger
Language:
Bulgarian , Dutch , English , French , German , Modern Greek (1453-) , Italian , Portuguese , Russian , Spanish , and Swahili (macrolanguage)
Description:
A part-of-speech tagger and lemmatizer for several languages.
Rights:
Not specified
Publisher:
University of Leipzig
Type:
corpus
Language:
Afrikaans , Albanian , Bulgarian , Catalan , Chinese , Croatian , Czech , Danish , Dutch , English , Esperanto , Estonian , Finnish , French , German , Hungarian , Icelandic , Indonesian , Italian , Japanese , Korean , Latin , Latvian , Lithuanian , Malay (macrolanguage) , Norwegian , Occitan (post 1500) , Romanian , Russian , Slovak , Slovenian , Spanish , Sundanese , Swedish , Tagalog , Turkish , Vietnamese , and Welsh
Description:
Collected from newspaper texts, webcrawling, etc.: words (+frequency), cooccurrences (+graph), left/right neighbours, example sentences
Rights:
Not specified