Skip to search
Skip to main content
Skip to first result
Search
Search Results
Type:
corpus
Language:
Arabic , Danish , Dutch , English , German , Modern Greek (1453-) , Italian , Japanese , Korean , Portuguese , Russian , Spanish , and Turkish
Description:
Large set of subtitles available for download in multiple languages. Can be used as parallel corpus.
Rights:
Not specified
Publisher:
Max Planck Institute for Psycholinguistics
Type:
lexicalConceptualResource
Language:
Dutch , English , and German
Rights:
Not specified
Creator:
Rüdiger, Jan Oliver
Publisher:
Jan Oliver Rüdiger
Type:
tool and toolService
Subject:
Corpus Linguisitics , NLP , conll , tei , XML , nlp , Natural Language Processing , linguistics , Linguistics , Computational Linguistics , corpus processing , tagger , POS tagger , lemmatization , text cleaning , CommonCrawl , epub , JSON , Twitter , Pandoc , Wikipedia , digital data , DTA , DSpin , MySQL , ElasticSearch , TextGrid , text corpora , TigerXML , and WeblichtXML
Language:
German , English , French , Italian , Dutch , Spanish , Polish , Arabic , Chinese , and Portuguese
Description:
Software for corpus linguists and text/data mining enthusiasts. The CorpusExplorer combines over 45 interactive visualizations under a user-friendly interface. Routine tasks such as text acquisition, cleaning or tagging are completely automated. The simple interface supports the use in university teaching and leads users/students to fast and substantial results. The CorpusExplorer is open for many standards (XML, CSV, JSON, R, etc.) and also offers its own software development kit (SDK).
Source code available at https://github.com/notesjor/corpusexplorer2.0
Rights:
Not specified
Publisher:
Center for Sprogteknologi, University of Copenhagen
Type:
toolService
Language:
Danish , Dutch , English , German , Modern Greek (1453-) , Icelandic , Norwegian , Russian , Slovenian , and Swedish
Description:
1) Fully automatic rule based lemmatization of inflected languages 2) Fully automatic training of lemmatization rules based on full form-lemma list
Rights:
Not specified
Publisher:
Joint Research Centre of the EU
Type:
corpus
Language:
Bulgarian , Czech , Danish , Dutch , English , Estonian , Finnish , French , German , Modern Greek (1453-) , Hungarian , Italian , Latvian , Maltese , Norwegian , Polish , Portuguese , Romanian , Slovak , Slovenian , Spanish , and Swedish
Description:
The largest parallel corpus, contains EU law, the Acquis Communautaire in 22 languages.
Rights:
Not specified
Publisher:
Max Planck Institute for Psycholinguistics
Type:
corpus
Language:
Dutch , English , French , and German
Description:
Language Acquisition corpus
Rights:
Not specified
Type:
corpus
Language:
Dutch , English , French , German , and Swedish
Description:
Corpus of the ESF Foreign Language Speakers project; almost perfect structurefor IEI; completely metadata described; lots of annotated audio recordings containing multimodal interaction;
Rights:
Not specified
Publisher:
Max Planck Institute for Psycholinguistics
Type:
corpus
Language:
Dutch , German , English , and French
Description:
Language Acquisition corpus
Rights:
Not specified
Creator:
Straková, Jana and Straka, Milan
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
service and toolService
Subject:
named entity recognition , NameTag , and WeblichtXML
Language:
Czech , German , English , Spanish , and Dutch
Description:
Metadata description of nametag (http://hdl.handle.net/11234/1-3633, https://lindat.mff.cuni.cz/services/nametag/) provided for weblicht.
Rights:
Not specified
Type:
corpus
Language:
Danish , Dutch , English , Finnish , French , German , Italian , Latin , Portuguese , Russian , Spanish , Swedish , and Telugu
Description:
Possibility to download or to browse free electronic books; Angebot: Download von und Online-Zugang zu frei verfügbaren E-Books; deutschsprachige Literatur stellt nur einen Teilbereich der verfügbaren E-Books dar
Rights:
Not specified