Skip to search
Skip to main content
Skip to first result
Search
Search Results
Type:
corpus
Language:
Arabic , Danish , Dutch , English , German , Modern Greek (1453-) , Italian , Japanese , Korean , Portuguese , Russian , Spanish , and Turkish
Description:
Large set of subtitles available for download in multiple languages. Can be used as parallel corpus.
Rights:
Not specified
Publisher:
Max Planck Institute for Psycholinguistics
Type:
lexicalConceptualResource
Language:
Dutch , English , and German
Rights:
Not specified
Creator:
Rüdiger, Jan Oliver
Publisher:
Jan Oliver Rüdiger
Type:
tool and toolService
Subject:
Corpus Linguisitics , NLP , conll , tei , XML , nlp , Natural Language Processing , linguistics , Linguistics , Computational Linguistics , corpus processing , tagger , POS tagger , lemmatization , text cleaning , CommonCrawl , epub , JSON , Twitter , Pandoc , Wikipedia , digital data , DTA , DSpin , MySQL , ElasticSearch , TextGrid , text corpora , TigerXML , and WeblichtXML
Language:
German , English , French , Italian , Dutch , Spanish , Polish , Arabic , Chinese , and Portuguese
Description:
Software for corpus linguists and text/data mining enthusiasts. The CorpusExplorer combines over 45 interactive visualizations under a user-friendly interface. Routine tasks such as text acquisition, cleaning or tagging are completely automated. The simple interface supports the use in university teaching and leads users/students to fast and substantial results. The CorpusExplorer is open for many standards (XML, CSV, JSON, R, etc.) and also offers its own software development kit (SDK).
Source code available at https://github.com/notesjor/corpusexplorer2.0
Rights:
Not specified
Publisher:
Center for Sprogteknologi, University of Copenhagen
Type:
toolService
Language:
Danish , Dutch , English , German , Modern Greek (1453-) , Icelandic , Norwegian , Russian , Slovenian , and Swedish
Description:
1) Fully automatic rule based lemmatization of inflected languages 2) Fully automatic training of lemmatization rules based on full form-lemma list
Rights:
Not specified
Publisher:
Max Planck Institute for Psycholinguistics
Type:
corpus
Language:
Dutch , English , French , and German
Description:
Language Acquisition corpus
Rights:
Not specified
Publisher:
ILK, Tilburg University and CNTS - Language Technology Group,
University of Antwerp
Type:
toolService
Language:
Dutch and English
Description:
MBSP is a set of linguistic tools based on the TiMBL and MBT memory based learning applications developed at CNTS and ILK. It provides tools for Part of Speech tagging, Chunking, Lemmatizing, Relation Finding, Named Entity Recognition, and (for medical language) Semantic tagging.
Rights:
Not specified
Type:
corpus
Language:
Dutch , English , French , German , and Swedish
Description:
Corpus of the ESF Foreign Language Speakers project; almost perfect structurefor IEI; completely metadata described; lots of annotated audio recordings containing multimodal interaction;
Rights:
Not specified
Publisher:
Max Planck Institute for Psycholinguistics
Type:
corpus
Language:
Dutch , German , English , and French
Description:
Language Acquisition corpus
Rights:
Not specified
Creator:
Straková, Jana and Straka, Milan
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
service and toolService
Subject:
named entity recognition , NameTag , and WeblichtXML
Language:
Czech , German , English , Spanish , and Dutch
Description:
Metadata description of nametag (http://hdl.handle.net/11234/1-3633, https://lindat.mff.cuni.cz/services/nametag/) provided for weblicht.
Rights:
Not specified
Publisher:
Universität Bamberg, World Language Documentation Centre
Format:
application/octet-stream
Type:
lexicalConceptualResource
Language:
Afrikaans , Arabic , Basque , Bulgarian , Catalan , Chinese , Czech , Danish , Dutch , English , Esperanto , Estonian , Finnish , French , Galician , Georgian , Modern Greek (1453-) , Hebrew , Hungarian , Icelandic , Indonesian , Interlingua (International Auxiliary Language Association) , Irish , Italian , Japanese , Khmer , Norwegian , Polish , Portuguese , Romanian , Russian , Serbian , Slovak , Spanish , Swedish , Turkish , Ukrainian , and Welsh
Rights:
GFDL or CC and http://www.omegawiki.org/Licensing