Type: lexicalConceptualResource - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Type lexicalConceptualResource

131. MorfFlex SK 170914

Creator:: Hajič, Jan and Hric, Jan
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text, computationalLexicon, and lexicalConceptualResource
Subject:: Slovak and morphological dictionary
Language:: Slovak
Description:: Slovak morphological dictionary modeled after the Czech one. It consists of (word form, lemma, POS tag) triples, reusing the Czech morphological system for POS tags and lemma descriptions.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

132. MorfoCzech

Creator:: Pelegrinová, Kateřina, Elšík, Viktor, Čech, Radek, and Mačutek, Ján
Publisher:: University of Ostrava
Type:: text, lexicon, and lexicalConceptualResource
Subject:: word segmentation, morphology, and morphological dictionary
Language:: Czech
Description:: A dictionary of morphologically segmented word forms in Czech. Rules of manual segmentation are described in Pelegrinová, K., Mačutek, J., Čech, R. (2021). The Menzerath-Altmann law as the relation between lengths of words and morphemes in Czech. Jazykovedný časopis, 72, 405-414. The dictionary is based on short stories, fairy tales, letters and studies written by Karel Čapek.
Rights:: Creative Commons - Attribution 4.0 International (CC BY 4.0), http://creativecommons.org/licenses/by/4.0/, and PUB

133. MorfoCzech 1.1

Creator:: Pelegrinová, Kateřina, Elšík, Viktor, Čech, Radek, and Mačutek, Ján
Publisher:: University of Ostrava
Type:: text, lexicon, and lexicalConceptualResource
Subject:: word segmentation, morphology, and morphological dictionary
Language:: Czech
Description:: A dictionary of morphologically segmented word forms in Czech. Rules of manual segmentation are described in Pelegrinová, K., Mačutek, J., Čech, R. (2021). The Menzerath-Altmann law as the relation between lengths of words and morphemes in Czech. Jazykovedný časopis, 72, 405-414. The dictionary is based on short stories, fairy tales, letters and studies written by Karel Čapek.
Rights:: Creative Commons - Attribution 4.0 International (CC BY 4.0), http://creativecommons.org/licenses/by/4.0/, and PUB

134. Morphdb.hu

Publisher:: Budapest University of Technology and Economics Media Research (BME MOKK)
Type:: lexicalConceptualResource
Language:: Hungarian
Description:: 100,000 lemmas
Rights:: Not specified

135. Motion Encoding Lexicalization Patterns: Portuguese and English Learners

Creator:: Costa-Silva, Jean
Publisher:: University of Georgia
Type:: text, other, and lexicalConceptualResource
Subject:: motion encoding, language acquisition, Portuguese, English, and lexicalization patterns
Language:: English
Description:: General Information: Data collector: Jean Costa Silva (University of Georgia) Date of collection: September-December 2022 Manner of collection: Online questionnaire via Qualtrics Funding: No
Rights:: Creative Commons - Attribution 4.0 International (CC BY 4.0), http://creativecommons.org/licenses/by/4.0/, and PUB

136. Multilingual Central Repository

Publisher:: Centro de Tecnologías y Aplicaciones del Lenguaje y del Habla (TALP)
Type:: lexicalConceptualResource
Subject:: lexical database
Language:: Basque, Catalan, English, Galician, and Spanish
Description:: Multilingual lexical database that follows the model proposed by the EuroWordNet project. The MCR integrates into the same EuroWordNet framework wordnets from five different languages (together with four English WordNet versions). It also integrates WordNet Domains and new versions of the Base Concepts and Top Concept Ontology. Overall, it contains 1,642,389 semantic relations between synsets, most of them acquired by automatic means. Information contained: semantics, synonyms, antonyms, definition, equivalents, example of use, morphology.
Rights:: Not specified

137. Multilingual static embeddings for Verbal Multiword Expressions trained on PARSEME raw corpora

Creator:: Estève, Louis Clément, Savary, Agata, and Lavergne, Thomas
Publisher:: Université Paris-Saclay, CNRS, Laboratoire Interdisciplinaire des Sciences du Numérique
Type:: text, computationalLexicon, and lexicalConceptualResource
Subject:: verbal multiword expressions, word embeddings, and word2vec
Language:: German, Modern Greek (1453-), Basque, French, Irish, Hebrew, Hindi, Italian, Polish, Portuguese, Romanian, Swedish, Turkish, and Chinese
Description:: This resource is a set of 14 vector spaces for single words and Verbal Multiword Expressions (VMWEs) in different languages (German, Greek, Basque, French, Irish, Hebrew, Hindi, Italian, Polish, Brazilian Portuguese, Romanian, Swedish, Turkish, Chinese). They were trained with the Word2Vec algorithm, in its skip-gram version, on PARSEME raw corpora automatically annotated for morpho-syntax (http://hdl.handle.net/11234/1-3367). These corpora were annotated by Seen2Seen, a rule-based VMWE identifier, one of the leading tools of the PARSEME shared task version 1.2. VMWE tokens were merged into single tokens. The format of the vector space files is that of the original Word2Vec implementation by Mikolov et al. (2013), i.e. a binary format. For compression, bzip2 was used.
Rights:: PARSEME Shared Task Raw Corpus Data (v. 1.2) Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-mwe-1.2-raw, and PUB

138. NAFIS Arabic Stemming Gold Standard Corpus

Creator:: Namly, Driss
Publisher:: Ibtikarat team
Type:: text, wordList, and lexicalConceptualResource
Subject:: corpus, stemming;, and Gold Standard Corpus
Language:: Arabic
Description:: Normalized Arabic Fragments for Inestimable Stemming (NAFIS) is an Arabic stemming gold standard corpus composed by a collection of texts, selected to be representative of Arabic stemming tasks and manually annotated.
Rights:: Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), http://creativecommons.org/licenses/by-nc/4.0/, and PUB

139. Narrangansett dictionary

Format:: application/pdf
Type:: lexicalConceptualResource
Description:: The Narrangansett corpus contains a cultural linguistic dictionary and grammatical information on the Narrangansett language, an extinct language of the USA.
Rights:: Not specified

140. Neologismos económicos en las lenguas románicas a través de la prensa

Publisher:: Institut Universitari de Lingüística Aplicada, Universitat Pompeu Fabra
Type:: lexicalConceptualResource
Subject:: terminology database
Language:: Catalan, French, Galician, Italian, Portuguese, Romanian, and Spanish
Description:: Multilingual terminological resource containing 3.875 entries from the Economics, Finance and Banking domains.
Rights:: Not specified

« Previous
Next »
1
2
…
10
11
12
13
14
15
16
17
18
…
22
23

131. MorfFlex SK 170914

132. MorfoCzech

133. MorfoCzech 1.1

134. Morphdb.hu

135. Motion Encoding Lexicalization Patterns: Portuguese and English Learners

136. Multilingual Central Repository

137. Multilingual static embeddings for Verbal Multiword Expressions trained on PARSEME raw corpora

138. NAFIS Arabic Stemming Gold Standard Corpus

139. Narrangansett dictionary

140. Neologismos económicos en las lenguas románicas a través de la prensa

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Coverage

Show values starting with

Creator

Show values starting with

Format

Language

Show values starting with

Publisher

Show values starting with

Rights

Show values starting with

Subject

Show values starting with

Type

Date

Original context has metadata only

Harvested from