Language: Arabic / Rights: PUB and http://creativecommons.org/licenses/by-nc-sa/4.0/ / Subject: lemmatization - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Language Arabic Rights PUB Rights http://creativecommons.org/licenses/by-nc-sa/4.0/ Subject lemmatization

1. CALEM (Comprehensive Arabic LEMmas)

Creator:: Namly, Driss, Bouzoubaa, Karim, and El Jihad, Abdelhamid
Publisher:: ALELM
Type:: text, lexicon, and lexicalConceptualResource
Subject:: lexicon, lemmatization, and stemming;
Language:: Arabic
Description:: Comprehensive Arabic LEMmas is a lexicon covering a large list of Arabic lemmas and their corresponding inflected word forms (stems) with details (POS + Root). Each lexical entry represents a lemma followed by all its possible stems and each stem is enriched by its morphological features especially the root and the POS. It is composed of 164,845 lemmas representing 7,200,918 stems, detailed as follow: 757 Arabic particles 2,464,631 verbal stems 4,735,587 nominal stems The lexicon is provided as an LMF conformant XML-based file in UTF8 encoding, which represents about 1,22 Gb of data. Citation: – Namly Driss, Karim Bouzoubaa, Abdelhamid El Jihad, and Si Lhoussain Aouragh. “Improving Arabic Lemmatization Through a Lemmas Database and a Machine-Learning Technique.” In Recent Advances in NLP: The Case of Arabic Language, pp. 81-100. Springer, Cham, 2020.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

2. Universal Dependencies 2.10 models for UDPipe 2 (2022-07-11)

Creator:: Straka, Milan
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: tool and toolService
Subject:: tokenizer, POS tagger, lemmatization, tagger, parser, and dependency parser
Language:: Afrikaans, Arabic, Belarusian, Bulgarian, Catalan, Czech, Church Slavic, Coptic, Welsh, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Faroese, Persian, Finnish, French, Old French (842-ca. 1400), Scottish Gaelic, Irish, Galician, Gothic, Ancient Greek (to 1453), Ancient Hebrew, Hebrew, Hindi, Croatian, Hungarian, Armenian, Western Armenian, Indonesian, Icelandic, Italian, Japanese, Korean, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Maltese, Dutch, Norwegian Nynorsk, Norwegian Bokmål, Old Russian, Nigerian Pidgin, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Telugu, Turkish, Uighur, Ukrainian, Urdu, Vietnamese, Gambian Wolof, Wolof, and Chinese
Description:: Tokenizer, POS Tagger, Lemmatizer and Parser models for 123 treebanks of 69 languages of Universal Depenencies 2.10 Treebanks, created solely using UD 2.10 data (https://hdl.handle.net/11234/1-4758). The model documentation including performance can be found at https://ufal.mff.cuni.cz/udpipe/2/models#universal_dependencies_210_models . To use these models, you need UDPipe version 2.0, which you can download from https://ufal.mff.cuni.cz/udpipe/2 .
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

3. Universal Dependencies 2.12 models for UDPipe 2 (2023-07-17)

Creator:: Straka, Milan
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: tool and toolService
Subject:: tokenizer, POS tagger, lemmatization, tagger, parser, and dependency parser
Language:: Afrikaans, Arabic, Belarusian, Bulgarian, Catalan, Czech, Church Slavic, Coptic, Welsh, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Faroese, Persian, Finnish, French, Old French (842-ca. 1400), Scottish Gaelic, Irish, Galician, Gothic, Ancient Greek (to 1453), Ancient Hebrew, Hebrew, Hindi, Croatian, Hungarian, Armenian, Western Armenian, Indonesian, Icelandic, Italian, Japanese, Korean, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Maltese, Dutch, Norwegian Nynorsk, Norwegian Bokmål, Old Russian, Nigerian Pidgin, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Telugu, Turkish, Uighur, Ukrainian, Urdu, Vietnamese, Gambian Wolof, Wolof, Chinese, Norwegian, Erzya, and Manx
Description:: Tokenizer, POS Tagger, Lemmatizer and Parser models for 131 treebanks of 72 languages of Universal Depenencies 2.12 Treebanks, created solely using UD 2.12 data (https://hdl.handle.net/11234/1-5150). The model documentation including performance can be found at https://ufal.mff.cuni.cz/udpipe/2/models#universal_dependencies_212_models . To use these models, you need UDPipe version 2.0, which you can download from https://ufal.mff.cuni.cz/udpipe/2 .
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

4. Universal Dependencies 2.4 Models for UDPipe (2019-05-31)

Creator:: Straka, Milan and Straková, Jana
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: tool and toolService
Subject:: tokenizer, POS tagger, lemmatization, tagger, parser, and dependency parser
Language:: Czech, Afrikaans, Arabic, Belarusian, Bulgarian, Catalan, Church Slavic, Coptic, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Persian, Finnish, French, Old French (842-ca. 1400), Irish, Galician, Gothic, Ancient Greek (to 1453), Hebrew, Hindi, Croatian, Hungarian, Armenian, Indonesian, Italian, Japanese, Kazakh, Korean, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Maltese, Dutch, Norwegian Nynorsk, Norwegian Bokmål, Old Russian, Polish, Portuguese, Romanian, Russian, Sanskrit, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Telugu, Turkish, Uighur, Ukrainian, Urdu, Vietnamese, Gambian Wolof, Wolof, and Chinese
Description:: Tokenizer, POS Tagger, Lemmatizer and Parser models for 90 treebanks of 60 languages of Universal Depenencies 2.4 Treebanks, created solely using UD 2.4 data (http://hdl.handle.net/11234/1-2988). The model documentation including performance can be found at http://ufal.mff.cuni.cz/udpipe/models#universal_dependencies_24_models . To use these models, you need UDPipe binary version at least 1.2, which you can download from http://ufal.mff.cuni.cz/udpipe . In addition to models itself, all additional data and value of hyperparameters used for training are available in the second archive, allowing reproducible training.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

5. Universal Dependencies 2.5 Models for UDPipe (2019-12-06)

Creator:: Straka, Milan and Straková, Jana
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: tool and toolService
Subject:: tokenizer, POS tagger, lemmatization, tagger, parser, and dependency parser
Language:: Czech, Afrikaans, Arabic, Belarusian, Bulgarian, Catalan, Church Slavic, Coptic, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Persian, Finnish, French, Old French (842-ca. 1400), Irish, Galician, Gothic, Ancient Greek (to 1453), Hebrew, Hindi, Croatian, Hungarian, Armenian, Indonesian, Italian, Japanese, Kazakh, Korean, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Maltese, Dutch, Norwegian Nynorsk, Norwegian Bokmål, Old Russian, Polish, Portuguese, Romanian, Russian, Sanskrit, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Telugu, Turkish, Uighur, Ukrainian, Urdu, Vietnamese, Gambian Wolof, Wolof, Chinese, and Scottish Gaelic
Description:: Tokenizer, POS Tagger, Lemmatizer and Parser models for 94 treebanks of 61 languages of Universal Depenencies 2.5 Treebanks, created solely using UD 2.5 data (http://hdl.handle.net/11234/1-3105). The model documentation including performance can be found at http://ufal.mff.cuni.cz/udpipe/models#universal_dependencies_25_models . To use these models, you need UDPipe binary version at least 1.2, which you can download from http://ufal.mff.cuni.cz/udpipe . In addition to models itself, all additional data and value of hyperparameters used for training are available in the second archive, allowing reproducible training.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

6. Universal Dependencies 2.6 models for UDPipe 2 (2020-08-31)

Creator:: Straka, Milan
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: tool and toolService
Subject:: tokenizer, POS tagger, lemmatization, tagger, parser, and dependency parser
Language:: Afrikaans, Arabic, Armenian, Belarusian, Bulgarian, Catalan, Czech, Church Slavic, Coptic, Welsh, Danish, German, Modern Greek (1453-), English, Estonian, Basque, Persian, Finnish, French, Old French (842-ca. 1400), Scottish Gaelic, Irish, Galician, Gothic, Ancient Greek (to 1453), Hebrew, Hindi, Croatian, Hungarian, Indonesian, Italian, Japanese, Korean, Latin, Latvian, Lithuanian, Literary Chinese, Marathi, Maltese, Dutch, Norwegian Nynorsk, Norwegian Bokmål, Old Russian, Nigerian Pidgin, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Northern Sami, Spanish, Serbian, Swedish, Tamil, Telugu, Turkish, Uighur, Ukrainian, Urdu, Vietnamese, Gambian Wolof, Wolof, and Chinese
Description:: Tokenizer, POS Tagger, Lemmatizer and Parser models for 99 treebanks of 63 languages of Universal Depenencies 2.6 Treebanks, created solely using UD 2.6 data (https://hdl.handle.net/11234/1-3226). The model documentation including performance can be found at https://ufal.mff.cuni.cz/udpipe/2/models#universal_dependencies_26_models . To use these models, you need UDPipe version 2.0, which you can download from https://ufal.mff.cuni.cz/udpipe/2 .
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB