Skip to search
Skip to main content
Skip to first result
Search
Search Results
Creator:
Ansari, Ebrahim , Žabokrtský, Zdeněk , Haghdoost, Hamid , and Nikravesh, Mahshid
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text , lexicon , and lexicalConceptualResource
Subject:
morphological analysis, and lemmatization
Language:
Persian
Description:
This dataset includes 45300 Persian word forms which are manually segmented into sequences of morphemes.
Rights:
Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) , http://creativecommons.org/licenses/by-nc-sa/4.0/ , and PUB
Creator:
Žabokrtský, Zdeněk , Bafna, Nyati , Bodnár, Jan , Kyjánek, Lukáš , Svoboda, Emil , Ševčíková, Magda , Vidra, Jonáš , Angle, Sachi , Ansari, Ebrahim , Arkhangelskiy, Timofey , Batsuren, Khuyagbaatar , Bella, Gábor , Bertinetto, Pier Marco , Bonami, Olivier , Celata, Chiara , Daniel, Michael , Fedorenko, Alexei , Filko, Matea , Giunchiglia, Fausto , Haghdoost, Hamid , Hathout, Nabil , Khomchenkova, Irina , Khurshudyan, Victoria , Levonian, Dmitri , Litta, Eleonora , Medvedeva, Maria , Muralikrishna, S. N. , Namer, Fiammetta , Nikravesh, Mahshid , Padó, Sebastian , Passarotti, Marco , Plungian, Vladimir , Polyakov, Alexey , Potapov, Mihail , Pruthwik, Mishra , Rao B, Ashwath , Rubakov, Sergei , Samar, Husain , Sharma, Dipti Misra , Šnajder, Jan , Šojat, Krešimir , Štefanec, Vanja , Talamo, Luigi , Tribout, Delphine , Vodolazsky, Daniil , Vydrin, Arseniy , Zakirova, Aigul , and Zeller, Britta
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text , lexicon , and lexicalConceptualResource
Subject:
universal segmentations , morphological segmentation , word segmentation , segmentation , morphology , morphemes , morphological dictionary , unisegments , morph , and multilingual
Language:
Czech , Catalan , German , English , Persian , Finnish , French , Serbo-Croatian , Croatian , Hungarian , Italian , Komi-Zyrian , Latin , Moksha , Mari (Russia) , Mongolian , Erzya , Polish , Portuguese , Russian , Spanish , Swedish , Tajik , Udmurt , Armenian , Bengali , Hindi , Malayalam , Marathi , and Kannada
Description:
Universal Segmentations (UniSegments) is a collection of lexical resources capturing morphological segmentations harmonised into a cross-linguistically consistent annotation scheme for many languages. The annotation scheme consists of simple tab-separated columns that stores a word and its morphological segmentations, including pieces of information about the word and the segmented units, e.g., part-of-speech categories, type of morphs/morphemes etc. The current public version of the collection contains 38 harmonised segmentation datasets covering 30 different languages.
Rights:
Universal Segmentations 1.0 License Terms , https://lindat.mff.cuni.cz/repository/xmlui/page/licence-unisegs-1.0 , and PUB