Creator: Zeman, Daniel / Language: Finnish - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Creator Zeman, Daniel Language Finnish

11. Deltacorpus 1.1

Creator:: Mareček, David, Yu, Zhiwei, Zeman, Daniel, and Žabokrtský, Zdeněk
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: part of speech, tagging, semi-supervised, and cross-language
Language:: Belarusian, Bosnian, Bulgarian, Czech, Serbo-Croatian, Croatian, Upper Sorbian, Macedonian, Polish, Russian, Slovak, Slovenian, Serbian, Ukrainian, Latvian, Lithuanian, Afrikaans, Danish, German, English, Faroese, Western Frisian, Swiss German, Icelandic, Limburgan, Luxembourgish, Low German, Dutch, Norwegian Nynorsk, Norwegian, Scots, Swedish, Yiddish, Aragonese, Asturian, Catalan, French, Galician, Haitian, Italian, Latin, Lombard, Neapolitan, Piemontese, Portuguese, Romanian, Spanish, Venetian, Walloon, Breton, Welsh, Scottish Gaelic, Irish, Modern Greek (1453-), Armenian, Albanian, Dimli (individual language), Persian, Gilaki, Kurdish, Tajik, Bengali, Bishnupriya, Gujarati, Fiji Hindi, Hindi, Marathi, Nepali (macrolanguage), Urdu, Amharic, Arabic, Egyptian Arabic, Hebrew, Estonian, Finnish, Hungarian, Basque, Georgian, Chuvash, Azerbaijani, Turkish, Uzbek, Kazakh, Tatar, Yakut, Korean, Mongolian, Telugu, Kannada, Malayalam, Tamil, Newari, Vietnamese, Indonesian, Javanese, Malagasy, Maori, Malay (macrolanguage), Pampanga, Sundanese, Tagalog, Waray (Philippines), Swahili (macrolanguage), Esperanto, Ido, Interlingua (International Auxiliary Language Association), and Volapük
Description:: Texts in 107 languages from the W2C corpus (http://hdl.handle.net/11858/00-097C-0000-0022-6133-9), first 1,000,000 tokens per language, tagged by the delexicalized tagger described in Yu et al. (2016, LREC, Portorož, Slovenia). Changes in version 1.1: 1. Universal Dependencies tagset instead of the older and smaller Google Universal POS tagset. 2. SVM classifier trained on Universal Dependencies 1.2 instead of HamleDT 2.0. 3. Balto-Slavic languages, Germanic languages and Romance languages were tagged by classifier trained only on the respective group of languages. Other languages were tagged by a classifier trained on all available languages. The "c7" combination from version 1.0 is no longer used.
Rights:: Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), http://creativecommons.org/licenses/by-sa/4.0/, and PUB

12. HamleDT 2.0

Creator:: Zeman, Daniel, Mareček, David, Mašek, Jan, Popel, Martin, Ramasamy, Loganathan, Rosa, Rudolf, Štěpánek, Jan, and Žabokrtský, Zdeněk
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: treebank, Stanford dependencies, Prague dependencies, harmonization, common annotation style, and Interset
Language:: Arabic, Bulgarian, Bengali, Catalan, Czech, Danish, German, Modern Greek (1453-), English, Spanish, Estonian, Basque, Persian, Finnish, Ancient Greek (to 1453), Hindi, Hungarian, Italian, Japanese, Latin, Dutch, Portuguese, Romanian, Russian, Slovak, Slovenian, Swedish, Tamil, Telugu, and Turkish
Description:: HamleDT 2.0 is a collection of 30 existing treebanks harmonized into a common annotation style, the Prague Dependencies, and further transformed into Stanford Dependencies, a treebank annotation style that became popular recently. We use the newest basic Universal Stanford Dependencies, without added language-specific subtypes.
Rights:: HamleDT 2.0 Licence Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-hamledt-2.0, and ACA

13. HamleDT 3.0

Creator:: Zeman, Daniel, Mareček, David, Mašek, Jan, Popel, Martin, Ramasamy, Loganathan, Rosa, Rudolf, Štěpánek, Jan, and Žabokrtský, Zdeněk
Publisher:: Charles University
Type:: text and corpus
Subject:: annotated corpus, morphology, syntax, dependency, treebank, harmonized annotation, and common annotation style
Language:: Arabic, Basque, Bengali, Bulgarian, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Modern Greek (1453-), Ancient Greek (to 1453), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Persian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Tamil, Telugu, and Turkish
Description:: HamleDT (HArmonized Multi-LanguagE Dependency Treebank) is a compilation of existing dependency treebanks (or dependency conversions of other treebanks), transformed so that they all conform to the same annotation style. This version uses Universal Dependencies as the common annotation style. Update (November 1017): for a current collection of harmonized dependency treebanks, we recommend using the Universal Dependencies (UD). All of the corpora that are distributed in HamleDT in full are also part of the UD project; only some corpora from the Patch group (where HamleDT provides only the harmonizing scripts but not the full corpus data) are available in HamleDT but not in UD.
Rights:: HamleDT 3.0 License Terms, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-hamledt-3.0, and PUB

14. IWPT 2020 Shared Task Data and System Outputs

Creator:: Zeman, Daniel, Bouma, Gosse, and Seddah, Djamé
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, enhanced universal dependencies, shared task, and parsing
Language:: Arabic, Bulgarian, Czech, Dutch, English, Estonian, Finnish, French, Italian, Latvian, Lithuanian, Polish, Russian, Slovak, Swedish, Tamil, and Ukrainian
Description:: This package contains data used in the IWPT 2020 shared task. It contains training, development and test (evaluation) datasets. The data is based on a subset of Universal Dependencies release 2.5 (http://hdl.handle.net/11234/1-3105) but some treebanks contain additional enhanced annotations. Moreover, not all of these additions became part of Universal Dependencies release 2.6 (http://hdl.handle.net/11234/1-3226), which makes the shared task data unique and worth a separate release to enable later comparison with new parsing algorithms. The package also contains a number of Perl and Python scripts that have been used to process the data during preparation and during the shared task. Finally, the package includes the official primary submission of each team participating in the shared task.
Rights:: Licence Universal Dependencies v2.5, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.5, and PUB

15. IWPT 2021 Shared Task Data and System Outputs

Creator:: Zeman, Daniel, Bouma, Gosse, and Seddah, Djamé
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, enhanced universal dependencies, shared task, and parsing
Language:: Arabic, Bulgarian, Czech, Dutch, English, Estonian, Finnish, French, Italian, Latvian, Lithuanian, Polish, Russian, Slovak, Swedish, Tamil, and Ukrainian
Description:: This package contains data used in the IWPT 2021 shared task. It contains training, development and test (evaluation) datasets. The data is based on a subset of Universal Dependencies release 2.7 (http://hdl.handle.net/11234/1-3424) but some treebanks contain additional enhanced annotations. Moreover, not all of these additions became part of Universal Dependencies release 2.8 (http://hdl.handle.net/11234/1-3687), which makes the shared task data unique and worth a separate release to enable later comparison with new parsing algorithms. The package also contains a number of Perl and Python scripts that have been used to process the data during preparation and during the shared task. Finally, the package includes the official primary submission of each team participating in the shared task.
Rights:: Licence Universal Dependencies v2.7, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.7, and PUB

16. Lingua::Interset 2.026

Creator:: Zeman, Daniel
Publisher:: Charles University, Faculty of Mathematics and Physics
Type:: tool and toolService
Subject:: morphology, part of speech, conversion, and tagset
Language:: Arabic, Bulgarian, Bengali, Catalan, Czech, Danish, German, Modern Greek (1453-), English, Spanish, Estonian, Basque, Persian, Finnish, Ancient Greek (to 1453), Hebrew, Hindi, Croatian, Japanese, Multiple languages, and Portuguese
Description:: Lingua::Interset is a universal morphosyntactic feature set to which all tagsets of all corpora/languages can be mapped. Version 2.026 covers 37 different tagsets of 21 languages. Limited support of the older drivers for other languages (which are not included in this package but are available for download elsewhere) is also available; these will be fully ported to Interset 2 in future. Interset is implemented as Perl libraries. It is also available via CPAN.
Rights:: Artistic License (Perl) 1.0, http://opensource.org/licenses/Artistic-Perl-1.0, and PUB

17. Treebanks for Unified Taxonomy of Deep Syntactic Relations

Creator:: Droganova, Kira and Zeman, Daniel
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: treebank and semantic roles
Language:: Czech, Spanish, Catalan, and Finnish
Description:: The datasets described in Droganova, Kira, and Daniel Zeman. "Towards a Unified Taxonomy of Deep Syntactic Relations." Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 2024. Four languages are included in this release. English PropBank is omitted due to its license terms.
Rights:: Licence Universal Dependencies v2.14, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-2.14, and PUB

18. Universal Dependencies 1.0

Creator:: Nivre, Joakim, Bosco, Cristina, Choi, Jinho, de Marneffe, Marie-Catherine, Dozat, Timothy, Farkas, Richárd, Foster, Jennifer, Ginter, Filip, Goldberg, Yoav, Hajič, Jan, Kanerva, Jenna, Laippala, Veronika, Lenci, Alessandro, Lynn, Teresa, Manning, Christopher, McDonald, Ryan, Missilä, Anna, Montemagni, Simonetta, Petrov, Slav, Pyysalo, Sampo, Silveira, Natalia, Simi, Maria, Smith, Aaron, Tsarfaty, Reut, Vincze, Veronika, and Zeman, Daniel
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Czech, German, English, Spanish, Finnish, French, Irish, Italian, Swedish, and Hungarian
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:: Universal Dependencies 1.0 License Set, https://lindat.mff.cuni.cz/repository/xmlui/page/license-ud-1.0, and PUB

19. Universal Dependencies 1.1

Creator:: Agić, Željko, Aranzabe, Maria Jesus, Atutxa, Aitziber, Bosco, Cristina, Choi, Jinho, de Marneffe, Marie-Catherine, Dozat, Timothy, Farkas, Richárd, Foster, Jennifer, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Goldberg, Yoav, Hajič, Jan, Johannsen, Anders Trærup, Kanerva, Jenna, Kuokkala, Juha, Laippala, Veronika, Lenci, Alessandro, Lindén, Krister, Ljubešić, Nikola, Lynn, Teresa, Manning, Christopher, Martínez, Héctor Alonso, McDonald, Ryan, Missilä, Anna, Montemagni, Simonetta, Nivre, Joakim, Nurmi, Hanna, Osenova, Petya, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Prokopidis, Prokopis, Pyysalo, Sampo, Seeker, Wolfgang, Seraji, Mojgan, Silveira, Natalia, Simi, Maria, Simov, Kiril, Smith, Aaron, Tsarfaty, Reut, Vincze, Veronika, and Zeman, Daniel
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency syntax, morphology, harmonized annotation, interset, universal tagset, stanford dependencies, and universal dependencies
Language:: Basque, Bulgarian, Croatian, Czech, Danish, English, Finnish, French, German, Modern Greek (1453-), Hebrew, Hungarian, Indonesian, Irish, Italian, Persian, Spanish, and Swedish
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This is the second release of UD Treebanks, Version 1.1.
Rights:: Licence Universal Dependencies v1.1, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.1, and PUB

20. Universal Dependencies 1.2

Creator:: Nivre, Joakim, Agić, Željko, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Bhat, Riyaz Ahmad, Bosco, Cristina, Bowman, Sam, Celano, Giuseppe G. A., Connor, Miriam, de Marneffe, Marie-Catherine, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Galbraith, Daniel, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Goldberg, Yoav, Gonzales, Berta, Guillaume, Bruno, Hajič, Jan, Haug, Dag, Ion, Radu, Irimia, Elena, Johannsen, Anders, Kanayama, Hiroshi, Kanerva, Jenna, Krek, Simon, Laippala, Veronika, Lenci, Alessandro, Ljubešić, Nikola, Lynn, Teresa, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, Mori, Shunsuke, Nurmi, Hanna, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Prokopidis, Prokopis, Pyysalo, Sampo, Ramasamy, Loganathan, Rosa, Rudolf, Saleh, Shadi, Schuster, Sebastian, Seeker, Wolfgang, Seraji, Mojgan, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Simov, Kiril, Smith, Aaron, Štěpánek, Jan, Suhr, Alane, Szántó, Zsolt, Tanaka, Takaaki, Tsarfaty, Reut, Uematsu, Sumire, Uria, Larraitz, Varga, Viktor, Vincze, Veronika, Žabokrtský, Zdeněk, Zeman, Daniel, and Zhu, Hanzhi
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, and Tamil
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
Rights:: Licence Universal Dependencies v1.2, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-1.2, and PUB

11. Deltacorpus 1.1

12. HamleDT 2.0

13. HamleDT 3.0

14. IWPT 2020 Shared Task Data and System Outputs

15. IWPT 2021 Shared Task Data and System Outputs

16. Lingua::Interset 2.026

17. Treebanks for Unified Taxonomy of Deep Syntactic Relations

18. Universal Dependencies 1.0

19. Universal Dependencies 1.1

20. Universal Dependencies 1.2

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Creator

Show values starting with

Language

Show values starting with

Publisher

Rights

Show values starting with

Subject

Show values starting with

Type

Date

Original context has metadata only

Harvested from