Creator: Zeman, Daniel / Language: Danish - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Creator Zeman, Daniel Language Danish

11. HamleDT 2.0

Creator:: Zeman, Daniel, Mareček, David, Mašek, Jan, Popel, Martin, Ramasamy, Loganathan, Rosa, Rudolf, Štěpánek, Jan, and Žabokrtský, Zdeněk
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: treebank, Stanford dependencies, Prague dependencies, harmonization, common annotation style, and Interset
Language:: Arabic, Bulgarian, Bengali, Catalan, Czech, Danish, German, Modern Greek (1453-), English, Spanish, Estonian, Basque, Persian, Finnish, Ancient Greek (to 1453), Hindi, Hungarian, Italian, Japanese, Latin, Dutch, Portuguese, Romanian, Russian, Slovak, Slovenian, Swedish, Tamil, Telugu, and Turkish
Description:: HamleDT 2.0 is a collection of 30 existing treebanks harmonized into a common annotation style, the Prague Dependencies, and further transformed into Stanford Dependencies, a treebank annotation style that became popular recently. We use the newest basic Universal Stanford Dependencies, without added language-specific subtypes.
Rights:: HamleDT 2.0 Licence Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-hamledt-2.0, and ACA

12. HamleDT 3.0

Creator:: Zeman, Daniel, Mareček, David, Mašek, Jan, Popel, Martin, Ramasamy, Loganathan, Rosa, Rudolf, Štěpánek, Jan, and Žabokrtský, Zdeněk
Publisher:: Charles University
Type:: text and corpus
Subject:: annotated corpus, morphology, syntax, dependency, treebank, harmonized annotation, and common annotation style
Language:: Arabic, Basque, Bengali, Bulgarian, Catalan, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Modern Greek (1453-), Ancient Greek (to 1453), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Persian, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish, Tamil, Telugu, and Turkish
Description:: HamleDT (HArmonized Multi-LanguagE Dependency Treebank) is a compilation of existing dependency treebanks (or dependency conversions of other treebanks), transformed so that they all conform to the same annotation style. This version uses Universal Dependencies as the common annotation style. Update (November 1017): for a current collection of harmonized dependency treebanks, we recommend using the Universal Dependencies (UD). All of the corpora that are distributed in HamleDT in full are also part of the UD project; only some corpora from the Patch group (where HamleDT provides only the harmonizing scripts but not the full corpus data) are available in HamleDT but not in UD.
Rights:: HamleDT 3.0 License Terms, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-hamledt-3.0, and PUB

13. Lingua::Interset 2.026

Creator:: Zeman, Daniel
Publisher:: Charles University, Faculty of Mathematics and Physics
Type:: tool and toolService
Subject:: morphology, part of speech, conversion, and tagset
Language:: Arabic, Bulgarian, Bengali, Catalan, Czech, Danish, German, Modern Greek (1453-), English, Spanish, Estonian, Basque, Persian, Finnish, Ancient Greek (to 1453), Hebrew, Hindi, Croatian, Japanese, Multiple languages, and Portuguese
Description:: Lingua::Interset is a universal morphosyntactic feature set to which all tagsets of all corpora/languages can be mapped. Version 2.026 covers 37 different tagsets of 21 languages. Limited support of the older drivers for other languages (which are not included in this package but are available for download elsewhere) is also available; these will be fully ported to Interset 2 in future. Interset is implemented as Perl libraries. It is also available via CPAN.
Rights:: Artistic License (Perl) 1.0, http://opensource.org/licenses/Artistic-Perl-1.0, and PUB

14. Slavic Forest, Norwegian Wood (scripts)

Creator:: Rosa, Rudolf, Zeman, Daniel, Mareček, David, and Žabokrtský, Zdeněk
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: suiteOfTools and toolService
Subject:: parsing, dependency parser, universal dependencies, and cross-lingual parsing
Language:: Czech, Slovak, Slovenian, Croatian, Danish, Swedish, and Norwegian
Description:: Tools and scripts used to create the cross-lingual parsing models submitted to VarDial 2017 shared task (https://bitbucket.org/hy-crossNLP/vardial2017), as described in the linked paper. The trained UDPipe models themselves are published in a separate submission (https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-1971). For each source (SS, e.g. sl) and target (TT, e.g. hr) language, you need to add the following into this directory: - treebanks (Universal Dependencies v1.4): SS-ud-train.conllu TT-ud-predPoS-dev.conllu - parallel data (OpenSubtitles from Opus): OpenSubtitles2016.SS-TT.SS OpenSubtitles2016.SS-TT.TT !!! If they are originally called ...TT-SS... instead of ...SS-TT..., you need to symlink them (or move, or copy) !!! - target tagging model TT.tagger.udpipe All of these can be obtained from https://bitbucket.org/hy-crossNLP/vardial2017 You also need to have: - Bash - Perl 5 - Python 3 - word2vec (https://code.google.com/archive/p/word2vec/); we used rev 41 from 15th Sep 2014 - udpipe (https://github.com/ufal/udpipe); we used commit 3e65d69 from 3rd Jan 2017 - Treex (https://github.com/ufal/treex); we used commit d27ee8a from 21st Dec 2016 The most basic setup is the sl-hr one (train_sl-hr.sh): - normalization of deprels - 1:1 word-alignment of parallel data with Monolingual Greedy Aligner - simple word-by-word translation of source treebank - pre-training of target word embeddings - simplification of morpho feats (use only Case) - and finally, training and evaluating the parser Both da+sv-no (train_ds-no.sh) and cs-sk (train_cs-sk.sh) add some cross-tagging, which seems to be useful only in specific cases (see paper for details). Moreover, cs-sk also adds more morpho features, selecting those that seem to be very often shared in parallel data. The whole pipeline takes tens of hours to run, and uses several GB of RAM, so make sure to use a powerful computer.
Rights:: GNU General Public License 2 or later (GPL-2.0), http://opensource.org/licenses/GPL-2.0, and PUB

20. Universal Dependencies 2.0 – CoNLL 2017 Shared Task Development and Test Data

Creator:: Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Antonsen, Lene, Aranzabe, Maria Jesus, Asahara, Masayuki, Ateyah, Luma, Attia, Mohammed, Atutxa, Aitziber, Badmaeva, Elena, Ballesteros, Miguel, Banerjee, Esha, Bank, Sebastian, Bauer, John, Bengoetxea, Kepa, Bhat, Riyaz Ahmad, Bick, Eckhard, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Burchardt, Aljoscha, Candito, Marie, Caron, Gauthier, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Cetin, Savas, Chalub, Fabricio, Choi, Jinho, Cho, Yongseok, Cinková, Silvie, Çöltekin, Çağrı, Connor, Miriam, de Marneffe, Marie-Catherine, de Paiva, Valeria, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Eli, Marhaba, Elkahky, Ali, Erjavec, Tomaž, Farkas, Richárd, Fernandez Alcalde, Hector, Foster, Jennifer, Freitas, Cláudia, Gajdošová, Katarína, Galbraith, Daniel, Garcia, Marcos, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grioni, Matias, Grūzītis, Normunds, Guillaume, Bruno, Habash, Nizar, Hajič, Jan, Hajič jr., Jan, Hà Mỹ, Linh, Harris, Kim, Haug, Dag, Hladká, Barbora, Hlaváčová, Jaroslava, Hohle, Petter, Ion, Radu, Irimia, Elena, Johannsen, Anders, Jørgensen, Fredrik, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Kayadelen, Tolga, Kettnerová, Václava, Kirchner, Jesse, Kotsyba, Natalia, Krek, Simon, Kwak, Sookyoung, Laippala, Veronika, Lambertino, Lorenzo, Lando, Tatiana, Lê Hồng, Phương, Lenci, Alessandro, Lertpradit, Saran, Leung, Herman, Li, Cheuk Ying, Li, Josie, Ljubešić, Nikola, Loginova, Olga, Lyashevskaya, Olga, Lynn, Teresa, Macketanz, Vivien, Makazhanov, Aibek, Mandl, Michael, Manning, Christopher, Manurung, Ruli, Mărănduc, Cătălina, Mareček, David, Marheinecke, Katrin, Martínez Alonso, Héctor, Martins, André, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Mendonça, Gustavo, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, More, Amir, Moreno Romero, Laura, Mori, Shunsuke, Moskalevskyi, Bohdan, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nainwani, Pinkey, Nedoluzhko, Anna, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikolaev, Vitaly, Nitisaroj, Rattima, Nurmi, Hanna, Ojala, Stina, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Perrier, Guy, Petrov, Slav, Piitulainen, Jussi, Pitler, Emily, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Rademaker, Alexandre, Real, Livy, Reddy, Siva, Rehm, Georg, Rinaldi, Larissa, Rituma, Laura, Rosa, Rudolf, Rovati, Davide, Saleh, Shadi, Sanguinetti, Manuela, Saulīte, Baiba, Sawanakunanon, Yanin, Schuster, Sebastian, Seddah, Djamé, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Shimada, Atsuko, Shohibussirri, Muh, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Smith, Aaron, Stella, Antonio, Strnadová, Jana, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Taji, Dima, Tanaka, Takaaki, Trosterud, Trond, Trukhina, Anna, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Urešová, Zdeňka, Uria, Larraitz, Uszkoreit, Hans, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Washington, Jonathan North, Yu, Zhuoran, Žabokrtský, Zdeněk, Zeman, Daniel, and Zhu, Hanzhi
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, Urdu, Northern Sami, Upper Sorbian, Russia Buriat, and Northern Kurdish
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This release contains the test data used in the CoNLL 2017 shared task on parsing Universal Dependencies. Due to the shared task the test data was held hidden and not released together with the training and development data of UD 2.0. Therefore this release complements the UD 2.0 release (http://hdl.handle.net/11234/1-1983) to a full release of UD treebanks. In addition, the present release contains 18 new parallel test sets and 4 test sets in surprise languages. The present release also includes the development data already released with UD 2.0. Unlike regular UD releases, this one uses the folder-file structure that was visible to the systems participating in the shared task.
Rights:: Licence Universal Dependencies v2.0, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.0, and PUB

11. HamleDT 2.0

12. HamleDT 3.0

13. Lingua::Interset 2.026

14. Slavic Forest, Norwegian Wood (scripts)

15. Universal Dependencies 1.1

16. Universal Dependencies 1.2

17. Universal Dependencies 1.3

18. Universal Dependencies 1.4

19. Universal Dependencies 2.0

20. Universal Dependencies 2.0 – CoNLL 2017 Shared Task Development and Test Data

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Creator

Show values starting with

Language

Show values starting with

Publisher

Rights

Show values starting with

Subject

Show values starting with

Type

Original context has metadata only

Harvested from