Creator: Popel, Martin / Language: English - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Creator Popel, Martin Language English

21. Optimal reference translation of English-Czech WMT2020

Creator:: Kloudová, Věra, Mraček, David, Bojar, Ondřej, and Popel, Martin
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: translational equivalence, reference translation, optimal reference translation, and WMT
Language:: Czech and English
Description:: We define "optimal reference translation" as a translation thought to be the best possible that can be achieved by a team of human translators. Optimal reference translations can be used in assessments of excellent machine translations. We selected 50 documents (online news articles, with 579 paragraphs in total) from the 130 English documents included in the WMT2020 news test (http://www.statmt.org/wmt20/) with the aim to preserve diversity (style, genre etc.) of the selection. In addition to the official Czech reference translation provided by the WMT organizers (P1), we hired two additional translators (P2 and P3, native Czech speakers) via a professional translation agency, resulting in three independent translations. The main contribution of this dataset are two additional translations (i.e. optimal reference translations N1 and N2), done jointly by two translators-cum-theoreticians with an extreme care for various aspects of translation quality, while taking into account the translations P1-P3. We publish also internal comments (in Czech) for some of the segments. Translation N1 should be closer to the English original (with regards to the meaning and linguistic structure) and female surnames use the Czech feminine suffix (e.g. "Mai" is translated as "Maiová"). Translation N2 is more free, trying to be more creative, idiomatic and entertaining for the readers and following the typical style used in Czech media, while still preserving the rules of functional equivalence. Translation N2 is missing for the segments where it was not deemed necessary to provide two alternative translations. For applications/analyses needing translation of all segments, this should be interpreted as if N2 is the same as N1 for a given segment. We provide the dataset in two formats: OpenDocument spreadsheet (odt) and plain text (one file for each translation and the English original). Some words were highlighted using different colors during the creation of optimal reference translations; this highlighting and comments are present only in the odt format (some comments refer to row numbers in the odt file). Documents are separated by empty lines and each document starts with a special line containing the document name (e.g. "# upi.205735"), which allows alignment with the original WMT2020 news test. For the segments where N2 translations are missing in the odt format, the respective N1 segments are used instead in the plain-text format.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

22. Optimal Reference Translations from English to Czech

Creator:: Zouhar, Vilém, Kloudová, Věra, Popel, Martin, and Bojar, Ondřej
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: translation, evaluation, and optimal reference translation
Language:: English and Czech
Description:: This corpus contains annotations of translation quality from English to Czech in seven categories on both segment- and document-level. There are 20 documents in total, each with 4 translations (evaluated by each annotator in paralel) of 8 segments (can be longer than one sentence). Apart from the evaluation, the annotators also proposed their own, improved versions of the translations. There were 11 annotators in total, on expertise levels ranging from non-experts to professional translators.
Rights:: Creative Commons - Attribution 4.0 International (CC BY 4.0), http://creativecommons.org/licenses/by/4.0/, and PUB

23. QTLeap WSD/NED corpus

Creator:: Agirre, Eneko, Branco, António, Popel, Martin, and Simov, Kiril
Publisher:: University of the Basque Country, UPV/EHU, Faculty of Science, Univeristy of Lisbon, FCUL, Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), and Bulgarian Academy of Sciences, IICT-BAS
Type:: text and corpus
Subject:: annotated corpus and multilingual
Language:: Basque, Bulgarian, Czech, English, Portuguese, and Spanish
Description:: This corpora is part of Deliverable 5.5 of the European Commission project QTLeap FP7-ICT-2013.4.1-610516 (http://qtleap.eu). The texts are Q&A interactions from the real-user scenario (batches 1 and 2). The interactions in this corpus are available in Basque, Bulgarian, Czech, English, Portuguese and Spanish. The texts have been automatically annotated with NLP tools, including Word Sense Disambiguation, Named Entity Disambiguation and Coreference resolution. Please check deliverable D5.6 in http://qtleap.eu/deliverables for more information.
Rights:: Attribution-NonCommercial-ShareAlike 3.0 Unported (CC BY-NC-SA 3.0), http://creativecommons.org/licenses/by-nc-sa/3.0/, and PUB

24. Synthetic part of CzEng 2.0

Creator:: Popel, Martin
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: parallel corpus
Language:: Czech and English
Description:: CzEng is a sentence-parallel Czech-English corpus compiled at the Institute of Formal and Applied Linguistics (ÚFAL). While the full CzEng 2.0 is freely available for non-commercial research purposes from the project website (https://ufal.mff.cuni.cz/czeng), this release contains only the original monolingual parts of news text (csmono 53M and enmono 79M sentences) with automatic (synthetic) translations by CUBBITT. See the attached README for additional details such as the file format.
Rights:: Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), http://creativecommons.org/licenses/by-sa/4.0/, and PUB

29. Universal Dependencies 2.0 – CoNLL 2017 Shared Task Development and Test Data

Creator:: Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Antonsen, Lene, Aranzabe, Maria Jesus, Asahara, Masayuki, Ateyah, Luma, Attia, Mohammed, Atutxa, Aitziber, Badmaeva, Elena, Ballesteros, Miguel, Banerjee, Esha, Bank, Sebastian, Bauer, John, Bengoetxea, Kepa, Bhat, Riyaz Ahmad, Bick, Eckhard, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Burchardt, Aljoscha, Candito, Marie, Caron, Gauthier, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Cetin, Savas, Chalub, Fabricio, Choi, Jinho, Cho, Yongseok, Cinková, Silvie, Çöltekin, Çağrı, Connor, Miriam, de Marneffe, Marie-Catherine, de Paiva, Valeria, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Eli, Marhaba, Elkahky, Ali, Erjavec, Tomaž, Farkas, Richárd, Fernandez Alcalde, Hector, Foster, Jennifer, Freitas, Cláudia, Gajdošová, Katarína, Galbraith, Daniel, Garcia, Marcos, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grioni, Matias, Grūzītis, Normunds, Guillaume, Bruno, Habash, Nizar, Hajič, Jan, Hajič jr., Jan, Hà Mỹ, Linh, Harris, Kim, Haug, Dag, Hladká, Barbora, Hlaváčová, Jaroslava, Hohle, Petter, Ion, Radu, Irimia, Elena, Johannsen, Anders, Jørgensen, Fredrik, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Kayadelen, Tolga, Kettnerová, Václava, Kirchner, Jesse, Kotsyba, Natalia, Krek, Simon, Kwak, Sookyoung, Laippala, Veronika, Lambertino, Lorenzo, Lando, Tatiana, Lê Hồng, Phương, Lenci, Alessandro, Lertpradit, Saran, Leung, Herman, Li, Cheuk Ying, Li, Josie, Ljubešić, Nikola, Loginova, Olga, Lyashevskaya, Olga, Lynn, Teresa, Macketanz, Vivien, Makazhanov, Aibek, Mandl, Michael, Manning, Christopher, Manurung, Ruli, Mărănduc, Cătălina, Mareček, David, Marheinecke, Katrin, Martínez Alonso, Héctor, Martins, André, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Mendonça, Gustavo, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, More, Amir, Moreno Romero, Laura, Mori, Shunsuke, Moskalevskyi, Bohdan, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nainwani, Pinkey, Nedoluzhko, Anna, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikolaev, Vitaly, Nitisaroj, Rattima, Nurmi, Hanna, Ojala, Stina, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Perrier, Guy, Petrov, Slav, Piitulainen, Jussi, Pitler, Emily, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Rademaker, Alexandre, Real, Livy, Reddy, Siva, Rehm, Georg, Rinaldi, Larissa, Rituma, Laura, Rosa, Rudolf, Rovati, Davide, Saleh, Shadi, Sanguinetti, Manuela, Saulīte, Baiba, Sawanakunanon, Yanin, Schuster, Sebastian, Seddah, Djamé, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Shimada, Atsuko, Shohibussirri, Muh, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Smith, Aaron, Stella, Antonio, Strnadová, Jana, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Taji, Dima, Tanaka, Takaaki, Trosterud, Trond, Trukhina, Anna, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Urešová, Zdeňka, Uria, Larraitz, Uszkoreit, Hans, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Washington, Jonathan North, Yu, Zhuoran, Žabokrtský, Zdeněk, Zeman, Daniel, and Zhu, Hanzhi
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, Urdu, Northern Sami, Upper Sorbian, Russia Buriat, and Northern Kurdish
Description:: Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This release contains the test data used in the CoNLL 2017 shared task on parsing Universal Dependencies. Due to the shared task the test data was held hidden and not released together with the training and development data of UD 2.0. Therefore this release complements the UD 2.0 release (http://hdl.handle.net/11234/1-1983) to a full release of UD treebanks. In addition, the present release contains 18 new parallel test sets and 4 test sets in surprise languages. The present release also includes the development data already released with UD 2.0. Unlike regular UD releases, this one uses the folder-file structure that was visible to the systems participating in the shared task.
Rights:: Licence Universal Dependencies v2.0, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.0, and PUB

30. Universal Dependencies 2.0 alpha (obsolete)

Creator:: Nivre, Joakim, Agić, Željko, Ahrenberg, Lars, Aranzabe, Maria Jesus, Asahara, Masayuki, Atutxa, Aitziber, Ballesteros, Miguel, Bauer, John, Bengoetxea, Kepa, Bhat, Riyaz Ahmad, Bick, Eckhard, Bosco, Cristina, Bouma, Gosse, Bowman, Sam, Candito, Marie, Cebiroğlu Eryiğit, Gülşen, Celano, Giuseppe G. A., Chalub, Fabricio, Choi, Jinho, Çöltekin, Çağrı, Connor, Miriam, Davidson, Elizabeth, de Marneffe, Marie-Catherine, de Paiva, Valeria, Diaz de Ilarraza, Arantza, Dobrovoljc, Kaja, Dozat, Timothy, Droganova, Kira, Dwivedi, Puneet, Eli, Marhaba, Erjavec, Tomaž, Farkas, Richárd, Foster, Jennifer, Freitas, Cláudia, Gajdošová, Katarína, Galbraith, Daniel, Garcia, Marcos, Ginter, Filip, Goenaga, Iakes, Gojenola, Koldo, Gökırmak, Memduh, Goldberg, Yoav, Gómez Guinovart, Xavier, Gonzáles Saavedra, Berta, Grioni, Matias, Grūzītis, Normunds, Guillaume, Bruno, Habash, Nizar, Hajič, Jan, Hà Mỹ, Linh, Haug, Dag, Hladká, Barbora, Hohle, Petter, Ion, Radu, Irimia, Elena, Johannsen, Anders, Jørgensen, Fredrik, Kaşıkara, Hüner, Kanayama, Hiroshi, Kanerva, Jenna, Kotsyba, Natalia, Krek, Simon, Laippala, Veronika, Lê Hồng, Phương, Lenci, Alessandro, Ljubešić, Nikola, Lyashevskaya, Olga, Lynn, Teresa, Makazhanov, Aibek, Manning, Christopher, Mărănduc, Cătălina, Mareček, David, Martínez Alonso, Héctor, Martins, André, Mašek, Jan, Matsumoto, Yuji, McDonald, Ryan, Missilä, Anna, Mititelu, Verginica, Miyao, Yusuke, Montemagni, Simonetta, More, Amir, Mori, Shunsuke, Moskalevskyi, Bohdan, Muischnek, Kadri, Mustafina, Nina, Müürisep, Kaili, Nguyễn Thị, Lương, Nguyễn Thị Minh, Huyền, Nikolaev, Vitaly, Nurmi, Hanna, Ojala, Stina, Osenova, Petya, Øvrelid, Lilja, Pascual, Elena, Passarotti, Marco, Perez, Cenel-Augusto, Perrier, Guy, Petrov, Slav, Piitulainen, Jussi, Plank, Barbara, Popel, Martin, Pretkalniņa, Lauma, Prokopidis, Prokopis, Puolakainen, Tiina, Pyysalo, Sampo, Rademaker, Alexandre, Ramasamy, Loganathan, Real, Livy, Rituma, Laura, Rosa, Rudolf, Saleh, Shadi, Sanguinetti, Manuela, Saulīte, Baiba, Schuster, Sebastian, Seddah, Djamé, Seeker, Wolfgang, Seraji, Mojgan, Shakurova, Lena, Shen, Mo, Sichinava, Dmitry, Silveira, Natalia, Simi, Maria, Simionescu, Radu, Simkó, Katalin, Šimková, Mária, Simov, Kiril, Smith, Aaron, Suhr, Alane, Sulubacak, Umut, Szántó, Zsolt, Taji, Dima, Tanaka, Takaaki, Tsarfaty, Reut, Tyers, Francis, Uematsu, Sumire, Uria, Larraitz, van Noord, Gertjan, Varga, Viktor, Vincze, Veronika, Washington, Jonathan North, Žabokrtský, Zdeněk, Zeldes, Amir, Zeman, Daniel, and Zhu, Hanzhi
Publisher:: Universal Dependencies Consortium
Type:: text and corpus
Subject:: treebank, dependency, syntax, morphology, harmonized annotation, interset, universal tagset, and stanford dependencies
Language:: Ancient Greek (to 1453), Arabic, Basque, Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Gothic, Modern Greek (1453-), Hebrew, Hindi, Hungarian, Indonesian, Irish, Italian, Japanese, Latin, Norwegian, Church Slavic, Persian, Polish, Portuguese, Romanian, Slovenian, Spanish, Swedish, Tamil, Catalan, Chinese, Galician, Kazakh, Latvian, Russian, Turkish, Coptic, Sanskrit, Slovak, Ukrainian, Uighur, Vietnamese, Belarusian, Korean, Lithuanian, and Urdu
Description:: This release contains errors in several files. Please use http://hdl.handle.net/11234/1-1983 instead.
Rights:: Licence Universal Dependencies v2.0, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-UD-2.0, and PUB

21. Optimal reference translation of English-Czech WMT2020

22. Optimal Reference Translations from English to Czech

23. QTLeap WSD/NED corpus

24. Synthetic part of CzEng 2.0

25. Universal Dependencies 1.2

26. Universal Dependencies 1.3

27. Universal Dependencies 1.4

28. Universal Dependencies 2.0

29. Universal Dependencies 2.0 – CoNLL 2017 Shared Task Development and Test Data

30. Universal Dependencies 2.0 alpha (obsolete)

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Creator

Show values starting with

Language

Show values starting with

Publisher

Rights

Show values starting with

Subject

Show values starting with

Type

Date

Original context has metadata only

Harvested from