Original context has metadata only: false / Rights: PUB - LINDAT/CLARIAH-CZ Catalog Search Results

965. Machine Translation Testsuite for Gender-Consistent Translation

Creator:: Aires, João Paulo
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: machine translation, testsuite, evaluation, and gender
Language:: English and Czech
Description:: Document-level testsuite for evaluation of gender translation consistency. Our Document-Level test set consists of selected English documents from the WMT21 newstest annotated with gender information. Czech unnanotated references are also added for convenience. We semi-automatically annotated person names and pronouns to identify the gender of these elements as well as coreferences. Our proposed annotation consists of three elements: (1) an ID, (2) an element class, and (3) gender. The ID identifies a person's name and its occurrences (name and pronouns). The element class identifies whether the tag refers to a name or a pronoun. Finally, the gender information defines whether the element is masculine or feminine. We performed a series of NLP techniques to automatically identify person names and coreferences. This initial process resulted in a set containing 45 documents to be manually annotated. Thus, we started a manual annotation of these documents to make sure they are correctly tagged. See README.md for more details.
Rights:: Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), http://creativecommons.org/licenses/by-nc/4.0/, and PUB

966. MADED

Creator:: ridouane, tachicart and karim, bouzoubaa
Publisher:: ALELM
Type:: text, computationalLexicon, and lexicalConceptualResource
Subject:: lexicon and moroccan arabic
Language:: Moroccan Arabic and Arabic
Description:: Moroccan Dialect Electronic Dictionary (MDED) is an electronic lexicon containing almost 15000 MSA entries. They are written in Arabic letters and translated to Moroccan Arabic dialect. In addition, MDED entries are annotated useful metadata such as POS, Origin and root. MDED can be useful in some advanced NLP applications such as Machine translation and morphological analyzer.
Rights:: Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), http://creativecommons.org/licenses/by-nc/4.0/, and PUB

967. Malach Center User Interface 1.0

Creator:: Kocián, Jiří and Obdržálek, Pavel
Publisher:: Malach Center for Visual History, Institute of Formal and Applied Linguistics, Charles University
Type:: tool and toolService
Subject:: digital humanities, database, search engine, Jews, Holocaust, History, Oral History, Visual History, and digital archive of cultural heritage
Description:: Source code of the first full and running version for the Malach Center User Interface, does not contain data or metadata fo the digital objects and resources.
Rights:: GNU General Public Licence, version 3, http://opensource.org/licenses/GPL-3.0, and PUB

968. Malayalam Visual Genome 1.0

Creator:: Parida, Shantipriya and Bojar, Ondřej
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: image and corpus
Subject:: multi-modal, neural machine translation, English-Malayalam Multimodal Corpus, Malayalam Image Captioning, and English-Malayalam Multimodal Translation
Language:: English and Malayalam
Description:: Data ------- Malayalam Visual Genome (MVG for short) 1.0 has similar goals as Hindi Visual Genome (HVG) 1.1: to support the Malayalam language. Malayalam Visual Genome 1.0 is the first multi-modal dataset in Malayalam for machine translation and image captioning. Malayalam Visual Genome 1.0 serves in "WAT 2021 Multi-Modal Machine Translation Task". Malayalam Visual Genome is a multimodal dataset consisting of text and images suitable for English-to-Malayalam multimodal machine translation task and multimodal research. We follow the same selection of short English segments (captions) and the associated images from Visual Genome as HGV 1.1 has. For MVG, we automatically translated these captions from English to Malayalam and manually corrected them, taking the associated images into account. The training set contains 29K segments. Further 1K and 1.6K segments are provided in development and test sets, respectively, which follow the same (random) sampling from the original Hindi Visual Genome. A third test set is called ``challenge test set'' and consists of 1.4K segments. The challenge test set was created for the WAT2019 multi-modal task by searching for (particularly) ambiguous English words based on the embedding similarity and manually selecting those where the image helps to resolve the ambiguity. The surrounding words in the sentence however also often include sufficient cues to identify the correct meaning of the ambiguous word. For MVG, we simply translated the English side of the test sets to Malayalam, again utilizing machine translation to speed up the process. Dataset Formats ---------------------- The multimodal dataset contains both text and images. The text parts of the dataset (train and test sets) are in simple tab-delimited plain text files. All the text files have seven columns as follows: Column1 - image_id Column2 - X Column3 - Y Column4 - Width Column5 - Height Column6 - English Text Column7 - Malayalam Text The image part contains the full images with the corresponding image_id as the file name. The X, Y, Width and Height columns indicate the rectangular region in the image described by the caption. Data Statistics ------------------- The statistics of the current release are given below. Parallel Corpus Statistics --------------------------------- Dataset Segments English Words Malayalam Words ---------- -------------- -------------------- ----------------- Train 28930 143112 107126 Dev 998 4922 3619 Test 1595 7853 5689 Challenge Test 1400 8186 6044 -------------------- ------------ ------------------ ------------------ Total 32923 164073 122478 The word counts are approximate, prior to tokenization. Citation ----------- If you use this corpus, please cite the following paper: @article{hindi-visual-genome:2019, title={{Hindi Visual Genome: A Dataset for Multimodal English-to-Hindi Machine Translation}}, author={Parida, Shantipriya and Bojar, Ond{\v{r}}ej and Dash, Satya Ranjan}, journal={Computaci{\'o}n y Sistemas}, volume={23}, number={4}, pages={1499--1505}, year={2019} }
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

969. Mandatory Exercise for Czech Girls

Creator:: Aktualita
Publisher:: Národní filmový archiv
Type:: video and clip
Subject:: gymnastika, dívky cvičící, tělocvična, akce Kuratorium pro výchovu mládeže, Kuratorium pro výchovu mládeže akce, Kuratorium, and Český zvukový týdeník Aktualita::1943
Language:: Czech
Description:: Segment from Český zvukový týdeník Aktualita (Czech Aktualita Sound Newsreel) issue no. 50B from 1943 shows how, as part of mandatory service, girls aged 10 to 18 had to exercise for two hours a week under the supervision of trained instructors of the Board of Trustees for the Education of Youth.
Rights:: http://creativecommons.org/licenses/by-nc-nd/4.0/, PUB, and Creative Commons - Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)

970. Manifestation for Reinhard Heydrich in Brno

Creator:: Aktualita
Publisher:: Národní filmový archiv
Type:: video and clip
Subject:: tryzna Heydrich Reinhard, akt pietní Heydrich Reinhard, projev veřejný, shromáždění veřejné, orlice říšskoněmecká, znak zemský Čechy, znak zemský Morava, Heydrichiáda, Places::Brno::Zelný trh, People::Moravec Emanuel (1893-1945), and Český zvukový týdeník Aktualita::1942/25
Language:: Czech
Description:: Segment from Československý zvukový týdeník Aktualita (Czechoslovak Aktualita Sound Newsreel) 1942, issue no. 25, depicts a public demonstration on Cabbage Market Square (Zelný trh) in Brno on 12 June 1942, which was to vociferously condemn the assassination of Acting Reich Protector Reinhard Heydrich. The gathering was attended by 12,000 people. A grandstand in the middle of the crowded square is decorated with the Imperial Eagle and the national emblems of Bohemia and Moravia. The main speaker, Minister of Education and People´s Enlightenment Emanuel Moravec, encourages Czech people to take into account the past. The segment concludes with the Czech anthem (authentic sound and singing) with images of people with arms raised in the Nazi salute.
Rights:: http://creativecommons.org/licenses/by-nc-nd/4.0/, PUB, and Creative Commons - Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)

961. Ludvík Kuba (painter)

962. Ludvík Souček (publisher, bookseller)

963. Ludvík Venclík (film professional)

964. Ludvík Veverka (actor)

965. Machine Translation Testsuite for Gender-Consistent Translation

966. MADED

967. Malach Center User Interface 1.0

968. Malayalam Visual Genome 1.0

969. Mandatory Exercise for Czech Girls

970. Manifestation for Reinhard Heydrich in Brno

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Coverage

Creator

Show values starting with

Language

Show values starting with

Publisher

Show values starting with

Rights

Show values starting with

Subject

Show values starting with

Type

Show values starting with

Date

Original context has metadata only

Harvested from