Source code of the first full and running version for the Malach Center User Interface, does not contain data or metadata fo the digital objects and resources.
Data
-------
Malayalam Visual Genome (MVG for short) 1.0 has similar goals as Hindi Visual Genome (HVG) 1.1: to support the Malayalam language. Malayalam Visual Genome 1.0 is the first multi-modal dataset in Malayalam for machine translation and image captioning.
Malayalam Visual Genome 1.0 serves in "WAT 2021 Multi-Modal Machine Translation Task".
Malayalam Visual Genome is a multimodal dataset consisting of text and images suitable for English-to-Malayalam multimodal machine translation task and multimodal research. We follow the same selection of short English segments (captions) and the associated images from Visual Genome as HGV 1.1 has. For MVG, we automatically translated these captions from English to Malayalam and manually corrected them, taking the associated images into account.
The training set contains 29K segments. Further 1K and 1.6K segments are provided in development and test sets, respectively, which follow the same (random) sampling from the original Hindi Visual Genome.
A third test set is called ``challenge test set'' and consists of 1.4K segments. The challenge test set was created for the WAT2019 multi-modal task by searching for (particularly) ambiguous English words based on the embedding similarity and manually selecting those where the image helps to resolve the ambiguity. The surrounding words in the sentence however also often include sufficient cues to identify the correct meaning of the ambiguous word. For MVG, we simply translated the English side of the test sets to Malayalam, again utilizing machine translation to speed up the process.
Dataset Formats
----------------------
The multimodal dataset contains both text and images.
The text parts of the dataset (train and test sets) are in simple tab-delimited plain text files.
All the text files have seven columns as follows:
Column1 - image_id
Column2 - X
Column3 - Y
Column4 - Width
Column5 - Height
Column6 - English Text
Column7 - Malayalam Text
The image part contains the full images with the corresponding image_id as the file name. The X, Y, Width and Height columns indicate the rectangular region in the image described by the caption.
Data Statistics
-------------------
The statistics of the current release are given below.
Parallel Corpus Statistics
---------------------------------
Dataset Segments English Words Malayalam Words
---------- -------------- -------------------- -----------------
Train 28930 143112 107126
Dev 998 4922 3619
Test 1595 7853 5689
Challenge Test 1400 8186 6044
-------------------- ------------ ------------------ ------------------
Total 32923 164073 122478
The word counts are approximate, prior to tokenization.
Citation
-----------
If you use this corpus, please cite the following paper:
@article{hindi-visual-genome:2019, title={{Hindi Visual Genome: A Dataset for Multimodal English-to-Hindi Machine Translation}}, author={Parida, Shantipriya and Bojar, Ond{\v{r}}ej and Dash, Satya Ranjan}, journal={Computaci{\'o}n y Sistemas}, volume={23}, number={4}, pages={1499--1505}, year={2019} }
Segment from Český zvukový týdeník Aktualita (Czech Aktualita Sound Newsreel) issue no. 50B from 1943 shows how, as part of mandatory service, girls aged 10 to 18 had to exercise for two hours a week under the supervision of trained instructors of the Board of Trustees for the Education of Youth.
Segment from Československý zvukový týdeník Aktualita (Czechoslovak Aktualita Sound Newsreel) 1942, issue no. 25, depicts a public demonstration on Cabbage Market Square (Zelný trh) in Brno on 12 June 1942, which was to vociferously condemn the assassination of Acting Reich Protector Reinhard Heydrich. The gathering was attended by 12,000 people. A grandstand in the middle of the crowded square is decorated with the Imperial Eagle and the national emblems of Bohemia and Moravia. The main speaker, Minister of Education and People´s Enlightenment Emanuel Moravec, encourages Czech people to take into account the past. The segment concludes with the Czech anthem (authentic sound and singing) with images of people with arms raised in the Nazi salute.
Segment from Československý zvukový týdeník Aktualita (Czechoslovak Aktualita Sound Newsreel) 1942, issue no. 27A, depicts a public manifestation held at the Chapel of Saint Anthony of Padua on Blatnická Mountain in Moravian Slovakia on 28 June 1942, which was to unequivocally condemn the assassination of Acting Reich Protector Reinhard Heydrich. The gathering was attended by 35,000 to 40,000 people. Footage of the participants in festive folk costumes. The main organizer of the event, Minister of Education and People´s Enlightenment Emanuel Moravec, delivers a speech from a grandstand (silent). People wearing Moravian-Slovakian folk costumes with arms raised in the Nazi salute.
Segment from Československý zvukový týdeník Aktualita (Czechoslovak Aktualita Sound Newsreel) 1942, issue no. 27A, captures a public manifestation held in Moravská Ostrava on 30 June 1942, which was to unequivocally condemn the assassination of Acting Reich Protector Reinhard Heydrich. The gathering was attended by more than 80,000 people. Footage of miners in their uniforms. Miner Karel Juříček and Minister of Education and People´s Enlightenment Emanuel Moravec (silent) speak from a grandstand. The segment concludes with the Czech anthem (authentic sound), supplemented with images of people with arms raised in the Nazi salute.
Segment from Československý zvukový týdeník Aktualita (Czechoslovak Aktualita Sound Newsreel) 1942, issue no. 25, captures a public manifestation held on Hlavní náměstí (present day náměstí Republiky) in Plzeň on 16 June 1942. The manifestation was to demonstrate condemnation of the assassination of Acting Reich Protector Reinhard Heydrich. The gathering was attended by approximately 60,000 people. The grandstand is decorated with the Reichsadler (Imperial Eagle) and the state emblems of Bohemia and Moravia. A speech by the Minister of Education and People's Enlightenment, Emanuel Moravec, follows. The segment concludes with the Czech anthem (authentic sound), supplemented with images of people with arms raised in the nazi salute.
Segment from Československý zvukový týdeník Aktualita (Czechoslovak Aktualita Sound Newsreel) 1942, issue no. 23, depicts the so-called Manifestation of the Czech People for the Reich, organised by the Protectorate Government on Old Town Square in Prague on 2 June 1942. The gathering of 65,000 people, whose attendance was "highly recommended", was to unequivocally condemn the assassination attempt on Acting Reich Protector, SS-Obergruppenführer Reinhard Heydrich as well as all of the activities of the Czechoslovak Government in Exile in Great Britain. Members of the Protectorate Government climb up a grandstand in front of Týn Church. High-angled shots of the crowded square. A speech by Prime Minister Jaroslav Krejčí. The footage shows Minister of Agriculture and Forestry Adolf Hrubý, Minister of Economic Affairs and Labour Walter Bertsch, Minister of the Interior Richard Bienert, Minister of Finance Josef Kalfus, Minister of Transport Jindřich Kamenický, and Minister of Education and People´s Enlightenment Emanuel Moravec. A speech by the General Secretary of the Board of Trustees for the Education of Youth in Bohemia and Moravia František Teuner, and concluding words by Minister of Education and People's Enlightenment Emanuel Moravec. A wide shot of the crowded square. The segment concludes with the Czech anthem (authentic sound and singing) with images showing people with arms raised in the Nazi salute.
Segment from Československý zvukový týdeník Aktualita (Czechoslovak Aktualita Sound Newsreel) 1942, issue no. 26, from a public demonstration held in Tábor on 20 June 1942, which was to unequivocally condemn the assassination of Acting Reich Protector Reinhard Heydrich. The gathering was attended by 30,000 people. A wide shot of the town of Tábor. A view of the crowded square. A grandstand in the middle of the crowded square is decorated with the Imperial Eagle and the national emblems of Bohemia and Moravia. The manifestation is opened by Minister of the Protectorate Government Jaroslav Krejčí. A member of the National Theatre´s drama company, actor Jiří Dohnal, delivers a manifesto of State President Emil Hácha, followed by a speech by Minister of Agriculture and Forestry Adolf Hrubý. The manifestation is concluded by Minister of Education and People´s Enlightenment Emanuel Moravec. The segment includes images of Tábor´s important monuments (the Town Hall Tower, Kotnov). The segment concludes with the Czech anthem (authentic sound) supplemented with images of people with arms raised in the Nazi salute.