Harvested from: LINDAT/CLARIAH-CZ repository / Language: English - LINDAT/CLARIAH-CZ Catalog Search Results

151. Hindi Visual Genome 1.1

Creator:: Parida, Shantipriya and Bojar, Ondřej
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: multilingual, neural machine translation, multi-modal, English-Hindi parallel corpus, image captioning, and image annotation
Language:: English and Hindi
Description:: Data ---- Hindi Visual Genome 1.1 is an updated version of Hindi Visual Genome 1.0. The update concerns primarily the text part of Hindi Visual Genome, fixing translation issues reported during WAT 2019 multimodal task. In the image part, only one segment and thus one image were removed from the dataset. Hindi Visual Genome 1.1 serves in "WAT 2020 Multi-Modal Machine Translation Task". Hindi Visual Genome is a multimodal dataset consisting of text and images suitable for English-to-Hindi multimodal machine translation task and multimodal research. We have selected short English segments (captions) from Visual Genome along with associated images and automatically translated them to Hindi with manual post-editing, taking the associated images into account. The training set contains 29K segments. Further 1K and 1.6K segments are provided in a development and test sets, respectively, which follow the same (random) sampling from the original Hindi Visual Genome. A third test set is called ``challenge test set'' consists of 1.4K segments and it was released for WAT2019 multi-modal task. The challenge test set was created by searching for (particularly) ambiguous English words based on the embedding similarity and manually selecting those where the image helps to resolve the ambiguity. The surrounding words in the sentence however also often include sufficient cues to identify the correct meaning of the ambiguous word. Dataset Formats -------------- The multimodal dataset contains both text and images. The text parts of the dataset (train and test sets) are in simple tab-delimited plain text files. All the text files have seven columns as follows: Column1 - image_id Column2 - X Column3 - Y Column4 - Width Column5 - Height Column6 - English Text Column7 - Hindi Text The image part contains the full images with the corresponding image_id as the file name. The X, Y, Width and Height columns indicate the rectangular region in the image described by the caption. Data Statistics ---------------- The statistics of the current release is given below. Parallel Corpus Statistics --------------------------- Dataset Segments English Words Hindi Words ------- --------- ---------------- ------------- Train 28930 143164 145448 Dev 998 4922 4978 Test 1595 7853 7852 Challenge Test 1400 8186 8639 ------- --------- ---------------- ------------- Total 32923 164125 166917 The word counts are approximate, prior to tokenization. Citation -------- If you use this corpus, please cite the following paper: @article{hindi-visual-genome:2019, title={{Hindi Visual Genome: A Dataset for Multimodal English-to-Hindi Machine Translation}}, author={Parida, Shantipriya and Bojar, Ond{\v{r}}ej and Dash, Satya Ranjan}, journal={Computaci{\'o}n y Sistemas}, volume={23}, number={4}, pages={1499--1505}, year={2019} }
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

152. Hunglish Corpus

Publisher:: Academy of Sciences and Budapest University of Technology and Economics Media Research (BME MOKK)
Type:: corpus
Subject:: parallel corpus
Language:: English and Hungarian
Description:: Billingual written general; 2 million sentences
Rights:: CC

153. ICLE International Corpus of Learner English

Publisher:: Centre for English Corpus Linguistics, Université catholique de Louvain
Type:: corpus
Language:: English
Description:: over 3 million words of writing by learners of English from 14 different mother tongue backgrounds
Rights:: Not specified

154. IDENTICv1.0

Creator:: Larasati, Septina Dian
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: Indonesian-English parallel corpus and parallel corpus
Language:: Indonesian and English
Description:: IDENTIC is an Indonesian-English parallel corpus for research purposes. The corpus is a bilingual corpus paired with English. The aim of this work is to build and provide researchers a proper Indonesian-English textual data set and also to promote research in this language pair. The corpus contains texts coming from different sources with different genres. and The research leading to these results has received funding from the European Commission’s 7th Framework Program under grant agreement no 238405 (CLARA) and by the grant LC536 Centrum Komputacni Lingvistiky of the Czech Ministry of Education.
Rights:: Attribution-NonCommercial-ShareAlike 3.0 Unported (CC BY-NC-SA 3.0), http://creativecommons.org/licenses/by-nc-sa/3.0/, and PUB

155. IDENTICv1.0-raw

Creator:: Larasati, Septina Dian
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: Indonesian-English parallel corpus and parallel corpus
Language:: Indonesian and English
Description:: Raw Text
Rights:: Creative Commons - Attribution 3.0 Unported (CC BY 3.0), http://creativecommons.org/licenses/by/3.0/, and PUB

156. IJS-ELAN

Type:: corpus
Language:: English and Slovenian
Description:: parallel, mixed text; 2x0.5 mil. words; TEI / morphosyntactic tags
Rights:: Not specified

157. Image Annotation Tool

Creator:: Roček, Martin
Publisher:: Charles University, Faculty of Arts
Type:: service and toolService
Subject:: tei and javascript
Language:: English
Description:: Image annotation tool is a web application that allows users to mark zones of interest in an image. These zones are then converted to TEI P5 code snippet that can be used in your document to connect the image and the text. This tool was developed to help students and teachers at the Faculty of Arts, Charles University to mark and annotate images of manuscripts.
Rights:: GNU General Public Licence, version 3, http://opensource.org/licenses/GPL-3.0, and PUB

158. Implemented Spelling Rules

Creator:: Saty, Ahmed, Aouragh, Si Lhoussain, and Bouzoubaa, Karim
Publisher:: Sudan University of Science and Technology
Type:: text, other, and languageDescription
Subject:: Implemented Spelling Rules
Language:: Arabic and English
Description:: The book [1] contains spelling rules classified into ten categories, each category containing many rules. This XML file presents our implemented rules classified with six category tags, as is the case in the book. We implemented 24 rules since the remaining rules require diacritical and morphological analysis that are outside the scope of our present work. References: [1] Dr.Fahmy Al-Najjar, 'Spelling rules in ten easy lessons', Al Kawthar Library,2008. Available: https://www.alukah.net/library/0/53498/%D9%82%D9%88%D8%A7%D8%B9%D8%AF-%D8%A7%D9%84%D8%A5%D9%85%D9%84%D8%A7%D8%A1-%D9%81%D9%8A-%D8%B9%D8%B4%D8%B1%D8%A9-%D8%AF%D8%B1%D9%88%D8%B3-%D8%B3%D9%87%D9%84%D8%A9-pdf/
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

159. Individual Textual Profiles of Hillary Clinton and Donald Trump

Creator:: Kvítková, Alena
Publisher:: Charles University, Faculty of Arts, Department of English Language and ELT Methodology
Type:: text and corpus
Subject:: idiolect, individual textual profile, Clinton, Trump, corpus, presidential debates, American president, candidates, Democrats, and Republicans
Language:: English
Description:: This corpus consists of full transcriptions of both Democratic and Republican 2016 presidential candidate debates, with a special focus on the idiolects of Hillary Clinton and Donald Trump against the background of the speeches of other candidates for the post of president of the United States. The transcriptions are sourced from the American Presidency Project at the University of California, Santa Barbara. Any use of the material requires a prior and explicit written permission by the project administrator (contact policy@ucsb.edu). This corpus material is now being shared with their kindly permission.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

160. INTERA Terminological Lexicon

Type:: lexicalConceptualResource
Language:: Bulgarian, English, Modern Greek (1453-), Serbian, and Slovenian
Description:: 17357 terms, XML
Rights:: Not specified

151. Hindi Visual Genome 1.1

152. Hunglish Corpus

153. ICLE International Corpus of Learner English

154. IDENTICv1.0

155. IDENTICv1.0-raw

156. IJS-ELAN

157. Image Annotation Tool

158. Implemented Spelling Rules

159. Individual Textual Profiles of Hillary Clinton and Donald Trump

160. INTERA Terminological Lexicon

Limit your search

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Show values starting with

Coverage

Show values starting with

Creator

Show values starting with

Format

Language

Show values starting with

Publisher

Show values starting with

Rights

Show values starting with

Subject

Show values starting with

Type

Show values starting with

Date

Original context has metadata only

Harvested from