Harvested from: LINDAT/CLARIAH-CZ repository / Rights: PUB / Type: corpus

1. Annotated Corpus of Czech Case Law for Reference Recognition Tasks

Creator:: Harašta, Jakub, Šavelka, Jaromír, Kasl, František, Kotková, Adéla, Loutocký, Pavel, Míšek, Jakub, Procházková, Daniela, Pullmannová, Helena, Semenišín, Petr, Šejnová, Tamara, Šimková, Nikola, Vosinek, Michal, Zavadilová, Lucie, and Zibner, Jan
Publisher:: Masaryk University, Brno
Type:: text and corpus
Subject:: reference recognition and legal texts
Language:: Czech
Description:: Annotated corpus of 350 decision of Czech top-tier courts (Supreme Court, Supreme Administrative Court, Constitutional Court). Every decision is annotated by two trained annotators and then manually adjudicated by one trained curator to solve possible disagreements between annotators. Adjudication was conducted non-destructively, therefore dataset contains all original annotations. Corpus was developed as training and testing material for reference recognition tasks. Dataset contains references to other court decisions and literature. All references consist of basic units (identifier of court decision, identification of court issuing referred decision, author of book or article, title of book or article, point of interest in referred document etc.), values (polarity, depth of discussion etc.).
Rights:: Creative Commons - Attribution 4.0 International (CC BY 4.0), http://creativecommons.org/licenses/by/4.0/, and PUB

2. Annotated Corpus of Czech Case Law for Reference Recognition Tasks (2019-06-25)

Creator:: Harašta, Jakub, Šavelka, Jaromír, Kasl, František, Kotková, Adéla, Loutocký, Pavel, Míšek, Jakub, Procházková, Daniela, Pullmannová, Helena, Semenišín, Petr, Šejnová, Tamara, Šimková, Nikola, Vosinek, Michal, Zavadilová, Lucie, and Zibner, Jan
Publisher:: Masaryk University, Brno
Type:: text and corpus
Subject:: reference recognition and legal texts
Language:: Czech
Description:: Annotated corpus of 350 decision of Czech top-tier courts (Supreme Court, Supreme Administrative Court, Constitutional Court). Every decision is annotated by two trained annotators and then manually adjudicated by one trained curator to solve possible disagreements between annotators. Adjudication was conducted non-destructively, therefore corpus (raw) contains all original annotations. Corpus was developed as training and testing material for reference recognition tasks. Dataset contains references to other court decisions and literature. All references consist of basic units (identifier of court decision, identification of court issuing referred decision, author of book or article, title of book or article, point of interest in referred document etc.), values (polarity, depth of discussion etc.).
Rights:: Creative Commons - Attribution 4.0 International (CC BY 4.0), http://creativecommons.org/licenses/by/4.0/, and PUB

3. Individual Textual Profiles of Hillary Clinton and Donald Trump

Creator:: Kvítková, Alena
Publisher:: Charles University, Faculty of Arts, Department of English Language and ELT Methodology
Type:: text and corpus
Subject:: idiolect, individual textual profile, Clinton, Trump, corpus, presidential debates, American president, candidates, Democrats, and Republicans
Language:: English
Description:: This corpus consists of full transcriptions of both Democratic and Republican 2016 presidential candidate debates, with a special focus on the idiolects of Hillary Clinton and Donald Trump against the background of the speeches of other candidates for the post of president of the United States. The transcriptions are sourced from the American Presidency Project at the University of California, Santa Barbara. Any use of the material requires a prior and explicit written permission by the project administrator (contact policy@ucsb.edu). This corpus material is now being shared with their kindly permission.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

4. Nottinghamer Korpus Deutscher YouTube-Sprache (The NottDeuYTSch Corpus)

Creator:: Cotgrove, Louis Alexander
Publisher:: University of Nottingham
Type:: text and corpus
Subject:: youth language, Computer-Mediated Communication, Digitally-Mediated Communication, CMC, DMC, online, YouTube, digital, emoji, translanguaging, multilingualism, and social media
Language:: German, English, Russian, Turkish, and Serbo-Croatian
Description:: The NottDeuYTSch corpus contains over 33 million words taken from approximately 3 million YouTube comments from videos published between 2008 to 2018 targeted at a young, German-speaking demographic and represents an authentic language snapshot of young German speakers. The corpus was proportionally sampled based on video category and year from a database of 112 popular German-speaking YouTube channels in the DACH region for optimal representativeness and balance and contains a considerable amount of associated metadata for each comment that enable further longitudinal cross-sectional analyses.
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

5. Nottinghamer Korpus Deutscher YouTube-Sprache (The NottDeuYTSch Corpus) (2022-07-27)

Creator:: Cotgrove, Louis Alexander
Publisher:: University of Nottingham
Type:: text and corpus
Subject:: youth language, Computer-Mediated Communication, Digitally-Mediated Communication, CMC, DMC, online, YouTube, digital, emoji, translanguaging, multilingualism, social media, digital humanities, and Web corpus
Language:: German, English, Russian, Turkish, and Serbo-Croatian
Description:: The NottDeuYTSch corpus contains over 33 million words taken from approximately 3 million YouTube comments from videos published between 2008 to 2018 targeted at a young, German-speaking demographic and represents an authentic language snapshot of young German speakers. The corpus was proportionally sampled based on video category and year from a database of 112 popular German-speaking YouTube channels in the DACH region for optimal representativeness and balance and contains a considerable amount of associated metadata for each comment that enable further longitudinal cross-sectional analyses.
Rights:: Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0), http://creativecommons.org/licenses/by-nc/4.0/, and PUB

6. Video699: lecture recordings and lecture materials

Creator:: Novotný, Vít
Publisher:: Faculty of Informatics, Masaryk University
Type:: video and corpus
Subject:: information retrieval, video, image, and XML
Language:: English and Czech
Description:: This is an XML dataset of 17 lecture recordings randomly sampled from the lectures recorded at the Faculty of Informatics, Brno, Czechia during 2010–2016. We drew a stratified sample of up to 25 video frames from each recording. In each video frame, we annotated lit projection screens and their condition. For each lit projection screen, we annotated lecture materials shown in the screen. The dataset contains 699 projection screen annotations, and 925 lecture materials.
Rights:: Open Data Commons Open Database License (ODbL), http://opendatacommons.org/licenses/odbl/summary/, and PUB

1. Annotated Corpus of Czech Case Law for Reference Recognition Tasks

2. Annotated Corpus of Czech Case Law for Reference Recognition Tasks (2019-06-25)

3. Individual Textual Profiles of Hillary Clinton and Donald Trump

4. Nottinghamer Korpus Deutscher YouTube-Sprache (The NottDeuYTSch Corpus)

5. Nottinghamer Korpus Deutscher YouTube-Sprache (The NottDeuYTSch Corpus) (2022-07-27)

6. Video699: lecture recordings and lecture materials

Limit your search

Show values starting with

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Creator

Show values starting with

Language

Publisher

Rights

Subject

Show values starting with

Type

Date

Original context has metadata only

Harvested from