figshare + Loughborough Research Repository + GRANTS Data + UP Research Data Repository2026 · Astronomical catalogue
Raw data<p dir="ltr">This PhD project investigates the significance of within-formant measurements for the vowels [i:], [ɪ], [e], [ə], [a:], [o], [u:], and [ʊ], for forensic speaker comparison. It contains six traditional PhD thesis chapters providing background information, as well as three research articles presenting analyses. Data was sourced from the Marwari language, spoken in Rajasthan, India, as a
ZivaHub + Deakin Research Online + DMU Figshare2026 · dataset · unknown
The stylistics of popular song lyricsBased on a specifically designed 1291-song corpus of every English-language UK number one single between January 1960 and December 2019, this thesis explores the linguistic style of popular song lyrics (PSL) by reference to previous research and cultural intuitions about the register. Using Antconc and Wmatrix corpus tools, it aims to provide a deeper understanding of the style of PSL, looking at
ZivaHub + Deakin Research Online + DMU Figshare + HKU DataHub + Swinburne Figshare + DaYta Ya Rona + SUNScholarData + figshare + Loughborough Research Repository + GRANTS Data + UP Research Data Repository2026 · Astronomical catalogue
Metaphoric Verb_Noun Collocations_Corpus Data<p dir="ltr">The shared dataset presents the Phase 1 outcome of a lexico-semantic analysis of metaphoric verb-noun collocations in Croatian. The end goal of the project is to compile a publicly accessible and searchable Repository of Affective and Lexico-semantic Norms for Croatian Metaphors (DigiMet). </p>
ZivaHub2026 · dataset
Oficina de Audiodescrição Didática de Imagens Estáticas: Abordagem prática e o uso de Inteligência Artificial Generativa<p dir="ltr">Material apresentado durante a segunda semana de Capacitação do Sistema de Bibliotecas da Universidade Federal do Ceará.<br><br></p>
ZivaHub2026 · dataset
Addendum to «Critical Analysis and Correction Plan for the ARC-22 Model»: Reproducibility Test Results<p dir="ltr">This addendum reports an independent reproducibility test of the ARC-22 model, based on the specification published in the critical paper (Kuul, 2026, DOI: 10.6084/m9.figshare.34013346). A Python parser was implemented according to the published rules and run on the ZL3b Voynich corpus (33,206 tokens) and the Timm–Schinner generated text (10,809 tokens). Three parser iterations were p
ZivaHub2026 · dataset
ReenKuul/arc22 — ARC-22 parser, code release (2026-09-28)<p dir="ltr">Python implementation of the ARC-22 parser for the Voynich Manuscript (EVA/ZL transcription). Includes compliance rate (CR), bigram entropy (H2, BPC), and LZ77 metrics. Two independent runs produced identical results. The reported CR values (44.38% Voynich, 50.50% generated) were not reproduced from the published specification. The model requires substantial revision and rebuilding. C
HKU DataHub + figshare + Loughborough Research Repository2026 · Astronomical catalogue
Shanghainese Word Frequency (v1.0): a spoken-weighted frequency list with rank uncertainty<p dir="ltr">A word frequency list for Shanghainese (Shanghai Wu, ISO 639-3 wuu): 2,524 ranked words, each with counts in conversational speech and in written Wu, a combined frequency per million (0.75 × spoken + 0.25 × written), document range, Gries' DP, and a bootstrap 95% rank interval.</p><p dir="ltr">Limits. The spoken corpus is small: 20 speakers, about 4 hours, 52,412 tokens. Ranks are rel
figshare + Loughborough Research Repository2026 · Astronomical catalogue
Making sense of artificial intelligence disruption in translation<p dir="ltr">research data</p>
figshare + Loughborough Research Repository2026 · Astronomical catalogue
Topic-emotion co-evolution in sixteen years of Wuhan rainstorm discourse on Weibo (2010-2025): topic assignments and aggregated tables<p>Topic assignments, topic taxonomy and aggregated topic x emotion tables of a study of topic-emotion co-evolution in sixteen years of Wuhan rainstorm discourse on Weibo (2010-2025). This record builds on the WuRain-Senti corpus record (<a href="https://doi.org/10.6084/m9.figshare.33302052">https://doi.org/10.6084/m9.figshare.33302052</a>), which carries the relevance and emotion labels themselve
figshare2026 · Astronomical catalogue
Keats Archaic Lexicon Corpus<p dir="ltr">This analysis-ready dataset supports the study of archaic vocabulary and semantic patterning in John Keats's poetry. It contains 135 retained works grouped into analysis years 1814-1819, with 106,937 tokens and an archaic subset of 2,433 occurrences representing 293 normalized forms. The deposit includes work-level plain texts, complete token data, a lexical frequency list, contextual
Teesside University Research Data Repository2026 · dataset
LLM-based proficiency assignment for learner corporaThis dataset accompanies the study on LLM-based proficiency assignment for learner corpora. The study addresses the lack of comprehensive proficiency assessment data in learner corpus research and tests whether large language models (LLMs) can provide useful proficiency information for large collections of learner writing. Two approaches were investigated: direct assignment of Common European Fram
IISH Dataverse2026 · dataset · unknown
TINTIN Corpus: Panel schemasData related to panel schemas annotated in the TINTIN Corpus with the schemes of Panel Templates, Perspective Taking, and Narrative Grammar.
TUDOdata2026 · dataset · unknown
Dataset Discourse Management Constructions in Wikipedia Talk PagesThis dataset forms the basis to the paper: Gillmann, M. (2024). Allostructions and stancetaking: a corpus study of the German discourse management constructions Wo/wenn wir gerade/schon dabei sind. Cognitive Linguistics, 35(1), 67-107. https://doi.org/10.1515/cog-2020-0117 Drawing on a corpus study of Wikipedia Talk pages, the paper presents a case study of German discourse management markers such
IISH Dataverse + ODISSEI Portal2026 · dataset · unknown
TINTIN Corpus: Panel ComplexityThis dataset reports annotation date from the TINTIN Corpus related to panel complexity, which is a metric aggregating classifications of Framing, Backgrounds, and Entities per Panel
Hugging Face Datasets2026 · Table · Parquet
ThaiTreesThaiTrees A 342M-token corpus of Thai drawn from news, Wikipedia, spoken transcripts and social media, automatically parsed under the Universal Dependencies framework. It is released as three artefacts: a raw text corpus, a frequency lexicon, and a dependency-parsed corpus in CoNLL-U. Dataset Summary ThaiTrees contains 341,967,133 tokens across 366,120 documents in four domains (news, Wikipedia, s
REDU - Unicamp Institutional Research Data Repository2026 · dataset · unknown
BrPoliCorpus: Brazilian political corpusThis is the version 1.0 of the package BrPoliCorpus (Brazilian Political Corpus). It is intended to be a free repository of open data regarding official documents of Brazilian Politics. For the current time, the following datasets are available: Inaugural Speeches: A set of Brazilian President's Inaugural Speeches. Updated until 01/01/2023 Parliamentary Floor: A set of all parliamentary discourses
Borealis2025 · dataset · unknown
Data in support of: A’ faighneachd Mhic-Talla [Asking the Echo]: A corpus-based approach to vernacular classification of Gaelic songsThis dataset is in support of the Master's thesis research entitled "A’ faighneachd Mhic-Talla [Asking the Echo]: A corpus-based approach to vernacular classification of Gaelic songs." The purpose of this research was to identify vernacular classifications of Gaelic songs. The original data for this study consisted of textual data from the all-Gaelic newspaper Mac-Talla (1892-1904) and was sourced
Teesside University Research Data Repository2024 · dataset
Formas pronominales de tratamiento en el español actual de Guatemala (Pronominal Forms of Address in Contemporary Guatemalan Spanish)The research analyzes the pronominal forms of address (PFA) in Guatemalan Spanish, considering sociolinguistic, stylistic, and dialectological aspects. The main objective is to identify the national-cultural specificity in the use of PFAs in Guatemala. The research is based on transcripts of recordings made in different regions of the country, interviews, and surveys with speakers from various soc
Borealis2024 · dataset · unknown
Read_Me Gess 2006 Maillardville corpus speaker informationReadme file for Gess 2006 Maillardville corpus speaker information
Teesside University Research Data Repository2023 · dataset
St. George's University List (SGUL)This dataset contains all of the files and data that we accessed and used to create out initial local (medical) academic word list (L-AWL). We cannot include the actual corpora in this dataset because many files (e.g., lecture PPT slides) were copyrighted by instructors and contain identifying information. The dataset is broken up into steps, and each step has a folder with data, supporting files,
IISH Dataverse2023 · dataset · unknown
Running through the Who, Where, and When - Data and PublicationUnderstanding visual narratives requires readers to track dimensions of time, spatial location, and characters across a sequence. Previous work has found situational changes across adjacent panels differ cross-culturally, but few works have examined such situational dimensions across extended sequences. We therefore investigated situational ‘runs’—uninterrupted sequences of the situational dimensi
IISH Dataverse2023 · dataset · unknown
Linguistic Typology of Motion Events in Visual Narratives - Data and PublicationLanguages use different strategies to encode motion. Some use particles or “satellites” to describe a path of motion (Satellite-framed or S-languages like English), while others typically use the main verb to convey the path information (Verb-framed or V-languages like French). We here ask: might this linguistic variation lead to differences in the way paths are depicted in visual narratives like
Borealis2022 · dataset · unknown
Read_Me Gess 2006 Maillardville corpus Free ConversationReadme file for Gess 2006 Maillardville corpus Free Conversation
Borealis2022 · dataset · unknown
Read_Me Gess 2006 Maillardville corpus Free Conversation transcriptionsReadme file for Gess 2006 Maillardville corpus Free Conversation.
Borealis2022 · dataset · unknown
Gess 2006 Maillardville corpus Free Conversation transcriptionsTranscriptions to accompany Gess 2006 Maillardville Free Conversation recordings.