ZivaHub + Deakin Research Online + DMU Figshare + HKU DataHub + Swinburne Figshare + DaYta Ya Rona + SUNScholarData + figshare + Loughborough Research Repository + GRANTS Data + UP Research Data Repository2026 · Astronomical catalogue
Metaphoric Verb_Noun Collocations_Corpus Data<p dir="ltr">The shared dataset presents the Phase 1 outcome of a lexico-semantic analysis of metaphoric verb-noun collocations in Croatian. The end goal of the project is to compile a publicly accessible and searchable Repository of Affective and Lexico-semantic Norms for Croatian Metaphors (DigiMet). </p>
Teesside University Research Data Repository2026 · dataset
A Parallel Dataset for Standard Bangla to Seven Regional DialectsThis dataset is developed for Standard Bangla to regional Bangla dialect translation. It consists of 2,500 Standard Bangla source sentences collected from the Vashantor dataset, along with corresponding English and Banglish representations and translations into seven regional varieties: Pabna, Noakhali, Jashore, Rangpur, Mymensingh, Barishal, and Chittagong. Pabna, Noakhali, Jashore, and Rangpur c
ZivaHub + Deakin Research Online2026 · dataset
Lógica Sintético-Contextual (LSC-FSL-CALC)<p dir="ltr">Esta obra en seis volúmenes presenta los fundamentos, el desarrollo formal, la validación y las aplicaciones computacionales de la <b>Lógica Sintético-Contextual (LSC-CALC)</b>. El sistema propone un marco analítico y formal que trasciende la lógica clásica al integrar de manera sistemática el contexto dinámico, la semántica situacional y la dimensión normativa en la estructura de inf
Teesside University Research Data Repository2026 · dataset
Coded English-French-Japanese Health Translation Dataset and Reliability DataThis dataset contains the materials used in the study “Semantic Drift, Metaphor Vulnerability, and Emotional Meaning in Health Translation: A Corpus-Based Comparison of English-French and English-Japanese Subtitles.” Supplementary Data 1 contains the coded results for all 217 health-related English-French-Japanese segments analyzed in the study. Supplementary Data 2 contains the 25% validation sub
ZivaHub2026 · dataset
Addendum to «Critical Analysis and Correction Plan for the ARC-22 Model»: Reproducibility Test Results<p dir="ltr">This addendum reports an independent reproducibility test of the ARC-22 model, based on the specification published in the critical paper (Kuul, 2026, DOI: 10.6084/m9.figshare.34013346). A Python parser was implemented according to the published rules and run on the ZL3b Voynich corpus (33,206 tokens) and the Timm–Schinner generated text (10,809 tokens). Three parser iterations were p
ZivaHub2026 · dataset
ReenKuul/arc22 — ARC-22 parser, code release (2026-09-28)<p dir="ltr">Python implementation of the ARC-22 parser for the Voynich Manuscript (EVA/ZL transcription). Includes compliance rate (CR), bigram entropy (H2, BPC), and LZ77 metrics. Two independent runs produced identical results. The reported CR values (44.38% Voynich, 50.50% generated) were not reproduced from the published specification. The model requires substantial revision and rebuilding. C
HKU DataHub + figshare + Loughborough Research Repository2026 · Astronomical catalogue
Shanghainese Word Frequency (v1.0): a spoken-weighted frequency list with rank uncertainty<p dir="ltr">A word frequency list for Shanghainese (Shanghai Wu, ISO 639-3 wuu): 2,524 ranked words, each with counts in conversational speech and in written Wu, a combined frequency per million (0.75 × spoken + 0.25 × written), document range, Gries' DP, and a bootstrap 95% rank interval.</p><p dir="ltr">Limits. The spoken corpus is small: 20 speakers, about 4 hours, 52,412 tokens. Ranks are rel
figshare + Loughborough Research Repository2026 · Astronomical catalogue
Multilingual Shopping Queries to Marketplace Queries (OneFindMe, 2026)<p>941 pairs mapping how shoppers phrase a product search <b>in their own language</b> to the short English query that marketplace listings (AliExpress) are indexed under. The target side is the <b>retrieval</b> form — short, noun-first, in the words sellers use — not a literal translation.</p><p>Useful for training or evaluating query rewriting / query translation for e-commerce retrieval, studyi
figshare + Loughborough Research Repository2026 · Astronomical catalogue
Neural evidence that dogs segment the speech they hear with a humanlike consonant bias<p dir="ltr">Across many human languages, consonants carry more lexical information than vowels. During speech segmentation, humans, unlike non-human primates, rely more on consonant than vowel patterns, despite vowels’ greater acoustic saliency. To investigate whether this consonant bias is human-unique or could also emerge in other species exposed to human speech we performed non-invasive EEG in
figshare + Loughborough Research Repository2026 · Astronomical catalogue
The Objective Projection Corpus: A Bilingual Turkish-English Resource for Six-Feature Craft Annotation in Narrative Text, with a Published Reliability Audit<p dir="ltr">We describe the Objective Projection Corpus, a bilingual Turkish–English resource built to support research on a specific narrative-craft distinction: whether an emotional or informational state is declared on the surface of a text (told) or must be reconstructed by the reader from physical, non-evaluative detail (shown). The corpus centers on 500 annotated scene pairs (300 Turkish, 2
figshare2026 · Astronomical catalogue
Prototypes AI images part 1<p dir="ltr">Prototypes AI images part 1</p>
figshare2026 · Astronomical catalogue
Prototypes AI images part 2<p dir="ltr">Prototypes AI images part 2</p>
figshare2026 · Astronomical catalogue
PANACEA dataset - Heterogeneous COVID-19 ClaimsThis dataset contains a heterogeneous set of True and False COVID claims and online sources of information for each claim.
figshare2026 · Astronomical catalogue
Code, Model Checkpoints, and Sample Audio Files (YouTube Dataset)<p dir="ltr">A sample audio dataset for review purposes from our YouTube dataset. We have also released our code and model checkpoints.</p>
Teesside University Research Data Repository2026 · dataset
Darbest Dataset: Universal Dependencies Treebank for Standard Sorani KurdishThe Darbest dataset is a Universal Dependencies (UD) treebank dataset for Standard Sorani Kurdish written in the Perso-Arabic script. It contains 69,000 annotated sentences and 1,205,855 tokens collected from nine textual domains. The corpus was collected from seven Kurdish online news websites and supplemented with texts from published books. Before preprocessing, the collected corpus contained 1
Teesside University Research Data Repository2026 · dataset
AranjiyyaCorpus: A span-annotated Arabic news dataset of anglicised style, calques, and borrowings across six domainsThe Arabic neologism Aranjiyya (عَرَنْجِيَّة) is a portmanteau of ʿArabiyya (Arabic) and Inkliziyya / Faranjiyya (English/foreign), used by contemporary Arabic editors and stylists to describe Arabic prose that retains Arabic vocabulary while importing the syntactic, stylistic, semantic, or lexical structures of English. The phenomenon is pervasive in translated news, press releases, technical wri
IISH Dataverse2026 · dataset · unknown
Investigating the Effect of Linguistic Features on Persuasion in Text through Computational and Experimental ApproachesThe present study investigates how persuasive intent is reflected in the linguistic structure of texts and whether such patterns remain stable across languages. Using a corpus of legal pleadings and judicial judgments from the International Court of Justice, persuasive and non-persuasive texts were compared in English and French.
Teesside University Research Data Repository2025 · dataset
Bangla Text Paraphrase Corpus for Natural Language ProcessingThis dataset is Bangla Paraphrase Sentence Pair Dataset (BPDS), contains pairs of Bangla sentences labeled as paraphrase (same meaning) or non-paraphrase (different meaning). The data has been collected from diverse Bangla sources including books, newspapers and literature articles, covering a wide range of topics and writing styles. It is designed for research in natural language processing tasks
IISH Dataverse2025 · dataset · unknown
Tutorial Package for: Text as Data in Economic AnalysisThis tutorial package, comprising both data and code, accompanies the article and is designed primarily to allow readers to explore the various vocabulary-building methods discussed in the paper. The article discusses how to apply computational linguistics techniques to analyze largely unstructured corporate-generated text for economic analysis. As a core example, we illustrate how textual analysi
e-cienciaDatos2025 · dataset · unknown
Guía de anotación de terminología financiera (FINTERM)The creation of this document of annotation guidelines is framed in the Spanish national project CLARA-FINT. It consists of annotation guidelines that establish some indications and rules to create a dataset. The dataset is made up of financial texts from the annual reports of the main Spanish listed companies. Usually, said reports are publicly available under their respective shareholders websit
e-cienciaDatos2025 · dataset · unknown
Annotation guidelines for endometriosis and menopause in the DIGITENDER projectThe project, 'DIGITENDER: Extracción terminológica automática y corpora de dominios específicos para la visibilización de los problemas de salud relacionados con la mujer,' aims to develop an automatic term extractor focused on endometriosis and menopause to enhance the visibility of women's health issues that are often overlooked. To achieve this, a language model is trained using annotated terms
e-cienciaDatos2025 · dataset · unknown
Lexicon of endometriosis and menopause terms in Spanish and English for fine-tuning language modelsThe project, 'DIGITENDER: Extracción terminológica automática y corpora de dominios específicos para la visibilización de los problemas de salud relacionados con la mujer,' aims to develop an automatic term extractor focused on endometriosis and menopause to enhance the visibility of women's health issues that are often overlooked. To achieve this, a language model is trained using annotated terms
e-cienciaDatos2025 · dataset · unknown
Automatic discourse markers extractorThis work is framed in the Spanish national project CLARA-FINT. The aim of this task within the project was to create an automatic discourse markers extractor for Spanish. In order to do so, the first step was to apply linguistic annotation on texts containing said markers. The next step involved the use of these annotations to fine-tune a model for the extraction of discourse markers task. This p
e-cienciaDatos2025 · dataset · unknown
Discourse markers: Annotation guidelinesThis work is framed in the Spanish national project CLARA-FINT. The aim of this task within the project was to create an automatic discourse markers extractor for Spanish. In order to do so, the first step was to create these Annotation Guidelines to apply linguistic annotation on texts containing said markers. The next step involved the use of these annotations to fine-tune a model for the extrac
e-cienciaDatos2025 · dataset · unknown
Automatic financial term extractorThe creation of this dataset is framed in the Spanish national project CLARA-FINT. The aim of this task within the project was to create an automatic financial term extractor for Spanish. In order to do so, the first step was to apply linguistic annotation on texts, namely annual reports from the main Spanish listed companies in the IBEX 35 index. The next step involved the use of these annotation