Hugging Face Datasets2026 · Text
LiveChatBenchLiveChatBench (v1.1) Korean → English translation benchmark built from Korean live-streaming chat messages. Each row pairs a Korean chat message with an English reference translation. Where a message relies on streaming slang, community jargon, or a streamer/game name, a short background note explains the term; it is an empty string when no note is needed. Fields Field Type Description background
Hugging Face Datasets2026 · Table · Parquet
Tutlait v1: Tamazight Speech with Arabic TranslationsTutlait v1: Tamazight Speech with Arabic Translations ~21 hours of spoken Tamazight (Amazigh / Berber) from 118 speakers, each clip paired with an Arabic (MSA) translation. It is built for speech-to-text translation (Tamazight audio → Arabic text) and for fine-tuning speech models such as Whisper on a low-resource language. Note: the text column is an Arabic translation of what was said, not a tra
Hugging Face Datasets2026 · Table · CSV · gated
Tiberian Biblical Hebrew → IPA (BHS-aligned)Tiberian Biblical Hebrew → IPA (BHS-aligned) Org: tiberianaiAuthors: John Locke, Teodor BorsCompanion model: tiberianai/tiberian-hebrew-ipa-byt5 Verse-level parallel corpus for distilling Masoretic Biblical Hebrew (niqqud and teʿamim) into Tiberian IPA, using Geoffrey Khan’s reconstruction of the Tiberian pronunciation tradition as encoded by our rule-based teacher. Gated (manual approval). The ca
Hugging Face Datasets2026 · Table · Parquet
Uzbek-law-LM: Authoritative Legal Corpus & SFT BenchmarkUzbek-law-LM: Authoritative Legal Corpus & SFT Benchmark Uzbek-law-LM is an engineered, legally grounded, and decontaminated Uzbek legal corpus designed for pre-training, continued pre-training, and supervised fine-tuning (SFT) of large language models serving the Republic of Uzbekistan's legislative, judicial, executive, and civic sectors. It provides a single, cohesive, authoritative legal resou
Hugging Face Datasets2026 · Text
The Bible in 1,004 LanguagesThe Bible in 1,004 Languages 14,497,397 verses across 1,253 translations in 1,004 languages, every verse keyed to the same chapter-and-verse address so that any two languages can be aligned by joining on book, chapter and verse. The Bible is the most widely translated text in existence, and for several hundred of the languages here it is the largest — sometimes the only — substantial digitised tex
Hugging Face Datasets2026 · dataset · gated
SLANGUAGE Tiv Parallel Corpus v1SLANGUAGE Tiv Parallel Corpus v1 Dataset Description The SLANGUAGE Tiv Parallel Corpus is the first commercially developed parallel English–Tiv text dataset, created by STECH Global Limited through the SAVA Languages (SLANGUAGE) project. Tiv is a tonal Bantoid language spoken by approximately 8 million people, primarily in Benue State, Nigeria. Despite its significant speaker population, Tiv remai
Hugging Face Datasets2026 · Table · Parquet
Vāgartha — Sanskrit Verses with Structured Explanationsवागर्थ · Vāgartha वागर्थाविव संपृक्तौ वागर्थप्रतिपत्तये ।जगतः पितरौ वन्दे पार्वतीपरमेश्वरौ ॥ "United as word and meaning are united, I bow to the parents of the world,Pārvatī and Parameśvara, that I may attain an understanding of word and meaning." — Kālidāsa, Raghuvaṃśa 1.1 Vāgartha — vāk (word) and artha (meaning) — is a corpus of 217,959 Sanskrit verses, each paired with a detailed, structured
Hugging Face Datasets2026 · Text
Last Translation BenchmarkLast Translation Benchmark Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. Standard benchmarks for machine translation evaluation are often either trivial (having few authentic mistakes) or unrealistic (overly synthetically contrived). Furthermore, automatic translation metrics become
Hugging Face Datasets2026 · Text · gated
ICON26:COILD-INDIC-MT — Dogri-PunjabiCOILD-INDIC-MT 2026 — Dogri–Punjabi Dataset This dataset is provided for the COILD-INDIC-MT 2026 Shared Task, co-located with ICON 2026. The shared task aims to foster research and innovation in Natural Language Processing (NLP) for Indian Languages. This repository contains data specifically for the: Dogri ↔ Punjabi language pair. 🔐 Access to the Dataset This is a restricted and gated dataset. Ac
Hugging Face Datasets2026 · Text
Entropy-Valley DatasetsEntropy-Valley Datasets 📄 Paper (arXiv:2608.22274) | 💻 GitHub | 🤗 Models This repository contains every data file read by Entropy-Valley (EV), the training-free target-length selector for masked diffusion machine translation introduced in "Length-Adaptive Decoding for Masked Diffusion Machine Translation" (EMNLP 2026 Main Conference). Masked diffusion language models decode by filling a fixed-size
Hugging Face Datasets2026 · Text
qvac/TranslatePsy-AfriSLM-Synthetic-MixTranslatePsy-AfriSLM Synthetic Mix TranslatePsy-AfriSLM Synthetic Mix is a quality-filtered synthetic parallel corpus for machine translation between English and 19 Sub-Saharan African languages. It contains 215,653,192 bidirectional training examples and was selected as the primary African translation component used to post-train the TranslatePsy-AfriSLM model family. The dataset accompanies the
Hugging Face Datasets2026 · Table · Parquet
Syrzyk Dataset🇺🇦 SyrzykAI Dataset: Паралельний корпус «Суржик ↔ Літературна українська мова» Опис датасету (Dataset Summary) Syrzyk Dataset — це ретельно верифікований паралельний корпус речень для задачі нормалізації тексту, виправлення русизмів та автоматичного виправлення суржику на нормативну літературну українську мову. Датасет створено в рамках проекту SyrzykAI для fine-tuning Seq2Seq моделей (в першу чер
Hugging Face Datasets2026 · Table · Parquet
Harbour MiniGUI Extended Code DatasetDataset Card for Harbour MiniGUI Extended / Dataset de Código Harbour MiniGUI Extended 🇬🇧 Dataset Summary This dataset contains 3,185 complete code examples written in Harbour, specifically focused on the MiniGUI Extended (HMG Extended) framework. Harbour is a modern, open-source implementation of the Clipper/xBase language, widely used for business and desktop applications. MiniGUI Extended is on
Hugging Face Datasets2026 · Table · Parquet
HAT: Hallucination Annotation for TranslationHAT: Hallucination Annotation for Translation 🧭 Table of Contents Overview Usage Data Creation Process Data Statistics Dataset Structure Paper Abstract Citation License 📘 Overview HAT (Hallucination Annotation for Translation) is a large-scale dataset for hallucination detection in machine translation (MT).It is released as part of our publication at ACL 2026 (paper). 350,959 span-level annotated
Hugging Face Datasets2026 · Table · Parquet
espnet/yodas3YODAS v3 Paper YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 a
Hugging Face Datasets2026 · Table · Parquet
PangeanicYueJa - Cantonese Japanese Parallel CorpusPangeanicYueJa - Cantonese Japanese Parallel Corpus PangeanicYueJa is a Cantonese-Japanese parallel corpus designed for machine translation, multilingual large language model (LLM) training, cross-lingual NLP research, retrieval-augmented generation (RAG), bilingual embeddings, instruction tuning, and multilingual AI systems. This release contains 55,000 Cantonese-Japanese sentence pairs sampled f
Hugging Face Datasets2026 · dataset
SignLink ASL Video Dictionary MatrixSignLink ASL Video Dictionary Matrix This dataset contains a processed matrix of American Sign Language (ASL) video clips used as the core dictionary lookup engine for SignLink, an offline voice-to-sign inference pipeline. 🤝 Attribution & Data Source The video assets in this dataset are sourced directly from the Sign-Language-Mocap-Archive created by StudioGalt. Original Creator: StudioGalt Source
Hugging Face Datasets2026 · Text
LingReason Chintang DataDataset Description This dataset contains Chintang data for the LingReason project, which is generated from using the code released in the LingReason GitHub repository. This dataset accompanies the paper Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?. Data Splits Split File Examples Description test_icl ctn_test_icl.json 344 Test set wit
Hugging Face Datasets2026 · Image
thuvienphapluat.vn /tnpl/ — Vietnamese Legal Terminology (bilingual VI<->EN, common-corpus edition)thuvienphapluat.vn /tnpl/ — Vietnamese Legal Terminology (bilingual VI ↔ EN) 🇻🇳 Tóm tắt. Bản thu thập đầy đủ chuyên mục Thuật ngữ pháp lý của THƯ VIỆN PHÁP LUẬT — mỗi dòng là một thuật ngữ pháp lý kèm định nghĩa tiếng Việt, lĩnh vực pháp luật, tình trạng hiệu lực, lịch sử cập nhật và tham chiếu chéo. Mỗi cột tiếng Việt có một cột song song tiếng Anh dịch bằng LLM. Đây là phiên bản common-corpus: b
Hugging Face Datasets2026 · Text
VotingBookletsVotingBooklets Dataset Summary VotingBooklets is a large-scale four-language parallel corpus extracted from the complete collection of Swiss federal voting booklets (Abstimmungsbüchlein), covering federal votes from June 1977 to March 2026. It contains aligned paragraph-level segments across German (de), French (fr), Italian (it), and Romansh Grischun (rm), and serves as a resource for low-resourc
Hugging Face Datasets2026 · Table · CSV
ChameleonHugging Face Datasets2026 · Table · Parquet
yaturk-7langYaTURK-7lang This dataset was used in the research paper No One-Size-Fits-All: Building Systems For Translation to Bashkir, Kazakh, Kyrgyz, Tatar and Chuvash Using Synthetic And Original Data. Dataset used for the online competition series "Machine Translation for Low-Resource Turkic Languages". This dataset was generated via Yandex.Translate. 📢 News – March 5, 2026 The dataset has been updated! A
Hugging Face Datasets2025 · Table · Parquet
Bulgarian Corpus 33BBulgarian Corpus 33B (BC-33B) Dataset Summary The Bulgarian Corpus 33B (BC-33B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Bulgarian. Comprising approximately 33.4 Billion tokens (measured with Qwen 2.5/Llama-3 tokenizer), it represents one of the largest open-source resources for Bulgarian LLM pretraining. The dataset is engineered for a moder
Hugging Face Datasets2025 · Table · CSV
NepTam🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes
Hugging Face Datasets2025 · Table · Parquet
My-En General Domain Parallel Corpus🇲🇲-🇬🇧 Myanmar-English General Text Translation Dataset 📚 Dataset Overview This dataset is a high-quality, parallel corpus designed for training robust and accurate Myanmar-English Machine Translation (MT) models. It focuses on General Domain texts, covering a wide range of everyday scenarios, literature, conversations, and descriptive narratives. Our primary goal in creating this dataset is to pro