Constarium
← Search

Data · dataset · 2026

Ali Data — Urdu & Roman Urdu Corpus

Listed in Hugging Face Datasets

Description

Ali Data — Urdu & Roman Urdu Corpus A large public-domain/open collection of Urdu (اردو) and Roman Urdu text for training Urdu language models. Stats 21,846,033 unique rows (deduplicated by exact text) Format: JSON Lines — each line: {"text": "...", "source": "..."} Scripts: Urdu (Arabic script) + Roman Urdu (Latin script) Sources Source Rows (approx) Hugging Face open datasets (news, QA, sentiment, translation, poetry, instructions)… See the full description on the dataset page: huggingface.co/datasets/PAILLM/ali-data.

Links

Where it is published

Catalogue records · 1

Topics

Stated by source
fill mask · text generation
Inferred from text
Text 75%
Provenance · 1 source records, 11 field assertions
SourceKeyLast seenRaw
Hugging Face DatasetsPAILLM/ali-data7 d agoJSON v1
FieldAssertionExtractorEvidence
access_levelsource · Hugging Faceconnector:huggingface@1.0.0/gated
concepts[field].local:field:computer-science-aimapping · Hugging Faceconnector:huggingface@1.0.0
concepts[modality].local:modality:textenrichment · Hugging Facekeyword-concept-rules@1.0.0title+description (75%)
concepts[task].hf_task:fill-masksource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
concepts[task].hf_task:text-generationsource · Hugging Faceconnector:huggingface@1.0.0/tags[task_categories:*]
created_datesource · Hugging Faceconnector:huggingface@1.0.0
descriptionsource · Hugging Faceconnector:huggingface@1.0.0/description
license_textsource · Hugging Faceconnector:huggingface@1.0.0
publication_datesource · Hugging Faceconnector:huggingface@1.0.0
titlesource · Hugging Faceconnector:huggingface@1.0.0/id
updated_datesource · Hugging Faceconnector:huggingface@1.0.0