Data · dataset · 2023
`trec-arabic`
Listed in Hugging Face Datasets
Description
Dataset Card for trec-arabic The trec-arabic dataset, provided by the ir-datasets package. For more information about the dataset, see the documentation. Data This dataset provides: docs (documents, i.e., the corpus); count=383,872 This dataset is used by: trec-arabic_ar2001, trec-arabic_ar2002 Usage from datasets import load_dataset docs = load_dataset('irds/trec-arabic', 'docs') for record in docs: record # {'doc_id': ..., 'text': ..., 'marked_up_doc':… See the full description on the dataset page: huggingface.co/datasets/irds/trec-arabic.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/irds/trec-arabic ↗
landing page · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/irds/trec-arabic ↗
metadata API · from Hugging Face
Topics
- Stated by source
- text retrieval
- From keywords
- Computer Science & AI
- Inferred from text
- Text 75%
Provenance · 1 source records, 9 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | irds/trec-arabic | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].local:modality:text | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[task].hf_task:text-retrieval | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |