Data · dataset · 2022
hugginglearners/malayalam_news
Listed in Hugging Face Datasets
The AI4Bharat-IndicNLP dataset is an ongoing effort to create a collection of large-scale, general-domain corpora for Indian languages.
Description
Currently, it contains 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora.
We create news article category classification datasets for 9 languages to evaluate the embeddings. We evaluate the IndicNLP embeddings on multiple evaluation tasks.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/hugginglearners/malayalam_news ↗
landing page · from Hugging Face
Documentation and papers
- arXiv:2005.00085 arxiv.org/abs/2005.00085 ↗
publication · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/hugginglearners/malayalam_news ↗
metadata API · from Hugging Face
Topics
- From keywords
- Computer Science & AI
- Inferred from text
- Artificial intelligence 69%
Provenance · 1 source records, 8 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | hugginglearners/malayalam_news | 11 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].anzsrc:group:4602 | enrichment · Hugging Face | taxonomy-embedding@1.1.0 | title+keywords+description (69%) |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |