Table · dataset · 2022
The Cross-lingual TRansfer Evaluation of Multilingual Encoders for Speech (XTREME-S) benchmark is a benchmark designed to evaluate speech representations across languages, tasks, domains and data regimes. It covers 102 languages from 10+ language families, 3 different domains and 4 task families: speech recognition, translation, classification and retrieval.
Listed in Hugging Face Datasets
FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark.
Description
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision.
Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: huggingface.co/datasets/google/fleurs.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/google/fleurs ↗
landing page · from Hugging Face
Documentation and papers
- arXiv:2205.12446 arxiv.org/abs/2205.12446 ↗
publication · from Hugging Face
- arXiv:2106.03193 arxiv.org/abs/2106.03193 ↗
publication · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/google/fleurs ↗
metadata API · from Hugging Face
Topics
- Stated by source
- audio · automatic speech recognition · text
- From keywords
- Computer Science & AI · Speech recognition
- Inferred from text
- Audio 65%
Related
- Possibly the same asThe Cross-lingual TRansfer Evaluation of Multilingual Encoders for Speech (XTREME-S) benchmark is a benchmark designed to evaluate speech representations across languages, tasks, domains and data regimes. It covers 102 languages from 10+ language families, 3 different domains and 4 task families: speech recognition, translation, classification and retrieval.
- Possibly the same asThe Cross-lingual TRansfer Evaluation of Multilingual Encoders for Speech (XTREME-S) benchmark is a benchmark designed to evaluate speech representations across languages, tasks, domains and data regimes. It covers 102 languages from 10+ language families, 3 different domains and 4 task families: speech recognition, translation, classification and retrieval.
Provenance · 1 source records, 12 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | google/fleurs | 7 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:audio | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].local:modality:audio | enrichment · Hugging Face | keyword-concept-rules@1.0.0 | title+description (65%) |
| concepts[task].hf_task:automatic-speech-recognition | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license | source · Hugging Face | connector:huggingface@1.0.0 | /tags[license:*] |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |