Data · dataset · 2026
2022 NIST Language Recognition Evaluation Test and Development Sets
Listed in UC Berkeley Library Dataverse
2022 NIST Language Recognition Evaluation Test and Development Sets was developed by the Linguistic Data Consortium (LDC) and the National Institute of Standards and Technology (NIST).
Description
This release contains the test and development data, metadata, answer keys, and documentation for the 2022 NIST Language Recognition Evaluation (LRE22). The source speech data is comprised of approximately 222 hours of conversational telephone speech (CTS) and broadcast narrowband speech (BNBS) in 14 languages: Afrikaans, Tunisian Arabic, Algerian Arabic, Libyan Arabic, South African English, Indian-accented South African English, North African French, Ndebele, Oromo, Tigrinya, Tsonga, Venda, Xhosa and Zulu.
The goals of NIST's Language Recognition Evaluation are to advance language recognition technologies, to facilitate technology development, and to measure the performance of current state-of-the-art technology. LRE22 emphasized language recognition for African languages, including low resource languages, and expanded the range of test segment durations. Further information about the 2022 evaluation can be found in the 2022 NIST Language Recognition Evaluation Plan.
Links
Where it is published
- Dataverse dataset page datasets.lib.berkeley.edu/dataset.xhtml?persistentId=doi%3A10.60503%2FD3%2FCQIMYN ↗
landing page · from datasets lib berkeley edu
- DOI doi.org/10.60503/d3/cqimyn ↗
DOI / persistent id · from datasets lib berkeley edu
Catalogue records · 1
- Dataverse API datasets.lib.berkeley.edu/api/datasets/:persistentId/?persistentId=doi%3A10.60503%2FD3%2… ↗
metadata API · from datasets lib berkeley edu
Topics
- Stated by source
- Social Sciences
- From keywords
- Social Science
- Inferred from text
- Audio 65%
Provenance · 1 source records, 9 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| UC Berkeley Library Dataverse | doi:10.60503/D3/CQIMYN | 4 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| concepts[field].dataverse_subject:social-sciences | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /subjects |
| concepts[field].local:field:social-science | mapping · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /subjects |
| concepts[modality].local:modality:audio | enrichment · datasets lib berkeley edu | keyword-concept-rules@1.0.0 | title+description (65%) |
| created_date | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | |
| description | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /description |
| publication_date | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | |
| title | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | /name |
| updated_date | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 | |
| version_label | source · datasets lib berkeley edu | connector:datasets_lib_berkeley_edu@1.0.0 |