Text · dataset · 2022
ScandiQA
Listed in Hugging Face Datasets
ScandiQA is a dataset of questions and answers in the Danish, Norwegian, and Swedish languages.
Description
All samples come from the Natural Questions (NQ) dataset, which is a large question answering dataset from Google searches. The Scandinavian questions and answers come from the MKQA dataset, where 10,000 NQ samples were manually translated into, among others, Danish, Norwegian, and Swedish.
However, this did not include a translated context, hindering the training of extractive question answering models. We merged the NQ dataset with the MKQA dataset, and extracted contexts as either "long answers" from the NQ dataset, being the paragraph in which the answer was found, or otherwise we extract the context by locating the paragraphs which have the largest cosine similarity to the question, and which contains the desired answer.
Read the rest (3 more)
Further, many answers in the MKQA dataset were "language normalised": for instance, all date answers were converted to the format "YYYY-MM-DD", meaning that in most cases these answers are not appearing in any paragraphs. We solve this by extending the MKQA answers with plausible "answer candidates", being slight perturbations or translations of the answer. With the contexts extracted, we translated these to Danish, Swedish and Norwegian using the DeepL translation service for Danish and Swedish, and the Google Translation service for Norwegian.
After translation we ensured that the Scandinavian answers do indeed occur in the translated contexts. As we are filtering the MKQA samples at both the "merging stage" and the "translation stage", we are not able to fully convert the 10,000 samples to the Scandinavian languages, and instead get roughly 8,000 samples per language. These have further been split into a training, validation and test split, with the former two containing roughly 750 samples.
The splits have been created in such a way that the proportion of samples without an answer is roughly the same in each split.
Links
Where it is published
- Hugging Face dataset page huggingface.co/datasets/alexandrainst/scandi-qa ↗
landing page · from Hugging Face
- DOI doi.org/10.57967/hf/6061 ↗
DOI / persistent id · from Hugging Face
Catalogue records · 1
- Hub API huggingface.co/api/datasets/alexandrainst/scandi-qa ↗
metadata API · from Hugging Face
Topics
- Stated by source
- question answering · text
- From keywords
- Computer Science & AI
Provenance · 1 source records, 10 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Hugging Face Datasets | alexandrainst/scandi-qa | 12 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · Hugging Face | connector:huggingface@1.0.0 | /gated |
| concepts[field].local:field:computer-science-ai | mapping · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[modality].hf_modality:text | source · Hugging Face | connector:huggingface@1.0.0 | |
| concepts[task].hf_task:question-answering | source · Hugging Face | connector:huggingface@1.0.0 | /tags[task_categories:*] |
| created_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| description | source · Hugging Face | connector:huggingface@1.0.0 | /description |
| license | source · Hugging Face | connector:huggingface@1.0.0 | /tags[license:*] |
| publication_date | source · Hugging Face | connector:huggingface@1.0.0 | |
| title | source · Hugging Face | connector:huggingface@1.0.0 | /id |
| updated_date | source · Hugging Face | connector:huggingface@1.0.0 |