Data · dataset · 2026
Hybrid Forgery Audio Dataset
Listed in ScienceDB
The hybrid forgery dataset contains two types of samples: traditional editing tampering and partial deepfake.
Description
The editing tampering portion is constructed based on the AISHELL-3 speech corpus, comprising speech samples from 174 different speakers. This dataset is built using the professional audio editing software CoolEdit Pro through manual cutting and splicing.
The specific tampering forms include segment replacement and insertion. The length of all tampered segments is controlled between 0.1 seconds and 3 seconds, covering various tampering scenarios ranging from extremely short transient anomalies to longer semantically inconsistent segments. The final generated audio samples have a total length ranging from 3 seconds to 8 seconds, consisting of 4,341 training samples, 1,600 validation samples, and 3,248 test samples.
Read the rest (2 more)
All samples have been manually verified and boundary-annotated.The deepfake samples are generated using speech synthesis technology based on Global Style Tokens (GST), simulating the forgery traces of algorithms such as speech synthesis and voice conversion. The length of both types of forged segments is strictly controlled within the range of 0.1 seconds to 3 seconds, and they are randomly spliced with real audio. The final total audio length is also constrained between 3 seconds and 8 seconds.
This dataset comprises 5,150 training samples, 2,000 validation samples, and 4,800 test samples, covering hybrid scenarios with multiple samples and multiple forgery types from the same set of 174 speakers.
Links
Where it is published
- DOI doi.org/10.57760/sciencedb.31778 ↗
DOI / persistent id · from scidb cn
Catalogue records · 1
- OAI-PMH record scidb.cn/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=10.57760%2… ↗
metadata API · from scidb cn
Topics
- From keywords
- Earth & Environmental Science · Engineering · Humanities · Life Sciences · Social Science
- Inferred from text
- Audio 75%
Provenance · 1 source records, 11 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| ScienceDB | 10.57760/sciencedb.31778 | 9 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| access_level | source · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:engineering | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:humanities | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[field].local:field:social-science | mapping · scidb cn | connector:scidb_cn@1.0.0 | |
| concepts[modality].local:modality:audio | enrichment · scidb cn | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · scidb cn | connector:scidb_cn@1.0.0 | /metadata/dc/description |
| license | source · scidb cn | connector:scidb_cn@1.0.0 | /metadata/dc/rights |
| publication_date | source · scidb cn | connector:scidb_cn@1.0.0 | |
| title | source · scidb cn | connector:scidb_cn@1.0.0 | /metadata/dc/title |