Data · dataset · 2025
Predicting Invasiveness of Lung Adenocarcinoma from Chest CT with Few-shot Vision-Language Ternary Classification Model
Listed in Teesside University Research Data Repository
This dataset contains the research data used in the study “Predicting Invasiveness of Lung Adenocarcinoma from Chest CT with Few-shot Vision–Language Ternary Classification Model.” It includes data from 848 patients with pathologically confirmed lung adenocarcinoma collected across four medical centers.
Description
The dataset supports a study evaluating the GPT-4o vision–language model for ternary classification of pure ground-glass nodules (pGGNs).
The input data for the GPT-4o model are provided in MP4 format and organized into three folders according to pathological subtype: preinvasive lesions (MP4_PRE; n = 333), minimally invasive adenocarcinomas (MP4_MIA; n = 376), and invasive adenocarcinomas (MP4_IAC; n = 139). To promote transparency and reproducibility, the dataset also includes two supplementary scripts, "dicm_to_nii.py" and "nii_to_mp4.py", which detail the anonymization and data conversion processes used in this study.
Read the rest (2 more)
These scripts demonstrate the step-by-step transformation from the original DICOM-format CT images to anonymized NIfTI (.nii.gz) files and subsequently to MP4-format videos used as model inputs. This workflow provides researchers with a clear reference for ensuring patient privacy protection when applying online vision–language models to medical imaging data. Due to Mendeley Data’s maximum storage capacity of 10 GB, we uploaded all video data used as inputs for the vision–language models (GPT-4o, Google Gemini 2.5 Pro, and Molmo), which together occupy 9.94 GB of space.
Accordingly, this dataset contains only the anonymized video data used for model analysis. The CT images of all patients (totaling 81.1GB) will be disclosed in other databases that can provide the corresponding capacity. Users of this dataset please cite the following publication: “Predicting Invasiveness of Lung Adenocarcinoma from Chest CT with Few-shot Vision–Language Ternary Classification Model.”
Links
Where it is published
- DOI doi.org/10.17632/h7b4ryrbzw.1 ↗
DOI / persistent id · from researchdata tees ac uk
Catalogue records · 1
- OAI-PMH record data.mendeley.com/oai?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai%3Adata… ↗
metadata API · from researchdata tees ac uk
Topics
- From keywords
- Artificial intelligence · Computer Science & AI · Earth & Environmental Science · Engineering · Humanities · Life Sciences · Social Science
- Inferred from text
- Computed tomography 50% · Image 65% · Imaging 75% · Video 75%
Provenance · 1 source records, 15 field assertions
| Source | Key | Last seen | Raw |
|---|---|---|---|
| Teesside University Research Data Repository | oai:data.mendeley.com/h7b4ryrbzw.1 | 8 d ago | JSON v1 |
| Field | Assertion | Extractor | Evidence |
|---|---|---|---|
| concepts[field].anzsrc:group:4602 | mapping · researchdata tees ac uk | vocabulary-mapper@1.0.0 | keywords['Artificial Intelligence'] |
| concepts[field].local:field:computer-science-ai | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:earth-environmental | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:engineering | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:humanities | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:life-sciences | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[field].local:field:social-science | mapping · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| concepts[modality].local:modality:ct | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (50%) |
| concepts[modality].local:modality:image | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (65%) |
| concepts[modality].local:modality:imaging | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| concepts[modality].local:modality:video | enrichment · researchdata tees ac uk | keyword-concept-rules@1.0.0 | title+description (75%) |
| description | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/description |
| license_text | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| publication_date | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | |
| title | source · researchdata tees ac uk | connector:researchdata_tees_ac_uk@1.0.0 | /metadata/dc/title |