MMIST Lung Dataset

Lung cancer remains one of the most prevalent and lethal malignancies. The MMIST Lung Dataset is a curated, multi-center, multi-modal, and longitudinal dataset comprising 1,365 lung cancer patients.

The dataset combines data from CPTAC-LSCC, CPTAC-LUAD, TCGA-LUSC, and TCGA-LUAD, including 1,026 patients from TCGA and 339 patients from CPTAC.

It includes both Lung Squamous Cell Carcinoma (LUSC/LSCC) and Lung Adenocarcinoma (LUAD) and integrates clinical information, transcriptomics, whole-slide images (WSI), CT and PET imaging, longitudinal follow-up, and treatment information.

Human lungs

Modalities Across the Dataset

The dataset intentionally reflects real-world clinical conditions, where not all modalities are available for every patient. The table below summarizes the number of patients and missingness for each modality.

Modality Patients Missingness Alive @12 months Deceased @12 months
Clinical 1365 0% 1169 (86%) 196 (14%)
Transcriptomics 1284 6% 1107 (86%) 177 (14%)
Follow-up 880 35% 777 (88%) 103 (12%)
Chemotherapy 277 80% 245 (88%) 32 (12%)
Radiation Therapy 923 32% 804 (87%) 119 (13%)
Surgery 291 79% 250 (86%) 41 (14%)
Immunotherapy 49 96% 45 (92%) 4 (8%)
WSI 1359 0.4% 1164 (86%) 195 (14%)
CT 71 95% 58 (82%) 13 (18%)
PET 33 98% 26 (79%) 7 (21%)

Dataset Source

The curated MMIST-Lung dataset, together with the associated files and resources, is available in the official GitHub repository.

Imaging Data

The dataset contains three imaging modalities: whole-slide histopathology images, CT scans, and PET scans. Multiple imaging instances may be available for the same patient.

Imaging Modality Patients Total Images / Volumes Median per Patient
WSI 1359 5427 slides 3 slides
CT 71 482 scans 5 scans
PET 33 144 volumes 4 scans

Whole-slide images can correspond to different specimen types, including diagnostic slides (DX), tissue-side slides (TS), and bottom-side slides (BS).

CT and PET data follow a hierarchical structure in which each patient may have multiple studies, with multiple series and 3D scans. Non-diagnostic series such as Localizer, Scout, and Reconstruction series were excluded during dataset curation.

Clinical Data

Clinical data are available for all 1,365 patients and include demographic characteristics and tumor diagnosis information.

Clinical variables from CPTAC and TCGA were harmonized across cohorts. Variables with more than 70% missing data were excluded, resulting in a final set of 15 clinical features: 7 demographic variables and 8 tumor diagnosis variables.

The clinical information includes AJCC staging variables, age, ethnicity, race, gender, and smoking-related information.

Transcriptomic Data

Transcriptomic data are available for 1,284 patients.

Bulk RNA expression data were collected from cBioPortal and initially contained raw expression counts for more than 60,000 genes.

For the benchmark experiments, the transcriptomic data were normalized using log counts per million and dimensionality reduction based on highly variable genes was performed, resulting in a subset of 4,096 genes.

Longitudinal Follow-Up

Longitudinal follow-up information is available for 880 patients, comprising 1,431 follow-up records.

These records include tumor status and the number of days to each follow-up visit, enabling longitudinal and dynamic survival analyses.

Patients have a median of two follow-up visits. The median time to the first follow-up is 298 days, while the second occurs at a median of 647 days.

Treatment Data

The dataset includes longitudinal treatment information covering chemotherapy, radiation therapy, surgery, and immunotherapy.

Treatment Patients Treatment Records
Chemotherapy 277 426
Radiation Therapy 923 1337
Surgery 291 525
Immunotherapy 49 58

Treatment records include temporal information such as days to treatment start, allowing treatment events to be incorporated into longitudinal patient trajectories.

Benchmark Tasks

MMIST Lung supports several survival prediction tasks designed to evaluate multi-modal learning under realistic missing-data conditions.

Disease-Specific Survival Prediction

Disease-specific survival is predicted using information available at diagnosis, combining clinical, transcriptomic, WSI, CT, and PET modalities.

12-Month Overall Survival Prediction

The dataset supports binary prediction of whether a patient survives beyond 12 months after diagnosis. Approximately 86% of the cohort survives beyond 12 months.

Dynamic Survival Prediction

Longitudinal clinical and treatment information is used to predict disease-specific survival at the patient's last known follow-up time.

Longitudinal Hazard Prediction

The dataset also supports discrete-time survival analysis, where the probability of an event is modeled across successive time intervals as new patient information becomes available.

Benchmark Results

The following results summarize the main benchmark performance reported for the MMIST Lung dataset.

Task Best Reported Performance
Disease-Specific Survival 70.36% Balanced Accuracy
12-Month Overall Survival 65.42% Balanced Accuracy
Dynamic Disease-Specific Survival 81.73 ± 0.37% Balanced Accuracy
Longitudinal Hazard Prediction 85.59 ± 4.03% C-index

Evaluation Protocol

The dataset was divided into five patient-level cross-validation folds, stratified according to 12-month survival status.

All scans, treatment records, follow-up information, and other modalities belonging to the same patient remain in the same fold, preventing patient-level information leakage.

Balanced accuracy is used as the primary evaluation metric for binary survival prediction due to the substantial class imbalance in the cohort.

Scientific Paper

For further information about the dataset, curation process, longitudinal information, and benchmark experiments, please refer to our paper:

Rita Cordeiro Mendes, Maria Rita Fonseca Verdelho, Carlos Santiago, and Catarina Barata.
Real-World Multi-Modal and Longitudinal Lung Cancer Dataset.

How to cite us?

Cite the paper where the dataset was first published and acknowledge MMIST:

"The results shown in this work used datasets collected from MMIST: https://Multi-Modal-IST.github.io/"

Get in Touch at

ana.c.fidalgo.barata@tecnico.ulisboa.pt