Modalities Across the Dataset
The dataset intentionally reflects real-world clinical
conditions, where not all modalities are available for
every patient. The table below summarizes the number of
patients and missingness for each modality.
| Modality |
Patients |
Missingness |
Alive @12 months |
Deceased @12 months |
| Clinical |
1365 |
0% |
1169 (86%) |
196 (14%) |
| Transcriptomics |
1284 |
6% |
1107 (86%) |
177 (14%) |
| Follow-up |
880 |
35% |
777 (88%) |
103 (12%) |
| Chemotherapy |
277 |
80% |
245 (88%) |
32 (12%) |
| Radiation Therapy |
923 |
32% |
804 (87%) |
119 (13%) |
| Surgery |
291 |
79% |
250 (86%) |
41 (14%) |
| Immunotherapy |
49 |
96% |
45 (92%) |
4 (8%) |
| WSI |
1359 |
0.4% |
1164 (86%) |
195 (14%) |
| CT |
71 |
95% |
58 (82%) |
13 (18%) |
| PET |
33 |
98% |
26 (79%) |
7 (21%) |
The curated MMIST-Lung dataset, together with the associated
files and resources, is available in the official GitHub repository.
The dataset contains three imaging modalities:
whole-slide histopathology images, CT scans, and PET scans.
Multiple imaging instances may be available for the same patient.
| Imaging Modality |
Patients |
Total Images / Volumes |
Median per Patient |
| WSI |
1359 |
5427 slides |
3 slides |
| CT |
71 |
482 scans |
5 scans |
| PET |
33 |
144 volumes |
4 scans |
Whole-slide images can correspond to different specimen
types, including diagnostic slides (DX), tissue-side slides
(TS), and bottom-side slides (BS).
CT and PET data follow a hierarchical structure in which
each patient may have multiple studies, with multiple series
and 3D scans. Non-diagnostic series such as Localizer,
Scout, and Reconstruction series were excluded during
dataset curation.
Clinical data are available for all 1,365 patients and
include demographic characteristics and tumor diagnosis
information.
Clinical variables from CPTAC and TCGA were harmonized
across cohorts. Variables with more than 70% missing data
were excluded, resulting in a final set of
15 clinical features:
7 demographic variables and 8 tumor diagnosis variables.
The clinical information includes AJCC staging variables,
age, ethnicity, race, gender, and smoking-related
information.
Transcriptomic data are available for
1,284 patients.
Bulk RNA expression data were collected from
cBioPortal and initially contained raw expression counts
for more than 60,000 genes.
For the benchmark experiments, the transcriptomic data
were normalized using log counts per million and
dimensionality reduction based on highly variable genes
was performed, resulting in a subset of
4,096 genes.
Longitudinal follow-up information is available for
880 patients, comprising
1,431 follow-up records.
These records include tumor status and the number of days
to each follow-up visit, enabling longitudinal and
dynamic survival analyses.
Patients have a median of two follow-up visits.
The median time to the first follow-up is 298 days,
while the second occurs at a median of 647 days.
The dataset includes longitudinal treatment information
covering chemotherapy, radiation therapy, surgery, and
immunotherapy.
| Treatment |
Patients |
Treatment Records |
| Chemotherapy |
277 |
426 |
| Radiation Therapy |
923 |
1337 |
| Surgery |
291 |
525 |
| Immunotherapy |
49 |
58 |
Treatment records include temporal information such as
days to treatment start, allowing treatment events to be
incorporated into longitudinal patient trajectories.
MMIST Lung supports several survival prediction tasks
designed to evaluate multi-modal learning under realistic
missing-data conditions.
Disease-Specific Survival Prediction
Disease-specific survival is predicted using information
available at diagnosis, combining clinical,
transcriptomic, WSI, CT, and PET modalities.
12-Month Overall Survival Prediction
The dataset supports binary prediction of whether a
patient survives beyond 12 months after diagnosis.
Approximately 86% of the cohort survives beyond
12 months.
Dynamic Survival Prediction
Longitudinal clinical and treatment information is used
to predict disease-specific survival at the patient's
last known follow-up time.
Longitudinal Hazard Prediction
The dataset also supports discrete-time survival
analysis, where the probability of an event is modeled
across successive time intervals as new patient
information becomes available.
The following results summarize the main benchmark
performance reported for the MMIST Lung dataset.
| Task |
Best Reported Performance |
|
Disease-Specific Survival
|
70.36% Balanced Accuracy
|
|
12-Month Overall Survival
|
65.42% Balanced Accuracy
|
|
Dynamic Disease-Specific Survival
|
81.73 ± 0.37% Balanced Accuracy
|
|
Longitudinal Hazard Prediction
|
85.59 ± 4.03% C-index
|
The dataset was divided into
five patient-level cross-validation folds,
stratified according to 12-month survival status.
All scans, treatment records, follow-up information,
and other modalities belonging to the same patient
remain in the same fold, preventing patient-level
information leakage.
Balanced accuracy is used as the primary evaluation
metric for binary survival prediction due to the
substantial class imbalance in the cohort.
For further information about the dataset, curation
process, longitudinal information, and benchmark
experiments, please refer to our paper:
Rita Cordeiro Mendes,
Maria Rita Fonseca Verdelho,
Carlos Santiago,
and Catarina Barata.
Real-World Multi-Modal and Longitudinal Lung Cancer Dataset.
How to cite us?
Cite the paper where the dataset was first published
and acknowledge MMIST:
"The results shown in this work used datasets collected
from MMIST:
https://Multi-Modal-IST.github.io/"
Get in Touch at
ana.c.fidalgo.barata@tecnico.ulisboa.pt