Summary
A retrospective study of 2,396 people with advanced NSCLC found clinical-data AI models outperformed single biomarkers and helped physicians predict treatment response and survival.
A large international study of people with advanced non-small-cell lung cancer (NSCLC) found that artificial-intelligence models using routine clinical and blood data predicted immunotherapy outcomes more consistently than individual biomarkers. In a separate usability exercise, explanations generated by the model improved physicians’ predictions of treatment response and survival.
The study analysed retrospective data from 2,396 adults with stage IIIC–IVB NSCLC treated with immunotherapy (IO), either alone or with chemotherapy, between September 2012 and October 2023. The patients came from six clinical centres in Italy, Germany, Greece, Israel, Spain and the United States.
The findings come from the retrospective phase of the I3LUNG project, whose prospective validation phase is ongoing.
Contents
- How the models were evaluated
- Routine clinical data performed consistently
- What the physician usability study found
- Why prospective testing matters
How the models were evaluated
The main analysis combined 2,075 patients who received immunotherapy as first-line treatment for metastatic disease or in later lines. The researchers divided the data into a training set of 1,550 patients, a test set of 274 patients and an external-validation cohort of 251 patients from the University of Chicago.
The models used clinical and blood features including performance status, smoking status, programmed death ligand 1 (PD-L1) expression, metastatic sites, the neutrophil-to-lymphocyte ratio and lactate dehydrogenase. Other analyses added computed tomography (CT) features, digital pathology and genomic information.
The researchers measured overall survival, survival at six and 24 months, and disease-control rate (DCR). DCR counted stable disease, partial response and complete response as disease control, while progressive disease was classified as no disease control. For classification outcomes, area under the curve (AUC) measures how well a model separates patients in different outcome groups; a value of 0.5 represents chance-level separation and higher values indicate better discrimination.
Routine clinical data performed consistently
The machine-learning clinical-and-blood-data models achieved AUC values of up to 0.77 in the test set. For example, the model’s AUC for predicting survival of at least 24 months was 0.74 in the test set and 0.60 in external validation. For six-month survival, the corresponding values were 0.69 and 0.64; for DCR, they were 0.71 and 0.55.
In the test set, the clinical-data models outperformed individual measures such as PD-L1, Eastern Cooperative Oncology Group performance status, lactate dehydrogenase and the neutrophil-to-lymphocyte ratio for 24-month survival prediction. The machine-learning model also produced a survival concordance index of 0.66 in the test set and 0.65 in external validation. The concordance index measures how well a model ranks patients by survival time.
Adding CT and pathology information produced higher results in some cross-validation analyses. For example, a multimodal model combining clinical data, digital pathology and radiomic CT features reached an AUC of 0.72 for 24-month survival, compared with 0.60 for the matched clinical-data model. However, these multimodal gains were not consistently reproduced in the independent test and external-validation cohorts. Deep-learning models also did not gain a clear benefit from adding other data types.
What the physician usability study found
The researchers then tested whether an explainable version of the clinical-data model could help doctors interpret individual cases. Twenty physicians—10 lung-cancer specialists and 10 non-specialists—assessed 100 real-world cases. They first predicted disease control and survival using clinical information, and then repeated the task with model outputs and SHAP-based explanations showing which features influenced each prediction.
For DCR, sensitivity increased from 0.72 without AI support to 0.87 with the model and explanations. Accuracy increased from 0.57 to 0.65, while specificity fell slightly. Both specialists and non-specialists showed higher sensitivity after using the explainable tool.
For survival estimation, the probability of a correct prediction increased by 36% across all physicians. The proportion of assessments matching the actual survival category rose from 16.5% to 22.5%. Agreement between specialists and non-specialists also improved when they used the model’s explanations.
These assessments were a structured usability experiment rather than decisions made during patient care. The study therefore measured whether the tool changed prediction performance, not whether using it improved treatment choices or patient outcomes.
Why prospective testing matters
The study’s retrospective design means the models were developed from existing records collected across centres with different patient populations, treatments and data completeness. Only 339 patients had all four data types available, limiting the strength of conclusions about multimodal AI. External-validation performance also varied, and the fairness analysis found differences in sensitivity between some clinical centres.
The authors are conducting prospective validation in more than 2,000 patients, including testing of both clinical-data and multimodal decision-support systems. They also describe a planned pragmatic randomised trial as part of a pathway toward regulatory approval and clinical implementation. Until those studies are completed, the strongest finding is that routine clinical data, combined with explainable AI, can support prediction of immunotherapy outcomes in this retrospective NSCLC dataset.