Skip to content
European Society for Medical OncologyEVALLM

Materials, methods and results

How the framework was agreed

A preparatory review established how clinician-facing AI systems are currently evaluated. A multidisciplinary panel then drafted candidate statements across eight domains and rated them in a structured consensus process, in two rounds in April and June 2026.

Preparatory literature review

What the evidence base looks like

PubMed and Web of Science were searched from 1 January 2010 to 20 January 2026, restricted to English-language articles. ClinicalTrials.gov was searched on the same date. This was not a systematic review: it was designed to inform statement development rather than to answer a predefined research question, and it was not registered.

  1. 792Records identifiedWeb of Science 375 · PubMed 321 · ClinicalTrials.gov 96
  2. 695Screened after duplicates removed97 duplicates removed using Covidence
  3. 172Taken to full extractionScreened by two reviewers
  4. 74Studies included98 excluded after full-text review
  5. 31Trials extracted in detailRegistered or published clinical trials

Why 98 studies were excluded after full-text review

  • Ineligible LLM type · 39
  • Patient-facing application · 25
  • Ineligible intervention · 18
  • Ineligible study design · 13
  • Ineligible outcomes · 2
  • Ineligible clinical setting · 1

Only 5 of the 30 extracted trials were conducted in an oncology population. Eighteen had a true control arm, and 22 specified validation requirements the reviewers judged strong. Structured safety assessment was the weakest area across all of them.

Consensus process

The rules, set before rating

Statements were rated on a nine-point scale, where 1 to 3 indicates disagreement, 4 to 6 uncertainty and 7 to 9 agreement. Consensus was defined a priori as at least 75% of ratings in the agreement band together with no more than 15% in the disagreement band. Statements below 50% would have been excluded; those between 50% and 74% were revised and carried into the next round. Between rounds, panel members received the anonymised distribution of scores.

Round 1 · April 2026

87

statements rated by 29 panel members

82 reached consensus, with no skipped items

Round 2 · June 2026

8

statements rated by 27 panel members

8 reached consensus, with no skipped items

The five statements that were revised

No statement fell below 50% agreement, so none were excluded. Each of these was rewritten on the basis of free-text feedback and rated again.

StatementRound 1Round 2
1.6Human reference standard62.1%85.2%
2.5Scope and institutional readiness69%92.6%
1.7Datasets and statistical reporting72.4%100%
3.1Subgroup performance reporting72.4%92.6%
6.2Efficiency72.4%92.6%

The three pairs that were merged

Free-text comments identified three pairs covering the same ground from different angles. These were merged for redundancy, not for low agreement: of the six components, one had reached 100% agreement in round 1 and another 96.6%.

  • 2.10 + 5.105.10Structured reporting of errors and safety concerns
  • 3.3 + 3.83.8Validation in the target language
  • 4.6 + 7.37.3Disclosure to patients

The disagreement criterion never determined an outcome. Across both rounds no statement attracted more than 7.4% of ratings in the 1 to 3 band, so every accept-or-revise decision was driven by the agreement criterion alone.

Results

Agreement by domain

The final framework contains 84 statements: 76 carry agreement figures from round 1 and 8 from round 2. 23 reached strong consensus at 96% or above and 61 reached consensus between 75% and 95%. Mean agreement was highest in safety and harm and in clinical accuracy, and lowest in bias and equity.

Mean agreement and number of strong-consensus statements in each of the eight domains.
DomainMean agreementStatementsStrong
1Clinical accuracy93.1%156
2Safety and harm93.4%113
3Bias and equity84.5%91
4Transparency91.9%83
5Robustness and monitoring90.6%114
6Usability and utility85.9%111
7Governance and ethics89.1%81
8Study design90.3%114

Bars run from 72% to 100%, the same scale used throughout the site. The lowest single statement in the framework is 6.5 at 75.9%.

Table 2

The 30 trials identified

Characteristics of the 30 clinical trials evaluating large language models in health care that were taken forward for detailed extraction.

Publication statusPublished 7/30 (23%), of which peer-reviewed 3 and pre-print 4; completed but not published 7/30 (23%); ongoing 16/30 (53%)
Study designEvaluation of clinical implementation 27/30 (90%); performance evaluation only 3/30 (10%)
Clinical conditionOncologic 5/30 (17%): colorectal cancer, gastrointestinal cancers, breast cancer (two trials), respiratory disease including lung cancer. Non-oncologic 25/30 (83%), most commonly general medicine and diagnostic reasoning (5), cardiology (4), documentation and workflow (3) and ophthalmology (3)
Control armTrue control arm 18/30 (60%); no control arm 12/30 (40%), comprising single-arm observational (5), comparison with historical or paired data (4) and within-patient before-after or case-crossover designs (3)
Performance criteria definedExplicitly defined 28/30 (93%). Most frequent measures: clinical accuracy 24/30 (80%), time or efficiency 18/30 (60%), quality scores 16/30 (53%), safety or harm 12/30 (40%), patient-reported outcomes 9/30 (30%), clinical outcomes such as survival 5/30 (17%)
Validation requirements definedStrong 22/30 (73%), comprising randomised design, blinded assessment, inter-rater reliability metrics or standardised rubrics; moderate 6/30 (20%); limited or unclear 2/30 (7%)
EvaluatorsPhysicians only 23/30 (77%); both physicians and patients 6/30 (20%); patients only 1/30 (3%)

Cross-cutting

Regulatory position

Clinician-facing AI systems occupy an uncertain regulatory position. In the European Union, software intended for diagnostic or treatment purposes falls under the Medical Devices Regulation or the In Vitro Diagnostic Regulation, and the EU AI Act adds a further layer whose interaction with those frameworks is still being clarified. In the United States, the Food and Drug Administration regulates software as a medical device, while clinical decision support software meeting certain criteria is excluded from device regulation. Many general-purpose systems currently used by clinicians are not regulated as medical devices in any jurisdiction, and some are used for purposes their developers explicitly disclaim.

EVALLM applies to systems that are not already regulated as approved medical devices, and it is not a substitute for regulatory approval. The panel recommended that regulatory classification be explicitly considered during evaluation and deployment.

Limitations

What this framework does not settle

  • 01

    The framework contains 84 statements, substantially more than previous ESMO guidance such as ELCAP or EBAI.

  • 02

    The preparatory literature review was not systematic. It was designed to inform statement development rather than to answer a predefined research question, and it was not registered.

  • 03

    The trial extraction is a snapshot of a fast-moving field. Trials registered after 20 January 2026 are not included, and the registry search covered ClinicalTrials.gov only, so trials registered elsewhere were missed.

  • 04

    EVALLM addresses ELCAP type 2 systems only. Patient-facing systems and background institutional systems raise related but distinct evaluation questions that this framework does not attempt to answer.

Scope

Where EVALLM sits

EVALLM covers clinician-facing AI systems used for clinical decision support in oncology, corresponding to ELCAP type 2. Patient-facing tools and background institutional systems are outside its current scope. Autonomous agents — systems that initiate clinical actions without a direct human prompt — are addressed by the separate ESMO EVAGENT workstream, which is also where legal accountability for such systems is directed.

ELCAP

Classifies LLM applications into three types and sets practical guidance for each

EVALLM

Specifies how clinician-facing systems should be evaluated

EBAI

Defines validation requirements where the output is a discrete biomarker

EVAGENT

Addresses systems that act without a direct human prompt