Skip to content
European Society for Medical OncologyEVALLM

ESMO Artificial Intelligence and Digital Oncology Committee

How do I know whether this system is good enough to use?

EVALLM answers that question for clinician-facing AI systems in oncology. It sets out what should be measured, and how, before such a system is adopted into practice — across clinical accuracy, safety, equity, transparency, robustness, utility, governance and study design.

The framework at a glance

Every statement was rated on a nine-point scale. Consensus required at least 75% of ratings in the 7 to 9 range and no more than 15% in the 1 to 3 range. No statement in the final framework falls below that line.

Agreement with each of the 84 EVALLM statements, grouped by domain. Values run from 75.9% to 100%; 23 statements sit at or above the 96% strong-consensus line.
Agreement per statement, domains 1 to 8Read the same 84 statements as a list
consensus statements
84consensus statementsacross eight evaluation domains
at strong consensus
23at strong consensus96% agreement or above
mean agreement
90.1%mean agreementhighest in safety and harm, lowest in bias and equity
without a counterpart
41without a counterpartthe rest map to TRIPOD+AI, TRIPOD-LLM, DECIDE-AI, CONSORT-AI or SPIRIT-AI

Three ways to use it

Build an evaluation checklist

Describe the system and its clinical tasks, and get the statements that apply to it. Record what the evidence shows, then export the result as a PDF, a spreadsheet or JSON.

Start the checklist

Assess a system for adoption

Twenty-eight questions in four sections: what the published evaluation shows, what your institution requires, what the vendor must supply, and what to watch once the system is in use.

Open the adoption checklist

Use the framework in code

EVALLM is published in a machine-readable format with an instruction file, so that a software engineering agent can scope an assessment and report which statements a system satisfies.

Get the machine-readable files

Eight evaluation domains

Open the framework explorer
  1. 01

    Clinical accuracy

    Whether the output of the system is correct. In oncology this is difficult, because the correct answer is itself often uncertain and depends on factors including patient preference, and because the clinical reasoning path matters as much as the final output. Six of fifteen statements reached strong consensus, the highest proportion of any domain.

    15 statements · mean 93.1%

  2. 02

    Safety and harm

    Safety is distinct from accuracy: a system can be accurate on average and still produce individual outputs that cause harm. The preparatory review found hallucination rates reported inconsistently, potentially harmful recommendations rarely documented, and no standardised harm classification scale in use.

    11 statements · mean 93.4%

  3. 03

    Bias and equity

    This domain achieved the lowest agreement in the framework, showing where additional evidence is most needed. Only one of nine statements reached strong consensus and four sit at 79.3%. Agreement was high on measuring subgroup performance and lower on the obligations that arise beyond the evaluation dataset.

    9 statements · mean 84.5%

  4. 04

    Transparency, explainability and reproducibility

    What must be disclosed for a published result to be interpretable at all: the exact system, version and date of access, the prompting strategy, the evidence behind each output, and whether repeated identical queries agree. The same model name can denote different systems at different dates.

    8 statements · mean 91.9%

  5. 05

    Robustness, reliability and post-deployment monitoring

    Performance drifts as medical knowledge evolves while training data stays static, so validation is not a single event. All four strong-consensus statements in this domain concern what happens after a system enters clinical use rather than before it; agreement was lower on the evidence required before deployment.

    11 statements · mean 90.6%

  6. 06

    Usability, workflow integration and clinical utility

    A system can be accurate, safe and stable and still provide no clinical utility. The single strong-consensus statement here is the one defining utility as the measurable effect of AI use on actual clinical decisions, rather than accuracy in isolation.

    11 statements · mean 85.9%

  7. 07

    Governance, legal and ethical considerations

    Institutional rules on what patient data may leave the institution, on responsibility for the decision, on regulatory classification, and on disclosure to patients. The panel also addressed unofficial or shadow use directly, through guidance and education rather than prohibition alone.

    8 statements · mean 89.1%

  8. 08

    Validation study design

    How the evaluation itself should be built: proportionate to clinical risk, pre-registered, powered a priori, screened for data contamination, and conducted in oncology rather than borrowed from other specialties. Statements are written at the level of clinical tasks so that they survive changes in the underlying technology.

    11 statements · mean 90.3%