ESMO Artificial Intelligence and Digital Oncology Committee
How do I know whether this system is good enough to use?
EVALLM answers that question for clinician-facing AI systems in oncology. It sets out what should be measured, and how, before such a system is adopted into practice — across clinical accuracy, safety, equity, transparency, robustness, utility, governance and study design.
The framework at a glance
Every statement was rated on a nine-point scale. Consensus required at least 75% of ratings in the 7 to 9 range and no more than 15% in the 1 to 3 range. No statement in the final framework falls below that line.
- consensus statements
- 84consensus statementsacross eight evaluation domains
- at strong consensus
- 23at strong consensus96% agreement or above
- mean agreement
- 90.1%mean agreementhighest in safety and harm, lowest in bias and equity
- without a counterpart
- 41without a counterpartthe rest map to TRIPOD+AI, TRIPOD-LLM, DECIDE-AI, CONSORT-AI or SPIRIT-AI
Three ways to use it
Build an evaluation checklist
Describe the system and its clinical tasks, and get the statements that apply to it. Record what the evidence shows, then export the result as a PDF, a spreadsheet or JSON.
Start the checklistAssess a system for adoption
Twenty-eight questions in four sections: what the published evaluation shows, what your institution requires, what the vendor must supply, and what to watch once the system is in use.
Open the adoption checklistUse the framework in code
EVALLM is published in a machine-readable format with an instruction file, so that a software engineering agent can scope an assessment and report which statements a system satisfies.
Get the machine-readable filesEight evaluation domains
Open the framework explorer- 01
Clinical accuracy
Whether the output of the system is correct. In oncology this is difficult, because the correct answer is itself often uncertain and depends on factors including patient preference, and because the clinical reasoning path matters as much as the final output. Six of fifteen statements reached strong consensus, the highest proportion of any domain.
15 statements · mean 93.1%
- 02
Safety and harm
Safety is distinct from accuracy: a system can be accurate on average and still produce individual outputs that cause harm. The preparatory review found hallucination rates reported inconsistently, potentially harmful recommendations rarely documented, and no standardised harm classification scale in use.
11 statements · mean 93.4%
- 03
Bias and equity
This domain achieved the lowest agreement in the framework, showing where additional evidence is most needed. Only one of nine statements reached strong consensus and four sit at 79.3%. Agreement was high on measuring subgroup performance and lower on the obligations that arise beyond the evaluation dataset.
9 statements · mean 84.5%
- 04
Transparency, explainability and reproducibility
What must be disclosed for a published result to be interpretable at all: the exact system, version and date of access, the prompting strategy, the evidence behind each output, and whether repeated identical queries agree. The same model name can denote different systems at different dates.
8 statements · mean 91.9%
- 05
Robustness, reliability and post-deployment monitoring
Performance drifts as medical knowledge evolves while training data stays static, so validation is not a single event. All four strong-consensus statements in this domain concern what happens after a system enters clinical use rather than before it; agreement was lower on the evidence required before deployment.
11 statements · mean 90.6%
- 06
Usability, workflow integration and clinical utility
A system can be accurate, safe and stable and still provide no clinical utility. The single strong-consensus statement here is the one defining utility as the measurable effect of AI use on actual clinical decisions, rather than accuracy in isolation.
11 statements · mean 85.9%
- 07
Governance, legal and ethical considerations
Institutional rules on what patient data may leave the institution, on responsibility for the decision, on regulatory classification, and on disclosure to patients. The panel also addressed unofficial or shadow use directly, through guidance and education rather than prohibition alone.
8 statements · mean 89.1%
- 08
Validation study design
How the evaluation itself should be built: proportionate to clinical risk, pre-registered, powered a priori, screened for data contamination, and conducted in oncology rather than borrowed from other specialties. Statements are written at the level of clinical tasks so that they survive changes in the underlying technology.
11 statements · mean 90.3%