Every evaluation of a clinician-facing AI system in oncology should compare AI performance with an appropriate human reference standard (e.g. an oncologist or multidisciplinary tumor board) performing the same task under similar conditions.
Table 1
The EVALLM framework
Eighty-four statements, reproduced as they were rated by the panel. Agreement is the percentage of panel members rating the statement 7 to 9 on a nine-point scale; round 1 had 29 raters and round 2 had 27. The classifications along actor, stage, applicability and thematic group aid navigation — they do not rank statements and carried no weight in the consensus process. How the consensus was reached.
84 of 84 statements
Clinical accuracy
15 shown · mean 93.1%Whether the output of the system is correct. In oncology this is difficult, because the correct answer is itself often uncertain and depends on factors including patient preference, and because the clinical reasoning path matters as much as the final output. Six of fifteen statements reached strong consensus, the highest proportion of any domain.
Human reference standard
In clinical situations where no established guideline exists (e.g. complex molecular tumor board decisions, rare cancers, or emerging biomarker-defined subgroups), AI outputs should be evaluated by a panel of subspecialty experts against the best available evidence rather than against a fixed guideline reference.
For clinical decision support use cases, AI outputs should be evaluated by multiple independent clinical experts (typically a multidisciplinary tumor board or equivalent panel) using a predefined scoring scheme appropriate to the clinical task. The number of raters and required level of independence should be determined based on task complexity, and inter-rater reliability should be reported using appropriate metrics (intraclass correlation coefficient or Fleiss kappa).
For biomarker interpretation tasks (e.g. actionability of a genomic alteration, interpretation of a complex molecular or multi-omic profile), AI accuracy should be benchmarked against molecular tumor board decisions or equivalent expert review.
Guidelines as benchmark
Guideline concordance (the proportion of AI-generated recommendations that align with current evidence-based guidelines such as ESMO guidelines) should be reported as one measure of clinical accuracy but should not be the sole benchmark.
As clinical guidelines may be incomplete, outdated, or absent for certain scenarios (e.g. rare molecular profiles, novel combinations, rapidly evolving treatment landscapes), the evaluation framework should also allow for AI recommendations to differ from current guidelines if they are supported by credible published evidence or expert clinical reasoning.
When an AI recommendation differs from existing guidelines, the evaluation should distinguish between (a) a genuinely incorrect or unsupported recommendation and (b) a recommendation that is concordant with emerging evidence, recent trial results, or expert consensus not yet captured in current guidelines.
Error and completeness
When AI systems are used for treatment recommendations, the evaluation should assess not only the top recommendation but also the completeness and appropriate ranking of all reasonable alternatives, including their supporting evidence.
Accuracy evaluation should distinguish between factual errors (wrong information), omission errors (missing critical information), and reasoning errors (correct facts but flawed clinical logic).
Evaluation of clinical accuracy should include assessment of output comprehensiveness, including whether clinically relevant options and supporting information are adequately captured.
Task and modality specificity
Clinical accuracy should be reported separately for each intended task (e.g. diagnosis, treatment recommendation, toxicity management, biomarker interpretation, clinical trial matching, molecular tumor board support), as performance may vary substantially across tasks.
For AI systems that process multimodal inputs (e.g. combining pathology images, radiology, genomic data, and clinical text), accuracy should be evaluated both for the integrated output and, where feasible, for each input modality individually to identify modality-specific weaknesses.
Datasets and statistical reporting
Evaluation datasets should be representative of the AI system's intended target population. For systems with a broad clinical scope, this should include a mix of common and rare cancer types, stages, and treatment lines. Accuracy metrics should be reported both overall and stratified by clinically meaningful subgroups (e.g. disease stage, treatment line, previous treatments, rare vs common presentation).
AI accuracy metrics should be reported with confidence intervals and, where applicable, inter-rater agreement (e.g. Cohen's kappa) among the human evaluators who serve as the reference standard, the proportion of cases with expert disagreement and the adjudication method.
AI system performance should be evaluated on both standardized clinical vignettes and real-world, unstructured clinical data (e.g. actual electronic health records with missing or inconsistent information) to assess robustness under realistic conditions. The scope of validation (e.g. per institution, per EHR system, or per clinical setting) should be specified and justified.
Safety and harm
11 shown · mean 93.4%Safety is distinct from accuracy: a system can be accurate on average and still produce individual outputs that cause harm. The preparatory review found hallucination rates reported inconsistently, potentially harmful recommendations rarely documented, and no standardised harm classification scale in use.
Error and harm measurement
Evaluation studies should report the hallucination rate, defined as outputs that contain fabricated clinical information not supported by the input data or established medical knowledge (e.g. citing a non-existent trial, inventing a drug dose, fabricating a molecular finding, or misinterpreting an image).
A standardized harm classification scale for clinical AI outputs should be adopted (e.g. no harm / minor harm / moderate harm / severe or potentially lethal harm), with definitions specific to oncology decision-making. Where applicable, established clinical scales (e.g. CTCAE) should inform the grading.
Safety evaluations should assess whether the AI system generates contraindicated or unsafe recommendations, for example suggesting a drug to which the patient has a documented allergy, recommending a treatment incompatible with organ function, or overlooking a critical imaging finding.
Evaluation studies should include analysis of outputs that could have caused patient harm if acted on without clinician review (a “near-miss” analysis).
Built-in safety behaviour
AI systems used for clinical decision support should communicate uncertainty in a clear, interpretable, and clinically meaningful way.
AI systems should incorporate built-in safety checks, including flagging of potential contraindications, prompting the clinician when critical data is missing from the input, and linking recommendations to their supporting evidence sources.
When an AI system is used to screen or triage clinical information (e.g. laboratory results, imaging reports), the system should define a minimum confidence level below which the output is automatically flagged for mandatory human review.
Testing under adverse conditions
AI safety should be tested under conditions of incomplete or contradictory input data (e.g. missing lab values, conflicting radiology and pathology reports, poor-quality imaging), because these situations are common in real-world practice and may trigger unreliable outputs.
Safety evaluations should assess the robustness of AI systems to adversarial or manipulative inputs, including prompt injection where relevant (e.g. prompt injection attacks).
Scope and institutional readiness
AI systems evaluated under this framework must function as decision support tools with human clinicians retaining final responsibility for all treatment decisions. Autonomous therapeutic decision-making by AI systems is outside the scope of this framework.
Institutions deploying clinical AI systems should ensure that adequate clinical expertise and fallback procedures are available in case the AI system becomes unavailable or produces unreliable outputs due to technical failure, system errors, or missed updates.
Bias and equity
9 shown · mean 84.5%This domain achieved the lowest agreement in the framework, showing where additional evidence is most needed. Only one of nine statements reached strong consensus and four sit at 79.3%. Agreement was high on measuring subgroup performance and lower on the obligations that arise beyond the evaluation dataset.
Subgroup performance reporting
AI evaluation studies should report performance stratified by clinically and demographically relevant patient variables, including at minimum age and sex. Where data are available, additional stratification should include race or ethnicity, socioeconomic factors, and geographic region. Missing demographic data should be reported transparently.
AI performance should be evaluated separately for common cancers and for rare or underrepresented cancer types, to identify performance gaps that may disadvantage patients with less common diseases.
The evaluation framework should assess whether AI outputs reflect biases present in training data, for example underrepresentation of certain ethnic groups or geographic regions in clinical trial populations, or systematic patterns in imaging datasets from specific institutions.
Setting and health-system context
AI-generated treatment recommendations should be evaluated for local applicability: whether the recommended treatments are actually available, approved, and reimbursed in the healthcare system where the tool is being used.
Evaluation datasets should include clinical scenarios from diverse healthcare settings (academic centers, community hospitals, low-resource settings) to test whether the AI system performs equitably across different levels of clinical infrastructure and data availability.
Language
AI systems intended for use in, or producing outputs across, a language different from their primary training or development language should be validated in the target language(s), with particular attention to high-stakes content such as treatment selection and toxicity grading.
Task-specific equity
When AI systems are used for clinical trial matching, evaluations should assess whether eligible patients are identified equitably across demographic and disease groups.
AI outputs should be evaluated for the ability to integrate patient preferences, values, and goals of care (e.g. tolerance for toxicity, preference for quality of life over maximal efficacy) when these are provided as part of the clinical input.
Acquisition and technical bias
For multimodal AI systems processing imaging data (e.g. pathology slides, radiology), bias evaluation should include assessment across different imaging equipment, staining protocols, and scanning conditions.
Transparency, explainability and reproducibility
8 shown · mean 91.9%What must be disclosed for a published result to be interpretable at all: the exact system, version and date of access, the prompting strategy, the evidence behind each output, and whether repeated identical queries agree. The same model name can denote different systems at different dates.
Disclosure of what was evaluated
Every published AI evaluation study should disclose the specific system(s) evaluated (including model name, version, and date of access), the prompting or input strategy employed, and any additional methods applied (e.g. fine-tuning on medical data, connecting the system to external knowledge sources, or integration of multiple data modalities).
For multimodal AI systems, the transparency requirements should extend to each input modality: which data types the system processes, how they are combined, and whether the system can indicate which input modality most influenced a given recommendation.
Evidence traceability and quality
AI outputs used for clinical decision support should include traceable references to the evidence sources on which the recommendation is based (e.g. specific guidelines, clinical trials, or publications).
Evaluations should assess not only whether evidence is cited, but also whether the cited evidence is relevant, accurate, and appropriate in strength for the recommendation made. A recommendation supported by a phase III randomized trial should be weighted differently from one supported only by a case report, preclinical data, or a source that the system fabricated.
Reproducibility and replication
Evaluation studies should report the reproducibility of AI outputs: whether repeated queries with identical clinical inputs produce consistent answers, or whether outputs vary meaningfully across runs.
Evaluation studies should make key evaluation materials available to support independent replication, including datasets, vignettes, prompts, and scoring rubrics, where legally and ethically feasible.
Deployment model choice
When both open-source (locally installed) and proprietary (cloud-based) systems are being considered, the evaluation should explicitly report trade-offs in transparency, data privacy, ability to customize the system, and clinical performance.
The choice between deploying an open-source (locally hosted) system vs. a cloud-based commercial system should be documented as an institutional governance decision, with explicit consideration of data privacy, regulatory compliance, and the level of institutional control over the system.
Robustness, reliability and post-deployment monitoring
11 shown · mean 90.6%Performance drifts as medical knowledge evolves while training data stays static, so validation is not a single event. All four strong-consensus statements in this domain concern what happens after a system enters clinical use rather than before it; agreement was lower on the evidence required before deployment.
Validation pathway
AI system performance should be validated in at least one external institution (external validation) before the tool is recommended for routine clinical use. The scope of external validation should account for relevant differences in clinical infrastructure, including the electronic health record system in use.
AI evaluation should follow a three-stage validation pathway: (1) retrospective evaluation on historical clinical cases, (2) controlled prospective evaluation in a clinical setting, (3) monitored real-world deployment with ongoing outcome tracking.
Drift and re-evaluation
After deployment, systematic monitoring should be in place to detect performance degradation over time, which may occur as medical knowledge evolves, new drugs are approved, or clinical guidelines are updated while the system's training data remains static.
A defined schedule for re-evaluation of deployed AI systems should be established, for example triggered by major guideline updates, new drug approvals, changes in standard-of-care, or at regular intervals (e.g. annually).
Post-deployment monitoring should, where feasible, link AI-supported clinical decisions to actual patient outcomes (e.g. treatment response rates, progression-free survival, adverse event rates) over time, to assess the real-world impact of the tool.
Version control
Clear rules should be established for when and how a deployed AI system is updated to a new version, and what level of re-validation is required before the new version replaces the previous one in clinical use. Institutions should have dedicated expertise available to manage this process.
Validation of a specific AI model version should not automatically extend to successor versions or substantially updated releases. Each major version update should undergo independent evaluation proportionate to the scope of the changes.
AI system outputs should demonstrate acceptable stability across different technical configurations (e.g. hardware, software versions, API updates) without clinically meaningful changes in recommendations.
Shared responsibility and reporting
The responsibility for continuous quality monitoring and safety surveillance of clinical AI systems should be shared between the vendor or developer and the deploying institution, with clearly defined roles for each party.
Vendors of clinical AI products should be required to support structured feedback collection from clinical users and to participate in post-deployment safety monitoring, analogous to existing post-market surveillance requirements for medical devices.
A structured reporting channel should be established for clinicians to report errors, unexpected outputs, and safety concerns to both institutional quality assurance and the AI system developer, enabling continuous post-deployment monitoring and early detection of systematic safety signals.
Usability, workflow integration and clinical utility
11 shown · mean 85.9%A system can be accurate, safe and stable and still provide no clinical utility. The single strong-consensus statement here is the one defining utility as the measurable effect of AI use on actual clinical decisions, rather than accuracy in isolation.
Efficiency
AI evaluation should include time-efficiency metrics: the time required for the clinician to obtain, review, and (if necessary) correct the AI output should be measured and compared to the time required for the same task without AI assistance.
Clinical utility evaluation should assess whether AI use improves efficiency for the intended task, in comparison with the standard AI-unassisted workflow for the intended task, defined to include time required, cognitive workload, or resource use. Efficiency gains should not be treated as a substitute for accuracy or clinical impact.
Effect on clinical decisions
Clinical utility should be evaluated not only as accuracy in isolation but as the measurable impact of AI use on actual clinical decisions (e.g. the proportion of cases where the tumor board decision changed after reviewing AI input).
Where feasible, prospective studies should assess the incremental value of AI by comparing decisions made with and without AI support.
For tasks where AI systems are used to augment expert reasoning in complex cases (e.g. molecular tumor boards, rare cancers, multi-line treatment decisions), the evaluation should specifically measure whether the AI adds clinically relevant information or perspectives that would not have been considered by the clinical team alone.
Reported outcomes
Clinician-reported outcomes (e.g. perceived usability, trust in the system, cognitive load, perceived impact on decision quality) should be measured using validated instruments as a secondary outcome in all prospective AI evaluation studies.
Where AI-supported decisions directly affect patient-facing interactions (e.g. treatment discussions, informed consent conversations), patient-reported outcomes (e.g. satisfaction with the consultation, understanding of the treatment plan, perceived quality of communication) should be assessed as an additional outcome measure.
Patients should be involved in the evaluation of AI systems where the outputs directly affect treatment discussions or care plans.
Workflow integration
Evaluation studies should assess the effect of AI integration on clinical workflow, including interoperability with existing systems and any additional documentation, verification, or administrative burden.
Human-system risks
The evaluation should consider the risk of de-skilling: whether prolonged reliance on AI support leads to measurable decline in clinicians' ability to reason independently when the tool is unavailable.
Evaluation studies should assess the risk of over-reliance (automatic bias) on AI, including clinician acceptance of incorrect outputs without adequate independent review.
Governance, legal and ethical considerations
8 shown · mean 89.1%Institutional rules on what patient data may leave the institution, on responsibility for the decision, on regulatory classification, and on disclosure to patients. The panel also addressed unofficial or shadow use directly, through guidance and education rather than prohibition alone.
Institutional policy and shadow use
Institutions should establish formal policies that either integrate validated AI tools into clinical workflows under defined governance, or clearly define acceptable boundaries for individual, unofficial use of external AI systems by clinicians.
Institutional governance should explicitly address unofficial or “shadow” use of AI in clinical practice through guidance, education, and risk mitigation. This should be acknowledged and addressed through institutional guidance and education, rather than through prohibition alone.
Institutions should define which categories of patient information may be submitted to external commercial systems vs. only to locally deployed institutional systems.
Responsibility and regulation
Responsibility for clinical decisions should remain with the treating clinician, regardless of whether an AI system was used in the decision-making process.
The regulatory classification of clinical AI systems should be explicitly considered in the evaluation and deployment framework.
Patient transparency and ethics
When AI systems are formally integrated into institutional clinical workflows, institutions should disclose to patients that AI-assisted decision support is being used as part of their care.
AI evaluation studies should address ethical review requirements, including informed consent for patients whose clinical data is used in AI evaluation or real-world deployment studies.
Forward reference
The legal accountability for autonomous AI agents (systems that can initiate clinical actions without a direct human prompt) is a distinct and unresolved issue that goes beyond the scope of clinician-facing decision support. This topic should be addressed through dedicated legal and regulatory guidance in the planned ESMO EVAGENT workstream.
Validation study design
11 shown · mean 90.3%How the evaluation itself should be built: proportionate to clinical risk, pre-registered, powered a priori, screened for data contamination, and conducted in oncology rather than borrowed from other specialties. Statements are written at the level of clinical tasks so that they survive changes in the underlying technology.
Design proportionality and strength
The study design used to evaluate an oncology AI system should be proportionate to the clinical risk, intended use, and stage of deployment of the system.
Randomized or otherwise well-controlled prospective studies should be considered the strongest design for demonstrating the clinical utility of high-impact AI decision-support systems.
Evaluation studies targeting hard clinical outcomes (e.g. progression-free survival, overall survival, adverse event rates) should be encouraged as the highest level of evidence, while acknowledging that such studies require large sample sizes and long follow-up.
Statistical planning and integrity
The risk of data contamination (i.e. the possibility that the AI system was trained on the same clinical cases used for evaluation) should be assessed and reported in all evaluation studies, particularly those using published clinical vignettes or board-style examination questions.
Evaluation studies should pre-register their primary outcomes, scoring rubrics, and statistical analysis plans in a public registry to reduce the risk of selective outcome reporting.
AI evaluation studies should include a priori sample size justification, with an explicit statement of the primary endpoint, target precision or statistical power, and expected performance levels.
Scalability and comparison
Scalable evaluation approaches that combine automated screening (e.g. using a validated second AI system as a structured reviewer) with targeted expert review may be acceptable, provided the automated screening component has been validated against human expert ratings.
Head-to-head comparison studies evaluating multiple AI systems on the same clinical task using standardized protocols should be encouraged, as these provide the most informative data for clinical and institutional decision-making.
Transferability across specialties
Evidence from AI evaluation studies conducted outside oncology (e.g. in general medicine, cardiology, or other specialties) should not be considered sufficient for recommending deployment in oncology without dedicated oncology-specific validation, given the unique complexity and risk profile of cancer treatment decisions.
Evaluation frameworks and study designs should be formulated at the level of clinical tasks and evaluation principles rather than specific technologies, so that they remain applicable as the underlying AI systems evolve.
Task-specific design
For AI-based clinical trial matching tools, prospective evaluation should assess effects on trial identification, time to enrolment, enrolment rate, and representativeness of enrolled patients.