Skip to content
European Society for Medical OncologyEVALLM

Table 1

The EVALLM framework

Eighty-four statements, reproduced as they were rated by the panel. Agreement is the percentage of panel members rating the statement 7 to 9 on a nine-point scale; round 1 had 29 raters and round 2 had 27. The classifications along actor, stage, applicability and thematic group aid navigation — they do not rank statements and carried no weight in the consensus process. How the consensus was reached.

84 of 84 statements

Domain 1

Clinical accuracy

15 shown · mean 93.1%

Whether the output of the system is correct. In oncology this is difficult, because the correct answer is itself often uncertain and depends on factors including patient preference, and because the clinical reasoning path matters as much as the final output. Six of fifteen statements reached strong consensus, the highest proportion of any domain.

Human reference standard

1.1

Every evaluation of a clinician-facing AI system in oncology should compare AI performance with an appropriate human reference standard (e.g. an oncologist or multidisciplinary tumor board) performing the same task under similar conditions.

93.1% rd 1Consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AIDECIDE-AI
1.4

In clinical situations where no established guideline exists (e.g. complex molecular tumor board decisions, rare cancers, or emerging biomarker-defined subgroups), AI outputs should be evaluated by a panel of subspecialty experts against the best available evidence rather than against a fixed guideline reference.

89.7% rd 1Consensus
EvaluatorPre-deploymentTreatment recommendationEVALLM-specific
1.6

For clinical decision support use cases, AI outputs should be evaluated by multiple independent clinical experts (typically a multidisciplinary tumor board or equivalent panel) using a predefined scoring scheme appropriate to the clinical task. The number of raters and required level of independence should be determined based on task complexity, and inter-rater reliability should be reported using appropriate metrics (intraclass correlation coefficient or Fleiss kappa).

85.2% rd 2Consensus
EvaluatorPre-deploymentAll systemsTRIPOD-LLM
1.11

For biomarker interpretation tasks (e.g. actionability of a genomic alteration, interpretation of a complex molecular or multi-omic profile), AI accuracy should be benchmarked against molecular tumor board decisions or equivalent expert review.

86.2% rd 1Consensus
EvaluatorPre-deploymentBiomarker interpretationEBAI

Guidelines as benchmark

1.2

Guideline concordance (the proportion of AI-generated recommendations that align with current evidence-based guidelines such as ESMO guidelines) should be reported as one measure of clinical accuracy but should not be the sole benchmark.

93.1% rd 1Consensus
EvaluatorPre-deploymentTreatment recommendationEVALLM-specific
1.3

As clinical guidelines may be incomplete, outdated, or absent for certain scenarios (e.g. rare molecular profiles, novel combinations, rapidly evolving treatment landscapes), the evaluation framework should also allow for AI recommendations to differ from current guidelines if they are supported by credible published evidence or expert clinical reasoning.

86.2% rd 1Consensus
EvaluatorPre-deploymentTreatment recommendationEVALLM-specific
1.5

When an AI recommendation differs from existing guidelines, the evaluation should distinguish between (a) a genuinely incorrect or unsupported recommendation and (b) a recommendation that is concordant with emerging evidence, recent trial results, or expert consensus not yet captured in current guidelines.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentTreatment recommendationEVALLM-specific

Error and completeness

1.9

When AI systems are used for treatment recommendations, the evaluation should assess not only the top recommendation but also the completeness and appropriate ranking of all reasonable alternatives, including their supporting evidence.

89.7% rd 1Consensus
EvaluatorPre-deploymentTreatment recommendationEVALLM-specific
1.10

Accuracy evaluation should distinguish between factual errors (wrong information), omission errors (missing critical information), and reasoning errors (correct facts but flawed clinical logic).

100% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific
1.15

Evaluation of clinical accuracy should include assessment of output comprehensiveness, including whether clinically relevant options and supporting information are adequately captured.

93.1% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific

Task and modality specificity

1.8

Clinical accuracy should be reported separately for each intended task (e.g. diagnosis, treatment recommendation, toxicity management, biomarker interpretation, clinical trial matching, molecular tumor board support), as performance may vary substantially across tasks.

100% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsTRIPOD-LLM
1.12

For AI systems that process multimodal inputs (e.g. combining pathology images, radiology, genomic data, and clinical text), accuracy should be evaluated both for the integrated output and, where feasible, for each input modality individually to identify modality-specific weaknesses.

89.7% rd 1Consensus
EvaluatorPre-deploymentMultimodal systemsEVALLM-specific

Datasets and statistical reporting

1.7

Evaluation datasets should be representative of the AI system's intended target population. For systems with a broad clinical scope, this should include a mix of common and rare cancer types, stages, and treatment lines. Accuracy metrics should be reported both overall and stratified by clinically meaningful subgroups (e.g. disease stage, treatment line, previous treatments, rare vs common presentation).

100% rd 2Strong consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AI
1.13

AI accuracy metrics should be reported with confidence intervals and, where applicable, inter-rater agreement (e.g. Cohen's kappa) among the human evaluators who serve as the reference standard, the proportion of cases with expert disagreement and the adjudication method.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AI
1.14

AI system performance should be evaluated on both standardized clinical vignettes and real-world, unstructured clinical data (e.g. actual electronic health records with missing or inconsistent information) to assess robustness under realistic conditions. The scope of validation (e.g. per institution, per EHR system, or per clinical setting) should be specified and justified.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsDECIDE-AI
Domain 2

Safety and harm

11 shown · mean 93.4%

Safety is distinct from accuracy: a system can be accurate on average and still produce individual outputs that cause harm. The preparatory review found hallucination rates reported inconsistently, potentially harmful recommendations rarely documented, and no standardised harm classification scale in use.

Error and harm measurement

2.1

Evaluation studies should report the hallucination rate, defined as outputs that contain fabricated clinical information not supported by the input data or established medical knowledge (e.g. citing a non-existent trial, inventing a drug dose, fabricating a molecular finding, or misinterpreting an image).

100% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsTRIPOD-LLM
2.2

A standardized harm classification scale for clinical AI outputs should be adopted (e.g. no harm / minor harm / moderate harm / severe or potentially lethal harm), with definitions specific to oncology decision-making. Where applicable, established clinical scales (e.g. CTCAE) should inform the grading.

86.2% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific
2.3

Safety evaluations should assess whether the AI system generates contraindicated or unsafe recommendations, for example suggesting a drug to which the patient has a documented allergy, recommending a treatment incompatible with organ function, or overlooking a critical imaging finding.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific
2.6

Evaluation studies should include analysis of outputs that could have caused patient harm if acted on without clinician review (a “near-miss” analysis).

93.1% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific

Built-in safety behaviour

2.4

AI systems used for clinical decision support should communicate uncertainty in a clear, interpretable, and clinically meaningful way.

93.1% rd 1Consensus
DeveloperPre-deploymentAll systemsDECIDE-AI
2.7

AI systems should incorporate built-in safety checks, including flagging of potential contraindications, prompting the clinician when critical data is missing from the input, and linking recommendations to their supporting evidence sources.

100% rd 1Strong consensus
DeveloperPre-deploymentAll systemsEVALLM-specific
2.12

When an AI system is used to screen or triage clinical information (e.g. laboratory results, imaging reports), the system should define a minimum confidence level below which the output is automatically flagged for mandatory human review.

89.7% rd 1Consensus
DeveloperDeployment decisionScreening or triage useEVALLM-specific

Testing under adverse conditions

2.8

AI safety should be tested under conditions of incomplete or contradictory input data (e.g. missing lab values, conflicting radiology and pathology reports, poor-quality imaging), because these situations are common in real-world practice and may trigger unreliable outputs.

93.1% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific
2.9

Safety evaluations should assess the robustness of AI systems to adversarial or manipulative inputs, including prompt injection where relevant (e.g. prompt injection attacks).

93.1% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific

Scope and institutional readiness

2.5

AI systems evaluated under this framework must function as decision support tools with human clinicians retaining final responsibility for all treatment decisions. Autonomous therapeutic decision-making by AI systems is outside the scope of this framework.

92.6% rd 2Consensus
EvaluatorPre-deploymentAll systemsELCAPEVAGENT
2.11

Institutions deploying clinical AI systems should ensure that adequate clinical expertise and fallback procedures are available in case the AI system becomes unavailable or produces unreliable outputs due to technical failure, system errors, or missed updates.

89.7% rd 1Consensus
InstitutionDeployment decisionAll systemsDECIDE-AI
Domain 3

Bias and equity

9 shown · mean 84.5%

This domain achieved the lowest agreement in the framework, showing where additional evidence is most needed. Only one of nine statements reached strong consensus and four sit at 79.3%. Agreement was high on measuring subgroup performance and lower on the obligations that arise beyond the evaluation dataset.

Subgroup performance reporting

3.1

AI evaluation studies should report performance stratified by clinically and demographically relevant patient variables, including at minimum age and sex. Where data are available, additional stratification should include race or ethnicity, socioeconomic factors, and geographic region. Missing demographic data should be reported transparently.

92.6% rd 2Consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AI
3.2

AI performance should be evaluated separately for common cancers and for rare or underrepresented cancer types, to identify performance gaps that may disadvantage patients with less common diseases.

79.3% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific
3.9

The evaluation framework should assess whether AI outputs reflect biases present in training data, for example underrepresentation of certain ethnic groups or geographic regions in clinical trial populations, or systematic patterns in imaging datasets from specific institutions.

89.7% rd 1Consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AI

Setting and health-system context

3.4

AI-generated treatment recommendations should be evaluated for local applicability: whether the recommended treatments are actually available, approved, and reimbursed in the healthcare system where the tool is being used.

79.3% rd 1Consensus
EvaluatorDeployment decisionTreatment recommendationEVALLM-specific
3.5

Evaluation datasets should include clinical scenarios from diverse healthcare settings (academic centers, community hospitals, low-resource settings) to test whether the AI system performs equitably across different levels of clinical infrastructure and data availability.

79.3% rd 1Consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AI

Language

3.8

AI systems intended for use in, or producing outputs across, a language different from their primary training or development language should be validated in the target language(s), with particular attention to high-stakes content such as treatment selection and toxicity grading.

81.5% rd 2Consensus
EvaluatorPre-deploymentCross-language useEVALLM-specific

Task-specific equity

3.6

When AI systems are used for clinical trial matching, evaluations should assess whether eligible patients are identified equitably across demographic and disease groups.

82.8% rd 1Consensus
EvaluatorPre-deploymentTrial matchingEVALLM-specific
3.7

AI outputs should be evaluated for the ability to integrate patient preferences, values, and goals of care (e.g. tolerance for toxicity, preference for quality of life over maximal efficacy) when these are provided as part of the clinical input.

79.3% rd 1Consensus
EvaluatorPre-deploymentTreatment recommendationEVALLM-specific

Acquisition and technical bias

3.10

For multimodal AI systems processing imaging data (e.g. pathology slides, radiology), bias evaluation should include assessment across different imaging equipment, staining protocols, and scanning conditions.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentImaging inputsEBAI
Domain 4

Transparency, explainability and reproducibility

8 shown · mean 91.9%

What must be disclosed for a published result to be interpretable at all: the exact system, version and date of access, the prompting strategy, the evidence behind each output, and whether repeated identical queries agree. The same model name can denote different systems at different dates.

Disclosure of what was evaluated

4.1

Every published AI evaluation study should disclose the specific system(s) evaluated (including model name, version, and date of access), the prompting or input strategy employed, and any additional methods applied (e.g. fine-tuning on medical data, connecting the system to external knowledge sources, or integration of multiple data modalities).

100% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsTRIPOD-LLM
4.9

For multimodal AI systems, the transparency requirements should extend to each input modality: which data types the system processes, how they are combined, and whether the system can indicate which input modality most influenced a given recommendation.

89.7% rd 1Consensus
DeveloperPre-deploymentMultimodal systemsEVALLM-specific

Evidence traceability and quality

4.2

AI outputs used for clinical decision support should include traceable references to the evidence sources on which the recommendation is based (e.g. specific guidelines, clinical trials, or publications).

89.7% rd 1Consensus
DeveloperPre-deploymentAll systemsEVALLM-specific
4.3

Evaluations should assess not only whether evidence is cited, but also whether the cited evidence is relevant, accurate, and appropriate in strength for the recommendation made. A recommendation supported by a phase III randomized trial should be weighted differently from one supported only by a case report, preclinical data, or a source that the system fabricated.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific

Reproducibility and replication

4.4

Evaluation studies should report the reproducibility of AI outputs: whether repeated queries with identical clinical inputs produce consistent answers, or whether outputs vary meaningfully across runs.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsTRIPOD-LLM
4.7

Evaluation studies should make key evaluation materials available to support independent replication, including datasets, vignettes, prompts, and scoring rubrics, where legally and ethically feasible.

89.7% rd 1Consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AI

Deployment model choice

4.5

When both open-source (locally installed) and proprietary (cloud-based) systems are being considered, the evaluation should explicitly report trade-offs in transparency, data privacy, ability to customize the system, and clinical performance.

89.7% rd 1Consensus
EvaluatorDeployment decisionOpen-source vs cloud choiceEVALLM-specific
4.8

The choice between deploying an open-source (locally hosted) system vs. a cloud-based commercial system should be documented as an institutional governance decision, with explicit consideration of data privacy, regulatory compliance, and the level of institutional control over the system.

82.8% rd 1Consensus
InstitutionDeployment decisionOpen-source vs cloud choiceEVALLM-specific
Domain 5

Robustness, reliability and post-deployment monitoring

11 shown · mean 90.6%

Performance drifts as medical knowledge evolves while training data stays static, so validation is not a single event. All four strong-consensus statements in this domain concern what happens after a system enters clinical use rather than before it; agreement was lower on the evidence required before deployment.

Validation pathway

5.1

AI system performance should be validated in at least one external institution (external validation) before the tool is recommended for routine clinical use. The scope of external validation should account for relevant differences in clinical infrastructure, including the electronic health record system in use.

79.3% rd 1Consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AIEBAI
5.2

AI evaluation should follow a three-stage validation pathway: (1) retrospective evaluation on historical clinical cases, (2) controlled prospective evaluation in a clinical setting, (3) monitored real-world deployment with ongoing outcome tracking.

82.8% rd 1Consensus
EvaluatorPre-deploymentAll systemsDECIDE-AI

Drift and re-evaluation

5.3

After deployment, systematic monitoring should be in place to detect performance degradation over time, which may occur as medical knowledge evolves, new drugs are approved, or clinical guidelines are updated while the system's training data remains static.

100% rd 1Strong consensus
InstitutionPost-deploymentAll systemsEBAI
5.4

A defined schedule for re-evaluation of deployed AI systems should be established, for example triggered by major guideline updates, new drug approvals, changes in standard-of-care, or at regular intervals (e.g. annually).

96.6% rd 1Strong consensus
InstitutionPost-deploymentAll systemsEVALLM-specific
5.7

Post-deployment monitoring should, where feasible, link AI-supported clinical decisions to actual patient outcomes (e.g. treatment response rates, progression-free survival, adverse event rates) over time, to assess the real-world impact of the tool.

89.7% rd 1Consensus
InstitutionPost-deploymentAll systemsDECIDE-AI

Version control

5.5

Clear rules should be established for when and how a deployed AI system is updated to a new version, and what level of re-validation is required before the new version replaces the previous one in clinical use. Institutions should have dedicated expertise available to manage this process.

96.6% rd 1Strong consensus
InstitutionPost-deploymentAll systemsEVALLM-specific
5.6

Validation of a specific AI model version should not automatically extend to successor versions or substantially updated releases. Each major version update should undergo independent evaluation proportionate to the scope of the changes.

93.1% rd 1Consensus
EvaluatorPost-deploymentAll systemsEVALLM-specific
5.11

AI system outputs should demonstrate acceptable stability across different technical configurations (e.g. hardware, software versions, API updates) without clinically meaningful changes in recommendations.

86.2% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific

Shared responsibility and reporting

5.8

The responsibility for continuous quality monitoring and safety surveillance of clinical AI systems should be shared between the vendor or developer and the deploying institution, with clearly defined roles for each party.

82.8% rd 1Consensus
Vendor · InstitutionPost-deploymentAll systemsEVALLM-specific
5.9

Vendors of clinical AI products should be required to support structured feedback collection from clinical users and to participate in post-deployment safety monitoring, analogous to existing post-market surveillance requirements for medical devices.

96.6% rd 1Strong consensus
VendorPost-deploymentAll systemsEBAI
5.10

A structured reporting channel should be established for clinicians to report errors, unexpected outputs, and safety concerns to both institutional quality assurance and the AI system developer, enabling continuous post-deployment monitoring and early detection of systematic safety signals.

92.6% rd 2Consensus
Institution · VendorPost-deploymentAll systemsEVALLM-specific
Domain 6

Usability, workflow integration and clinical utility

11 shown · mean 85.9%

A system can be accurate, safe and stable and still provide no clinical utility. The single strong-consensus statement here is the one defining utility as the measurable effect of AI use on actual clinical decisions, rather than accuracy in isolation.

Efficiency

6.1

AI evaluation should include time-efficiency metrics: the time required for the clinician to obtain, review, and (if necessary) correct the AI output should be measured and compared to the time required for the same task without AI assistance.

86.2% rd 1Consensus
EvaluatorPre-deploymentAll systemsDECIDE-AI
6.2

Clinical utility evaluation should assess whether AI use improves efficiency for the intended task, in comparison with the standard AI-unassisted workflow for the intended task, defined to include time required, cognitive workload, or resource use. Efficiency gains should not be treated as a substitute for accuracy or clinical impact.

92.6% rd 2Consensus
EvaluatorPre-deploymentAll systemsDECIDE-AI

Effect on clinical decisions

6.6

Clinical utility should be evaluated not only as accuracy in isolation but as the measurable impact of AI use on actual clinical decisions (e.g. the proportion of cases where the tumor board decision changed after reviewing AI input).

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific
6.7

Where feasible, prospective studies should assess the incremental value of AI by comparing decisions made with and without AI support.

89.7% rd 1Consensus
EvaluatorPre-deploymentAll systemsCONSORT-AI
6.10

For tasks where AI systems are used to augment expert reasoning in complex cases (e.g. molecular tumor boards, rare cancers, multi-line treatment decisions), the evaluation should specifically measure whether the AI adds clinically relevant information or perspectives that would not have been considered by the clinical team alone.

82.8% rd 1Consensus
EvaluatorPre-deploymentComplex or MTB casesEVALLM-specific

Reported outcomes

6.3

Clinician-reported outcomes (e.g. perceived usability, trust in the system, cognitive load, perceived impact on decision quality) should be measured using validated instruments as a secondary outcome in all prospective AI evaluation studies.

79.3% rd 1Consensus
EvaluatorPre-deploymentAll systemsDECIDE-AI
6.4

Where AI-supported decisions directly affect patient-facing interactions (e.g. treatment discussions, informed consent conversations), patient-reported outcomes (e.g. satisfaction with the consultation, understanding of the treatment plan, perceived quality of communication) should be assessed as an additional outcome measure.

89.7% rd 1Consensus
EvaluatorPre-deploymentPatient-facing outputsDECIDE-AI
6.11

Patients should be involved in the evaluation of AI systems where the outputs directly affect treatment discussions or care plans.

79.3% rd 1Consensus
EvaluatorPre-deploymentPatient-facing outputsEVALLM-specific

Workflow integration

6.5

Evaluation studies should assess the effect of AI integration on clinical workflow, including interoperability with existing systems and any additional documentation, verification, or administrative burden.

75.9% rd 1Consensus
EvaluatorDeployment decisionAll systemsDECIDE-AI

Human-system risks

6.8

The evaluation should consider the risk of de-skilling: whether prolonged reliance on AI support leads to measurable decline in clinicians' ability to reason independently when the tool is unavailable.

82.8% rd 1Consensus
EvaluatorPost-deploymentAll systemsEVALLM-specific
6.9

Evaluation studies should assess the risk of over-reliance (automatic bias) on AI, including clinician acceptance of incorrect outputs without adequate independent review.

89.7% rd 1Consensus
EvaluatorPre-deploymentAll systemsDECIDE-AI
Domain 7

Governance, legal and ethical considerations

8 shown · mean 89.1%

Institutional rules on what patient data may leave the institution, on responsibility for the decision, on regulatory classification, and on disclosure to patients. The panel also addressed unofficial or shadow use directly, through guidance and education rather than prohibition alone.

Institutional policy and shadow use

7.1

Institutions should establish formal policies that either integrate validated AI tools into clinical workflows under defined governance, or clearly define acceptable boundaries for individual, unofficial use of external AI systems by clinicians.

89.7% rd 1Consensus
InstitutionDeployment decisionAll systemsELCAP
7.2

Institutional governance should explicitly address unofficial or “shadow” use of AI in clinical practice through guidance, education, and risk mitigation. This should be acknowledged and addressed through institutional guidance and education, rather than through prohibition alone.

93.1% rd 1Consensus
InstitutionDeployment decisionAll systemsEVALLM-specific
7.6

Institutions should define which categories of patient information may be submitted to external commercial systems vs. only to locally deployed institutional systems.

100% rd 1Strong consensus
InstitutionDeployment decisionAll systemsELCAP

Responsibility and regulation

7.4

Responsibility for clinical decisions should remain with the treating clinician, regardless of whether an AI system was used in the decision-making process.

89.7% rd 1Consensus
ClinicianDeployment decisionAll systemsELCAP
7.5

The regulatory classification of clinical AI systems should be explicitly considered in the evaluation and deployment framework.

86.2% rd 1Consensus
InstitutionDeployment decisionAll systemsEBAI

Patient transparency and ethics

7.3

When AI systems are formally integrated into institutional clinical workflows, institutions should disclose to patients that AI-assisted decision support is being used as part of their care.

81.5% rd 2Consensus
InstitutionDeployment decisionAll systemsELCAP
7.7

AI evaluation studies should address ethical review requirements, including informed consent for patients whose clinical data is used in AI evaluation or real-world deployment studies.

82.8% rd 1Consensus
EvaluatorPre-deploymentAll systemsSPIRIT-AI

Forward reference

7.8

The legal accountability for autonomous AI agents (systems that can initiate clinical actions without a direct human prompt) is a distinct and unresolved issue that goes beyond the scope of clinician-facing decision support. This topic should be addressed through dedicated legal and regulatory guidance in the planned ESMO EVAGENT workstream.

89.7% rd 1Consensus
InstitutionDeployment decisionAll systemsEVAGENT
Domain 8

Validation study design

11 shown · mean 90.3%

How the evaluation itself should be built: proportionate to clinical risk, pre-registered, powered a priori, screened for data contamination, and conducted in oncology rather than borrowed from other specialties. Statements are written at the level of clinical tasks so that they survive changes in the underlying technology.

Design proportionality and strength

8.1

The study design used to evaluate an oncology AI system should be proportionate to the clinical risk, intended use, and stage of deployment of the system.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsDECIDE-AI
8.2

Randomized or otherwise well-controlled prospective studies should be considered the strongest design for demonstrating the clinical utility of high-impact AI decision-support systems.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsCONSORT-AI
8.3

Evaluation studies targeting hard clinical outcomes (e.g. progression-free survival, overall survival, adverse event rates) should be encouraged as the highest level of evidence, while acknowledging that such studies require large sample sizes and long follow-up.

86.2% rd 1Consensus
EvaluatorPre-deploymentAll systemsCONSORT-AI

Statistical planning and integrity

8.6

The risk of data contamination (i.e. the possibility that the AI system was trained on the same clinical cases used for evaluation) should be assessed and reported in all evaluation studies, particularly those using published clinical vignettes or board-style examination questions.

93.1% rd 1Consensus
EvaluatorPre-deploymentAll systemsTRIPOD-LLM
8.7

Evaluation studies should pre-register their primary outcomes, scoring rubrics, and statistical analysis plans in a public registry to reduce the risk of selective outcome reporting.

79.3% rd 1Consensus
EvaluatorPre-deploymentAll systemsSPIRIT-AI
8.8

AI evaluation studies should include a priori sample size justification, with an explicit statement of the primary endpoint, target precision or statistical power, and expected performance levels.

93.1% rd 1Consensus
EvaluatorPre-deploymentAll systemsTRIPOD+AI

Scalability and comparison

8.5

Scalable evaluation approaches that combine automated screening (e.g. using a validated second AI system as a structured reviewer) with targeted expert review may be acceptable, provided the automated screening component has been validated against human expert ratings.

89.7% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific
8.9

Head-to-head comparison studies evaluating multiple AI systems on the same clinical task using standardized protocols should be encouraged, as these provide the most informative data for clinical and institutional decision-making.

82.8% rd 1Consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific

Transferability across specialties

8.10

Evidence from AI evaluation studies conducted outside oncology (e.g. in general medicine, cardiology, or other specialties) should not be considered sufficient for recommending deployment in oncology without dedicated oncology-specific validation, given the unique complexity and risk profile of cancer treatment decisions.

82.8% rd 1Consensus
EvaluatorDeployment decisionAll systemsEVALLM-specific
8.11

Evaluation frameworks and study designs should be formulated at the level of clinical tasks and evaluation principles rather than specific technologies, so that they remain applicable as the underlying AI systems evolve.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentAll systemsEVALLM-specific

Task-specific design

8.4

For AI-based clinical trial matching tools, prospective evaluation should assess effects on trial identification, time to enrolment, enrolment rate, and representativeness of enrolled patients.

96.6% rd 1Strong consensus
EvaluatorPre-deploymentTrial matchingEVALLM-specific