Language Selection

Get healthy now with MedBeds!
Click here to book your session

Protect your whole family with Orgo-Life® Quantum MedBed Energy Technology® devices.

Advertising by Adpathway

         

 Advertising by Adpathway

Evidence Under Fire: New Debate Over How Medication Optimisation Tools Should Be Judged

4 hours ago 12

PROTECT YOUR DNA WITH QUANTUM TECHNOLOGY

Orgo-Life the new way to the future

  Advertising by Adpathway

When older patients take a dozen or more medications each day, clinicians need tools that tell them which drugs are doing more harm than good. The most widely used of these instruments, the STOPP/START criteria, have shaped deprescribing and prescribing-check practice across Europe for two decades. But a pointed exchange published in European Geriatric Medicine has reignited a question that sits at the heart of evidence-based geriatrics: what standard of proof should such medication optimisation tools be held to, and who gets to decide when that standard has been met? In a letter to the editor, researchers from KU Leuven and University Hospitals Leuven reflect on two recent publications, one by Boland and colleagues appraising the references behind STOPP/START version 3, and a response from O’Mahony and colleagues defending the criteria, and argue that the dispute exposes a deeper methodological fault line in how clinical decision-support tools are validated.

The controversy began when Boland and collaborators systematically examined the individual references cited in support of the STOPP/START version 3 criteria, a 2023 revision of the screening tool for older persons’ prescriptions. Their appraisal, published in the same journal, examined the quality of the evidence underpinning each recommendation, classifying studies according to established hierarchies of evidence such as the GRADE framework, which rates the certainty of evidence from high to very low and ties the strength of a recommendation to that rating. The findings were sobering. A substantial proportion of the criteria rested on studies with weak designs, and only a minority were anchored in randomised controlled trials or systematic reviews. For a tool that clinicians use to justify starting or stopping drugs in frail patients, the appraisal implied that many recommendations were grounded more in expert consensus than in high-certainty trials.

O’Mahony and colleagues, the architects of the STOPP/START criteria, responded forcefully. They argued that applying randomised-trial standards to prescribing criteria for older people misreads the nature of the evidence base in geriatric medicine. Very few trials enrol the oldest, most polymedicated, most comorbid patients in whom these tools are deployed; much of the relevant knowledge comes from observational studies, pharmacovigilance data and accumulated clinical experience. In their view, the absence of high-grade trial evidence does not invalidate a criterion that is biologically plausible, clinically coherent and consistent with decades of observed practice. Expecting each criterion to meet the evidentiary bar of a drug-approval dossier, they suggested, would paralyse the development of guidance in a field where such trials are frequently infeasible or unethical.

In their letter, Van der Linden, Shao and Tournoy take a middle path that is nonetheless critical of the status quo. They acknowledge the practical constraints O’Mahony’s team describes, but contend that the debate reveals an unresolved structural problem: medication optimisation tools are being updated and disseminated without an agreed, transparent framework for grading the evidence behind each criterion. The Leuven authors point to their own experience developing the RASP_CARDIO list, a consensus-validated screening tool for cardiovascular pharmacotherapy in geriatric patients derived from an adjusted STOPP list. That project, they note, required explicit consensus methods to validate each recommendation, illustrating an alternative model in which the provenance and strength of every criterion is documented from the outset rather than defended after criticism.

Technical questions about evidence appraisal are far from academic. A medication optimisation tool functions as a decision-support algorithm: it maps patient characteristics, drugs and clinical events to recommendations. Like any algorithm, its reliability depends on the quality of the data and rules encoded within it. When a criterion such as ‘avoid long-term benzodiazepines in older adults’ is supported by high-certainty evidence, clinicians can apply it with confidence. When a more nuanced criterion depends on a single retrospective cohort or a small case series, the tool effectively embeds an assumption rather than a demonstrated effect. Users of the tool, and the health systems that build electronic prescribing alerts around it, generally cannot see that distinction unless the underlying evidence grading is published alongside the criteria themselves.

The Leuven authors observe that this problem is not unique to prescribing tools. A comparative analysis of 50 European Society of Cardiology guidelines published between 2011 and 2022 found that only a minority of recommendations across cardiology practice were supported by the highest level of evidence, and that many rested on class IIb or expert-opinion grounds. Even in cardiology, arguably the most trial-rich specialty in medicine, the pyramid of evidence narrows sharply near the top. If high-certainty evidence is scarce even there, the argument goes, geriatric medicine, whose patient populations are chronically under-represented in trials, cannot reasonably be judged by the same yardstick, and blanket criticism of consensus-based criteria risks throwing away genuinely useful guidance.

Yet the letter also insists that scarcity of trials does not exempt a tool from scrutiny; it changes the kind of scrutiny required. Where randomised evidence is unavailable, the Leuven team argues, developers should make the evidential basis of each criterion explicit: whether it derives from meta-analysis, observational data, pharmacological reasoning or expert consensus, and how confident the developers are that applying the criterion improves patient outcomes. Tools should also be evaluated on the outcomes that matter. A screening list can flag potentially inappropriate prescriptions with perfect sensitivity, but if applying it does not reduce falls, hospitalisations or adverse drug events, and does not preserve quality of life, its clinical value remains unproven. Outcome-based validation studies, not reference appraisals alone, are ultimately what justify clinical use.

The exchange also touches on the economics and governance of guideline production. STOPP/START version 3 was produced by a large European expert panel, and its criteria are increasingly embedded in electronic health record systems, prescribing software and national deprescribing programmes across multiple countries. The letter’s authors note that once a tool reaches this level of adoption, its recommendations acquire de facto regulatory weight: a flagged prescription may be challenged, deprescribed or denied reimbursement on the basis of a criterion whose evidentiary pedigree few end users will ever examine. In their view, this asymmetry between the influence of a tool and the transparency of its foundations is precisely why systematic appraisal efforts such as Boland’s, though unwelcome to the tool’s developers, perform a necessary public service. The appropriate response to criticism, they suggest, is not to lower the bar for appraisal but to build evidence grading into the tool itself, version by version.

What emerges from the three-way exchange is less a verdict on STOPP/START than a blueprint for the next generation of medication optimisation instruments. The Leuven authors propose that developers adopt explicit, published evidence-grading schemes; that consensus methods be documented so that panels’ reasoning is reproducible; that criteria be linked wherever possible to outcome data; and that revision cycles incorporate external appraisal as a feature rather than a threat. They also stress that geriatricians should keep using validated tools in practice, since unsupported prescribing is far riskier than criteria supported by imperfect evidence, while recognising that imperfect evidence is an argument for transparency and continued validation, not complacency. As polypharmacy grows with ageing populations and electronic prescribing systems automate the application of criteria at scale, the question the letter poses, how evidence should be judged for medication optimisation tools, will shape whether these algorithms earn the trust that clinical practice, and the patients they serve, already place in them.

Subject of Research: Evidence standards for medication optimisation tools such as the STOPP/START criteria in geriatric prescribing

Article Title: How should evidence be judged for medication optimisation tools? Reflections on Boland et al. and O’Mahony et al.

Article References: How should evidence be judged for medication optimisation tools? Reflections on Boland et al. and O’Mahony et al.. (n.d.). https://doi.org/10.1007/s41999-026-01545-4

Image Credits: AI Generated

DOI: 10.1007/s41999-026-01545-4

Keywords: STOPP/START criteria, medication optimisation, deprescribing, geriatrics, polypharmacy, evidence-based medicine, GRADE framework, inappropriate prescribing, clinical decision support, consensus validation, pharmacotherapy, older adults

Cite Scienmag News
APA MLA Chicago

Copy citation Download RIS

Tags: clinical decision supportconsensus validationdeprescribingevidence-based medicinegeriatricsGRADE frameworkinappropriate prescribingmedication optimisationolder adultspharmacotherapypolypharmacySTOPP/START criteria

Read Entire Article

         

        

Start the new Vibrations with a Medbed Franchise today!  

Protect your whole family with Quantum Orgo-Life® devices

  Advertising by Adpathway