Thursday, July 29, 2021

New GRADE guidance on assessing imprecision in a network meta-analysis

Imprecision is one of the major domains of the GRADE framework and is used to assess whether to rate down the certainty of evidence related to an outcome of interest. In a traditional ("pairwise") meta-analysis which compares two intervention groups, exposures, or tests against one another, two considerations are made: the confidence interval around the absolute estimate of effect, and the optimal information size (OIS). If the bounds of the confidence interval cross a threshold for a meaningful effect, and/or if optimal information size given the sample size in the meta-analysis is not met, then one should consider rating down for imprecision.

In the context of small sample sizes, confidence intervals around an effect may be fragile - meaning they could be changed substantially with additional information. Therefore, the consideration of OIS along with the bounds of the confidence interval helps address this concern when rating the certainty of evidence to develop a clinical recommendation. This is typically done by assessing whether the sample size of the meta-analysis meets that determined by a traditional power analysis for a given effect size.

However, in a network meta-analysis, both direct and indirect comparisons are made across various interventions or tests. Thus, especially if the inclusion of indirect comparisons changes the overall estimate of effect, considering only the sample size involved in the direct comparisons would be misleading. 


A new GRADE guidance paper lays out how to assess imprecision in the context of a network meta-analysis:

  • If the 95% confidence interval crosses a decision-making threshold, rate down for imprecision. Thresholds should be ideally set a priori. It may be considered to rate down by two or even three levels depending on the degree of imprecision and the resulting communication of the certainty of evidence. For example, if imprecision is the only concern for an outcome, rating down by two instead of one level would be the difference between saying that a certain intervention or test "likely" or "probably" increases or decreases a given outcome, versus whether it simply "may" have this effect.
  • If the 95% confidence interval does not cross a decision-making threshold, consider whether the effect size may be inflated. If a point estimate is far away enough from a threshold, even a relatively wide CI may not cross it. Further, relatively large effect sizes from smaller pools of evidence can be reduced with future research. 
    • In the case of a large effect size, consider whether OIS is met. If the number of patients contributing to a NMA does not meet this number, consider rating down by one, two, or three levels depending on the severity of the width of the CI. 
    • If the upper-limit of a confidence interval using relative risk is 3 or more times higher than the lower-limit, OIS has likely not been met. Similarly, upper-to-lower-limit comparisons of odds ratios exceeding 2.5 have likely not met OIS.
  • Alternatively, when the effect size is both modest, plausible, and does not cross a threshold, one likely does not need to rate down for imprecision. 
  • Avoid "double dinging" for imprecision if this limitation has already been addressed by rating down elsewhere.

Brignardello-Peterson R, Guyatt GH, Mustafa RA, et al. (2021). GRADE guidelines 33. Addressing imprecision in a network meta-analysis. J Clin Epidemiol (in-press). 

Manuscript available at the publisher's website here.





Friday, July 16, 2021

New GRADE concept paper identifies challenges and solutions to use of GRADE in public health contexts

The GRADE framework can be applied across a variety of different fields, not the least of which is public health. Public health, as the authors of a new GRADE concept paper define it, is concerned with "preventing disease, prolonging life, and promoting health through the organized efforts of society" and comprises three key domains: health protection, health services, and health improvement. However, the field of public health also has unique challenges in the application of GRADE that require addressing. 

To dig deeper into these challenges and design a plan of action for solutions and guidance, the GRADE Public Health group conducted a scoping review to better understand published accounts of the barriers, challenges, and facilitators to the adoption and application of GRADE in public health contexts, presenting the results of nine identified articles. Of these, five major challenges were identified:

  • Incorporating diverse perspectives 
  • Selecting and prioritizing outcomes
  • Interpreting outcomes and identifying a threshold for decision-making
  • Assessing certainty of evidence from diverse sources (e.g., nonrandomized studies)
  • Addressing implications for decision-makers, including concerns about conditional recommendations
The article then discusses proposed solutions and a work plan to address these key challenges.


Forthcoming GRADE public health guidance articles, collaborations with the GRADE Evidence-to-Decision working group, and the adaptation of GRADE training materials to nonhealth and policy audiences will help guide those in public health contexts in meeting the unique needs presented for rigorous guideline development. Additional promotion of existing GRADE guidance, such as the consideration of equity in the evidence-to-decision process, may help guideline developers within specific challenges related to selecting and prioritizing outcomes or identifying thresholds for decision-making. Ongoing guidance from the GRADE group for Non-Randomizes Studies and the use of ROBINS-I may further improve the application of GRADE in settings where observational evidence is dominant. 

Hilton Boon, M., Thomson, H., Shaw, B., et al. (2021). Challenges in applying the GRADE approach in public health guidelines and systematic reviews: A concept article from the GRADE public health group. J Clin Epidemol 135:42-53.

Article available at the publisher's website here










 

Friday, June 25, 2021

Scholars at 14th GRADE Workshop Discuss the Unique Challenges of Sparse Evidence, Guideline Collaborations, and Financial Incentives in Healthcare

During the 14th GRADE Guideline Development Workshop held virtually last month, the Evidence Foundation had the pleasure of welcoming three new scholars with the opportunity to attend the workshop free of charge. As part of the scholarship, each recipient presented to the workshop attendees about their current or proposed project related to evidence-based medicine and reducing bias in healthcare.

This spring's lot of three scholars was nothing short of incredibly impressive. Ifeoluwa Babatunde, a PhD student in clinical research at Case Western Reserve University, discussed the unique challenges of developing a guideline on the management of patients undergoing patent foramen ovale (PFO) closure for the Society for Cardiovascular Angiography and Interventions (SCAI). The synthesis of evidence for this question is hampered by controversies and limited evidence as well as complications due to comorbidities and age differences in the populations of interest. Babatunde discussed her interest in attending the workshop to learn more about the appropriate use of observational and indirect evidence to better answer questions related to PFO closure.

"The GRADE workshop helped me to see systematic review methodology from a deeper and more critical perspective," said Babatunde. "GRADE offers a very comprehensive yet succinct and transparent framework for developing and ascertaining the certainty of evidence in guidelines. Hence I feel better equipped to tackle challenges that arise from creating reviews and guidelines regarding conditions and populations with sparse RCTs."


Next, Dr. Pichamol Jirapinyo, the Director of Bariatric Endoscopy Fellowship at Brigham and Women's Hospital and instructor at Harvard Medical School, discussed her work on an international joint guideline development effort between the American Society for Gastrointestinal Endoscopy (ASGE) and the European Society of Gastrointestinal Endoscopy (ESGE) to produce recommendations for endoscopic and bariatric metabolic therapy (EBMT) in patients with obesity. EBMT is one of several possible management routes for obesity, alongside pharmacological and surgical options. The project will aim to answer several questions, including how patients should be managed before and after EBMT, and regarding the safety and efficacy of both gastric and small bowel EBMT.

“The GRADE workshop provided me a great framework on how to apply GRADE methodology to systematic review and meta-analysis to rigorously develop a guideline," said Dr. Jirapinyo. 'In addition to learning about the GRADE methodology itself, I found the workshop to be tremendously helpful with providing practical tips on how to run a guideline task force successfully and efficiently.”  

Finally, Dr. Lillian Lai, a research fellow in the Department of Urology at the University of Michigan, presented an intriguing discussion of financial incentives in clinical decision-making in urology. The surveillance and management of localized prostate cancer, for instance, has several different options ranging from active surveillance (which is less costly) to prostatectomy (which is more costly). Regardless of the reported health outcomes of these approaches, there is little financial incentive to conduct surveillance as opposed to surgery. The project's goal is to use health services research methods to understand how urologists response to large financial incentives, and then create financial incentives and remove financial disincentives for the promotion of guideline-concordant practices. 

"I gained invaluable knowledge on how to use the GRADE approach to rate the certainty of evidence and strength of recommendations," said Dr. Lai. "Going through the guideline development tool with experts in small groups was particularly useful for me to understand what a guideline recommendation means and entails. This workshop came at a critical time in the backdrop of COVID, and the ever-changing landscape of medicine where patients and providers need to make timely and informed decisions together."

If you are interested in learning more about GRADE and attending the workshop as a scholarship recipient, applications for our upcoming virtual workshop in October are now open. The deadline to apply is July 31, 2021. Details can be found here. 

Wednesday, June 9, 2021

Evidence Foundation scholar spotlight: Georgios Schoretsanitis

Last fall, Dr. Georgios Schoretsanitis attended the 13th (and first-ever virtual) GRADE guideline development workshop as a scholar of the Evidence Foundation. As such, he presented to the rest of the workshop attendees on his work developing guidelines for therapeutic drug monitoring to optimize and tailor treatment for psychotherapeutic medications. Beginning in 2017, a series of recommendations for reference ranges for two commonly prescribed antipsychotic medications was developed, followed this year by an international joint consensus statement on blood levels to optimize antipsychotic treatment in clinical practice.

Dr. Schoretsanitis now has an exciting update on his project.

"My main research interest is therapeutic drug monitoring, also known as TDM, which refers to the quantification and interpretation of medication levels in the blood (plasma or serum) of the patient treated with psychotropic agents," says Dr. Schoretsanitis. "The aim of TDM in clinical practice is to improve treatment response and safety outcomes. Apart from analyzing TDM clinical routine data, I have also been working as a member of the TDM taskforce of the German Association of Neuropsychopharmacology and Phaarmacopsychiatry (Arbeitsgemeinschaft für Neuropsychopharmakologie und Pharmacopsychiatrie; AGNP) involved in systematic reviews of TDM literature, which provide so-called therapeutic reference ranges for medication levels. These ranges may orient clinicians during dose selection. Attending the virtual GRADE workshop in October 2020 provided me much of inspiration, but also knowledge of well-established methodological tools for the assessment of quality of evidence.

Hereafter, in the TDM task force of AGNP, we adopted a GRADE-oriented approach in assessing TDM literature as we are reviewing new TDM evidence on commonly prescribed antipsychotics under the supervision of Prof. Gerhard Gründer, Department of Molecular Neuroimaging, Central Institute of Mental Health, Medical Faculty Mannheim, University of Heidelberg, Mannheim, Germany. This type of approach is more standardized and follows GRADE guidelines. Ultimately, this work will enhance methodological rigidity for the next Consensus guidelines for therapeutic drug monitoring in neuropsychopharmacology [last update 2018; Hiemke et al, Pharmacopsychiatry]. I strongly encourage researchers involved in systematic reviews or assessment of evidence quality to attend the GRADE workshop which enables a major upgrade of related skills and knowledge."



Stay tuned for future updates from other past Evidence Foundation scholars like Dr. Schoretsanitis and the exciting work they are doing to improve the application of GRADE methodology and evidence-based medicine.

If you are interested in learning more about GRADE and attending the workshop as a scholarship recipient, applications for our upcoming virtual workshop in October are now open. The deadline to apply is July 31, 2021. Details can be found here. 


Thursday, June 3, 2021

The Systematic Survey Behind a Collection of Minimal Important Differences (MIDs) Across the Patient-Reported Outcome Literature

Patient-reported outcome measures, or PROMs, allow clinicians and researchers to directly elicit information about treatments that are important to patients, such as side effects or improvements in pain, function, or quality of life. In order to interpret changes in PROMs relative to a clinical recommendation, however, a minimal important difference (MID) - or the smallest possible change in the outcome that would mandate a change in the patient's management - must be determined.

Luckily, a wealth of published studies exist to provide a library of MIDs for a wide range of outcomes, and recently, a review published by Carrasco-Labra and colleagues synthesized these works together. The study included any empirical reports of MIDs in adolescents or adults that used an anchor-based approach, in which MIDs are based on an observed change related to an external criterion rather than the distribution of a particular patient sample. As such, anchor-based MIDs tend to be more directly applicable across patient populations. Ultimately, a collection of 585 studies reporting on 5,324 MID estimates across 526 distinct PROMs was presented.

About two-thirds (66%) collected MIDs related to patients' improvement, whereas about one-third (31%) addressed MIDs related to worsening or assumed the MIDs for improvement or worsening would be the same. Most (88%) were based on a longitudinal design in which patients' reported outcomes and satisfaction were measured at multiple timepoints. The most common types of anchors used were global ratings of change (59%), change in disease-related outcome (23%), and comparison with another group (11%), whereas the most common sources of anchor information were self-report (83%) proxy-reported (9%) and laboratory data (3%). 


MIDs are essential in interpreting the magnitude of an effect from a study or systematic review of evidence, especially when assessing imprecision as part of GRADE. They can also allow researchers to conduct "responder analyses" based on subsets of patients who experience a change in an outcome beyond a given MID. Finally, reporting mean differences in units of MIDs as part of a systematic review can standardize interpretation of an effect size in a way that may be less problematic to interpret than a traditional standardized mean difference (SMD).

The work corresponds to PROMID, a project to develop an inventory of MIDs across the literature, which can be accessed at https://promid.mcmaster.ca/. 



Carrasco-Labra A, Devji T, Qasim A, et al. (2021). Minimal important difference estimates for patient-reported outcomes: A systematic survey. J Clin Epidemiol 133:61-71.

Manuscript available at the publisher's website here. 










Friday, May 14, 2021

Reliability of Risk of Bias Assessments of Non-randomized Studies Improves After Customized Training

We previously reported on a paper published in 2020 assessing the inter-rater reliability (IRR) and inter-consensus reliability (ICR) of the Risk of Bias in Non-Randomized Studies of Interventions (ROBINS-I) tool, developed in 2016, and the Risk of Bias instrument for NRS of Exposures (ROB-NRSE) tool, developed in 2018. This paper found that reliability generally tended to be poor for these tools, while risk of bias assessments took evaluators, on average, 48 minutes for the ROBINS-I tool and almost 37 minutes for the ROB-NRSE.

Now, a new publication from the same group has examined the effect of training on the reliability of these tools. An international team of reviewers with a median of 5 years of experience with risk of bias assessment first applied the ROBINS-I and ROB-NRSE tools to a list of 44 non-randomized studies of interventions and exposures, respectively, using only the 53 pages of publicly available guidance. Then, the reviewers received an abridged and customized training document which was tailored specifically to the topic area of the reviews, included simplified guidance for assessing risk of bias, and also provided additional guidance related to more advanced concepts. The reviewers then re-assessed the studies' risk of bias after a several-weeks-long wash-out period.



Changes in the inter-rater reliability (IRR) for the ROBINS-I (top) and ROB-NRSE tools (bottom) from before and after a customized training intervention.


The training intervention improved the IRR of the ROBINS-I tool, generally improving the range of within-domain reliability while the reliability of the overall bias rating improved from "poor" to "fair." Meanwhile, the ICR improved substantially, with the overall rating's reliability improving from "poor" to "near perfect." Improvements were also observed after training in the application of the ROB-NRSE tool, with IRR of the overall bias improving significantly from "slight" to "near perfect" while its ICR improved from "poor" to "near perfect." For both tools, the pre-to-post-intervention correlations between reviewers' scores were poor, suggesting that the training did have an impact on these measures independent of a simple learning effect. While customized training was associated with a decrease in evaluator burden for the ROBINS-I tool, this did not hold true for the ROB-NRSE.

The findings of this analysis suggest that the use of a customized, shortened guidance tool specifically tailored to the topical content of a review, including simplified guidance for decision-making within each domain, can improve the reliability of resulting risk of bias assessments. The authors suggest that future reviewers create such guidance based on the specific needs and considerations of their topic area, and publish these tools along with the review.

Jeyaraman MM, Robson RC, Copstein L et al. (2021). Customized guidance/training improved the psychometric properties of methodologically rigorous risk of bias instruments for non-randomized studies. J Clin Epidemiol, in-press.

Manuscript available here. 































Tuesday, May 4, 2021

Restricting Systematic Search to English-only is a Viable Shortcut in Most, but Perhaps Not All Topics in Medicine

In the limitations sections of systematic reviews on any topic, it is not uncommon for the authors to discuss how language limitations within their search may have restricted the breadth of evidence presented. For instance, if the reviewers speak only English, the review is likely limited to publications and journals in that language. But how much of a difference does such a limitation make in terms of the overall conclusions of a systematic review? According to a new paper in the Journal of Clinical Epidemiology, probably not much - but it may depend on the specific topic of medicine under investigation.

While other methods reviews have previously examined this question, Dobrescu and colleagues extended the range of topics to methods reviews that included systematic reviews within the realm of complementary and alternative medicine, yielding four reviews previously unexamined by prior studies. Specifically, the authors looked for methods reviews comparing the restriction of literature searches to English-only versus unrestricted searches and whose primary outcomes compared differences in treatment effect estimates, certainty of evidence ratings, or conclusions based on the language restrictions enforced. 

The search yielded eight studies investigating the impact of language restrictions in anywhere from 9 to 147 systematic reviews in medicine. Overall, the exclusion of non-English articles had a greater impact on estimates of treatment effects and the statistical significance of findings in reviews of complementary and alternative medicine versus conventional medicine topics. Most commonly, the exclusion of non-English studies led to a loss of statistical significance in these topic areas.

Overall, the methods studies examined found that the exclusion of non-English studies of conventional medicine topics led to small to moderate changes in the estimate of effect; however, exclusion of non-English studies shrank the observed effect size in complementary and alternative medicine topics by 63 percent. Two studies examined whether language restricted influenced authors' overall conclusions, generally finding no effect.

The figure above shows the frequency of languages of the excluded reviews examined.

The authors conclude that when it comes to systematic reviews of conventional medicine topics, their findings are in line with those of previous methods studies which demonstrate little to no effect of language restrictions and suggest that restricting a search to English-only should not greatly impact the findings or conclusions of a review. However, the effect appears greater in the realm of complementary and alternative medicine, perhaps due to the greater proportion of non-English studies published in this field. Thus, systematic reviewers attempting to synthesize the evidence on an alternative medicine topic should be cognizant of their choices regarding language restriction and the potential implications they may have on their ultimate findings.

Dobrescu A, Nussbaumer SB, Klerings I et al. (2021). Restricting evidence syntheses of interventions to English-language publications is a viable methodological shortcut for most medical topics: A systematic review: Excluding English-language publications a valid shortcut. J Clin Epidemiol, epub ahead of print.

Manuscript available from publisher's website here.