Wednesday, January 22, 2020

Research Shorts: Rating the certainty in evidence in the absence of a single estimate of effect

Contributed by Madelin Siedler, 2019/2020 U.S. GRADE Network Research Fellow

When a pooled estimate from a meta-analysis of several studies is not present to guide the rating of evidence in these domains, how should one make a final determination of the certainty of evidence using GRADE? 


Evidence from a 30,000-foot view

In their 2017 paper published in Evidence-Based Medicine, Murad and colleagues describe methods for applying GRADE when bodies of evidence are either sparse or too disparate to pool. A systematic review, for instance, may only provide a narrative synthesis of the current evidence given these limitations. When a neat estimate of effect presented as part of a tidy forest plot is not available, it is necessary to use one’s best judgment to rate the domains by taking a broader view. In these cases, Murad et al. recommend the following approach:
  • Risk of Bias: Judge the risk of bias across all studies that include the outcome of interest.
  • Inconsistency: Consider the direction and size of the estimates of effect from each study. Generally, do they all tell the same story, or do they vary considerably?
  • Indirectness: Make an overall judgment about the amount of directness or indirectness of the body of evidence, given your specific question (always consider your population, intervention, outcome, and comparator[s] of interest). Generally, are the studies synthesized answering questions similar to yours? Or might the dissimilarities be enough to lower your trust in the estimate of effect as it pertains to your question?
  • Imprecision: Examine the total information size of all studies (number of events for binary outcomes, or number of participants for continuous outcomes) as well as each study’s reported confidence interval for this outcome. If there are fewer than 400 total events or participants, or if the confidence intervals from most studies - or the largest - include no effect, imprecision is likely present.
  • Publication bias: Suspect publication bias if there is a small number of only positive studies, or if data were reported in trial registries but never published.
As always, one may consider rating up the quality of evidence from an observational study if a large magnitude of effect, a dose-response gradient, or plausible residual confounding that would increase the certainty of effect are present in the majority of studies examined.


Murad MH, Mustafa RA, Schünemann HJ, Sultan S, Santesso N. Rating the certainty in evidence in the absence of a single estimate of effect. BMJ Evidence-Based Medicine. 2017 Jun 1;22(3):85-7.

Manuscript available here on publisher's site.

Monday, January 20, 2020

Research Shorts: Assessing the certainty of evidence in the importance of outcomes or values and preferences

Contributed by Madelin Siedler, 2019/2020 U.S. GRADE Network Research Fellow

The rating of outcomes in terms of their importance is a key aspect of GRADE guideline development. So is, of course, the rating of the certainty of evidence that will inform clinical decision-making. However, it is often difficult to rate the certainty of evidence of the importance of outcomes – assuming there is any evidence to draw from at all. In their July 2019 article published in the Journal of Clinical Epidemiology, Zhang and colleagues describe the ways to assess the certainty of a body of evidence used to determine the relative importance of outcomes.



The GRADE domains that present the most challenges when rating the certainty of evidence are inconsistency and imprecision. Assuming there is more than one study, assessment of inconsistency should include judging the amount of variance across studies’ reported importance of outcomes, exploring potential sources for this inconsistency (such as differences in populations or instruments used) and rating down when inconsistency is not explained by these. Imprecision should take into consideration the sample size first. In fact, in cases where there is no available quantitative synthesis, sample size may be the only consideration. In other cases, assuming information size meets a pre-defined threshold, the evidence may still be rated down if the confidence intervals of relative importance outcomes cross a pre-defined decision-making threshold.


Y. Zhang et al. (2019)/Journal of Clinical Epidemiology

The authors warn against attempts to rate the certainty of evidence in the variability of outcome importance – in other words, how much the perceived importance of any outcome varies from one individual to the next. If both inconsistency and imprecision are ruled out as potential sources of observed variance, then true variability may exist. In these cases, guideline panels should consider the formation of a conditional recommendation based on differences in values and preferences.

The article also provides guidance for assessing publication bias and rating up.


Zhang Y, Coello PA, Guyatt GH, Yepes-Nuñez JJ, Akl EA, Hazlewood G, Pardo-Hernandez H, Etxeandia-Ikobaltzeta I, Qaseem A, Williams Jr JW, Tugwell P. GRADE guidelines: 20. Assessing the certainty of evidence in the importance of outcomes or values and preferences—inconsistency, imprecision, and other domains. Journal of clinical epidemiology. 2019 Jul 1;111:83-93.

Manuscript available here on publisher's site.

Thursday, December 19, 2019

Research Shorts: A Theoretical Framework and Competency-Based Approach to Training in Guideline Development


Contributed by Madelin Siedler, 2019/2020 U.S. GRADE Network Research Fellow

As an increasing number of organizations are developing clinical guidelines, expectations for the quality and trustworthiness of these guidelines are on the rise as well. Thus, there is an increased need for guideline-producing organizations to identify, train, and hire or contract with individuals who are adequately skilled in guideline development methods, and for a clear delineation of the knowledge and skill sets required of these individuals.


Recently, seven members of the U.S. GRADE Network and GRADE Working Group co-published an article establishing a framework of core competencies for individuals serving on guideline development panels in a variety of roles, from panel content experts to methodologists. The paper outlines the minimal knowledge, skills, and expertise that would allow an individual to perform tasks adequately. The framework describes three major domains of competency for guideline development: facilitating the development of guideline structure and setup, making judgments about the quality or certainty of evidence, and transforming evidence into recommendations.


Within each core competency, there are multiple sub-competencies and various educational “milestones” which further clarify the required skill. These milestones track to the five stages of educational development established by the Dreyfus model: novice, advanced beginner, competent, proficient, and expert. The authors note that the required level of expertise related to these milestones will vary depending on the role of the individual in developing the guideline. While some sub-competencies related to rating of the certainty of evidence follow the GRADE approach specifically, much of the competency-based framework can be applied universally to other guideline development methodologies. The authors encourage future research efforts to validate, assess, and refine the proposed milestones for their widespread use in efforts to train the next generation of guideline developers.


Sultan S, Morgan RL, Murad MH, Falck-Ytter Y, Dahm P, Schünemann HJ, Mustafa RA. A Theoretical Framework and Competency-Based Approach to Training in Guideline Development. Journal of general internal medicine. 2019 Nov 14:1-7.

Manuscript available here on publisher's site.

Sunday, November 3, 2019

Fall 2019 - Scholarship recipients

Contributed by Madelin Siedler, 2019/2020 U.S. GRADE Network Research Fellow

The Eleventh GRADE Guideline Development Workshop was held in Orlando, Florida, this past September. This workshop welcomed 52 participants, including several international attendees from Canada and Brazil. Among these participants were three recipients of a scholarship provided by the Evidence Foundation, which covers the cost of registration. Coming from diverse backgrounds ranging from organizational work to evidence synthesis to policymaking, scholars Faduma Gure, Eric Linskens, and Christian Kershaw presented their proposals for new innovations or opportunities for improving the application and implementation of evidence-based medicine.


Faduma Gure, MSc, a knowledge translation and research specialist for the Association of Ontario Midwives, discussed the challenges of incorporating client perspectives to midwifery guidelines for the organization, which utilizes the GRADE approach. Pregnancy through the post-partum period is often a tumultuous time filled with decision-making, Gure explained, and more can be done to better understand the values and preferences of midwifery clients and to employ an equity lens when formulating recommendations. Gure proposed a solution that includes the development of an equity advisory group consisting of key stakeholders representative of Ontario’s population of midwifery clients. The organization could then involve these stakeholders through the entire guideline development process - from the initial setting of research priorities to the ultimate formulation of recommendations - and elicit important feedback about patient values and preferences as well as the potential impacts of a guideline on various communities.

Eric Linskens, BSc, serves the Minneapolis Veterans Affairs Evidence Synthesis Program and the Minnesota Agency for Healthcare Research and Quality (AHRQ) Evidence-based Practice Center. Linskens presented his current work applying the GRADE approach to existing systematic reviews which did not originally use GRADE. Linskens discussed some of the challenges that his team has faced as part of the initiative, as well as innovative solutions to these issues. For instance, a systematic review conducted by AHRQ may automatically rate down for inconsistency due simply to the inclusion of a sole study, whereas in the GRADE approach, this would not be the case. Additionally, an existing review may break down one clinical question into multiple smaller analyses of comparators or sub-populations, whereas it would be more clinically relevant to use these data to create one larger recommendation. To best solve these issues, Linskens noted, it is important to consider the end-user of any given review or guideline so that their needs can be best met. Additionally, transparently reporting all judgments around the analyses is key. Regarding his time at the workshop, Linskens said, “[i]t was very helpful to work through examples with the GRADE workshop facilitators in small group sessions. They answered our questions as they came up.”

Christian Kershaw, PhD, is a molecular neuroscientist who now works as a health policy analyst for CGS, a Medicare fee-for-service contractor. Dr. Kershaw used her personal experience transitioning from bench science to policymaking to inspire her presentation on the utility of cross-functional teams in medicine and healthcare policy. To develop a cross-functional team, Dr. Kershaw explained, it is best to identify a problem that would best be solved by a group of individuals with heterogeneous skills and backgrounds that would each uniquely serve a common goal or purpose. As an example, Kershaw discussed the development of a team to standardize the way information is used to form coverage decisions as part of the 21st Century Cures Act. The team is comprised of a medical doctor to understand the need for and content of the policies; an outreach and education specialist to understand their legal implications; and a basic research scientist to compile and assess the information. Leveraging individual team members’ strengths and encouraging innovation are keys to success when working in a cross-functional team. “I was impressed with the versatility of the GRADE framework,” Kershaw noted. “It was very informative to learn all of the different ways that the conference attendees were using GRADE to suit their projects.”

If interested in applying for a scholarship to future GRADE workshops, more details can be found here: https://evidencefoundation.org/scholarships.html. Please note the deadline for applications to our next workshop in Phoenix, AZ will be December 4, 2019.

Tuesday, May 14, 2019

Spring 2019 - Scholarship Recipients

Contributed by Madelin Siedler, 2018/2019 U.S. GRADE Network Research Fellow

Recently, we held the Tenth GRADE Guideline Development Workshop in Denver, Colorado. This workshop was one of the largest groups to date, with 51 participants traveling to the Mile-High City from as far away as Poland and Korea. During the workshop, participants focused on learning and applying the GRADE approach for diagnostic test accuracy. 

Two participants attended as recipients of the scholarship program funded by the U.S. GRADE Network and Evidence Foundation. This scholarship covers the cost of registration for workshop attendees who are newer to GRADE and have never attended a formal GRADE workshop. Scholars Janice Tufte and Dr. Irbaz bin Riaz presented on their innovative ideas for improving the development, implementation, or dissemination of guidelines with the aim of reducing bias in healthcare recommendations.

Scholarship recipients: Dr. Irbaz bin Riaz (L) and Ms. Janice Tufte (R), 
with scholarship coordinator, Dr. Shahnaz Sultan

Tufte, an independent consultant who leads patient-public partnership initiatives, presented on the unique opportunities of using patient partners during the development of GRADE guidelines. Patient partners are representatives of the patient population whom the guideline aims to serve. As part of a guideline panel, they offer fresh perspectives, ground the guideline development process with lived experience, and help the panel to identify and address differences in priorities among stakeholders.

Throughout her presentation, Tufte provided ways to improve how patient partners are involved in the guideline process, such as creating one-pagers and glossaries that cover the basics of GRADE methodology and inquiring beforehand about specific accommodations that might be needed in order to enhance the patient’s participation in the panel. “It was an honor to attend the GRADE Workshop in Denver as a Patient Partner Scholar,” said Tufte. “I felt like I was treated like a colleague where we were all learning together how to use GRADE tools to share best evidence within our individual systems and guidelines work.”

Dr. bin Riaz, an oncologist at Mayo Clinic, presented on a framework for developing living systematic reviews and guidelines to inform clinical decision-making, especially in topic areas undergoing rapid change. It can take several years for a systematic review and resulting clinical recommendations to be developed, Dr. bin Riaz explained. In the meantime, new drug approvals or indications, changes in drug labeling, or new information about potential risks and benefits of a treatment option can arise. As opposed to traditional, static documents, living systematic reviews and guidelines are continually updated as new evidence or important decision-making information comes to light. The ultimate goal of such an approach is to facilitate a more timely translation of medical knowledge into clinical practice, allowing patients and their providers to come to decisions informed by the totality of current evidence.

If interested in applying for a scholarship to future GRADE workshops, more details can be found here: https://evidencefoundation.org/scholarships.html. Please note the deadline for applications to our next workshop in Orlando, Florida will be July 1, 2019.





Wednesday, January 30, 2019

Research Shorts: Borrowing of strength from indirect evidence

Contributed by M. Hassan Murad, M.D.

Network meta-analysis (NMA) is supposed to increase precision (because it includes more studies in the analysis); however, this is not always the case. This empirical study evaluated 915 possible treatment comparisons. The study used the recently proposed borrowing of strength statistic, which quantifies the percentage reduction in the uncertainty of the effect estimate when adding indirect evidence to an NMA



When only one study contributed direct evidence, NMAs resulted in reduced precision and no appreciable improvements in precision in 57.5% and 12.7% comparisons, respectively. When at least two studies contributed direct evidence, NMAs provided increased precision in 66.4% comparisons. The bottom-line is that, in sparse networks (i.e., networks that mostly consist of indirect evidence), NMA may not improve precision as much as stakeholders want and expect.

Reference: Lin L, Xing A, Kofler MJ, Murad MH. Borrowing of strength from indirect evidence in 40 network meta-analyses. Journal of clinical epidemiology. 2018 Oct 17.

Stating the “Obvious”: A Primer on Good Practice Statements in GRADE Guidelines

Contributed by Madelin Siedler, 2018/2019 U.S. GRADE Network Research Fellow


Stating the “Obvious”: A Primer on Good Practice Statements in GRADE Guidelines

One of the benefits of the GRADE approach is that it provides a framework for the development of evidence-based recommendations that are clear and actionable for practicing clinicians even when only lower-quality evidence is available. However, in some particular instances, caution is warranted when developing a recommendation based on low-quality evidence or inference. “Good practice statements” are one such instance. 

The term “good practice statement” is sometimes interchanged for “motherhood statement”.” In either case, the practice being recommended is usually something that is already commonly accepted as beneficial or practical advice. It could even be seen as irrefutably “good” as motherhood and apple pie (hence the term). The nature of these types of statements is such that the action is seen as so obviously beneficial that it would be unduly onerous to conduct a review to demonstrate its efficacy.

An example of a good practice statement is the first recommendation from the American Gastroenterological Association (AGA)’s 2015 guideline on the management of asymptomatic pancreatic cysts, which reads, “The AGA recommends that before starting any pancreatic cyst surveillance program, patients should have a clear understanding of programmatic risks and benefits” (Vege et al., 2015).

How to Spot a Good Practice Statement
An easy way to identify a good practice statement is to restate the recommendation as its inverse: for instance, “patients should not have a clear understanding of programmatic risks and benefits.” If this “unstated alternative” is absurd or clearly does not conform to ethical norms, the original statement is likely a good practice statement (Guyatt et al., 2016).

The Problem with Good Practice Statements
The GRADE Working Group recommends that good practice statements be used sparingly, if at all. Because good practice guidance is typically based on several linked sources of indirect evidence, there is no way to tell whether the benefits of the proposed recommended action are as truly obvious or incontestable as they seem. And if they are (if the inverse of the proposed recommendation would be absurd or unethical) then they are likely unwarranted and can dilute the strength of the guideline as a whole. 

Sometimes, good practice statements may even appear as graded recommendations in a guideline (a decision that’s not recommended by the GRADE working group for the reasons above). In this case, the guideline authors may be tempted to make a strong recommendation based on low-quality evidence, which should ideally be a rare occurrence and based on well-defined criteria (Guyatt et al., 2015). 

Practical Advice for Dealing with Good Practice Statements
The GRADE Working Group recommends using the following checklist to determine whether a good practice statement is warranted:
  1. Is the statement clear and actionable? 
  2. Is the message really necessary in regards to actual health practice? 
  3. After consideration of all relevant health outcomes and potential downstream consequences, will implementing the good practice statement result in large net positive consequences? 
  4. Is collecting and summarizing the evidence a poor use of a guideline panel’s limited time and energy? 
  5. Is there a well-documented clear and explicit rationale connecting the indirect evidence?
If the answer is 'yes' to all five questions, a good practice statement may be warranted for inclusion in a guideline document. When done correctly, good practice advice statements should appear as ungraded recommendations, meaning no formal rating of quality of evidence or strength of recommendation should be given (Guyatt et al., 2015).

However, many potential good practice statements will be eliminated through the use of this checklist. For instance, careful consideration of Question #3 could lead the panel to realize that the assumed net positive of a specific action may not be so obvious after all. In this case, the guideline panel should consider whether a thorough review of the evidence should be conducted and formal grading methods applied.