Showing posts with label Research shorts. Show all posts
Showing posts with label Research shorts. Show all posts

Monday, October 17, 2022

Use of an Evidence-to-Decision Framework is Associated with Better Reporting, More Thorough Consideration of Recommendations

In the guideline development process, a panel should use a defined framework to consider multiple aspects of a clinical decision, including but not limited to the certainty of the underlying evidence, potential impact on resource use, or variability in the values and preferences of patients and other stakeholders. Such frameworks include the GRADE Evidence-to-Decision (EtD) format as well as others such as the "decision-making triangle" and Guidance for Priority-Setting in Health care (GPS Health). 

To better understand the prevalence and use of these various frameworks within guidelines, Meneses-Echavez and colleagues systematically searched for guidelines and related guideline production manuals published between 2003 and May 2020. Items were screened and extracted by two independent authors, with a total of 68 full text documents included and analyzed.

Of these documents, most (93%) reported using a structured framework to assess the certainty of evidence, about half (53%) of which used GRADE or adapted systems based on GRADE (10%). Similarly, 88% of documents reported using a framework to rate the strength of recommendations, with about half (51%) using the GRADE approach. However,  only about two-thirds (66%) of the included documents explicitly stated the process for formulating resulting recommendations. 

Finally, the GRADE framework  was most commonly used for the evidence-to-decision, being cited in 42% of the included articles, with other reported frameworks including NICE (8%), SIGN (8%) and USPSTF (4%). Articles using the GRADE EtD framework reported considering more criteria than those using alternative approaches. The most commonly used criteria across documents included desirable effects (72%), undesirable effects (73%), and the certainty of evidence of effects (73%); the least commonly applied criteria were acceptability (28%), certainty of the evidence of required resources (25%), and equity (16%). 



The use of any EtD framework was associated with a greater likelihood of incorporating perspectives (odds ratio: 2.8; from 0.6-13.8) and subgroup considerations (odds ratio:7.2; from 0.9-57.9), as was the use of GRADE compared to other EtDs (odds ratios: 1.4 and 8.4). These differences also affected whether justifications were reported for each judgment as well as the inclusion of notes to consider for the implementation of recommendations and for monitoring and evaluating recommendations.  

The authors conclude that guidance documents stand to benefit from the more explicit reporting of how recommendations are formulated, from the initial grading of the certainty of underlying evidence to the consideration of how recommendations will affect various criteria such as resource use and equity. These changes, in the words of the authors, may "enhance transparency and credibility, enabling end users to determine how much confidence they can have in the recommendations; facilitate later adaptation to contexts other than the ones in which they were originally developed; and improve usability and communicability of the EtD frameworks." 

Meneses-Echaves JF, Bidonde J, Yepes-Nuñez JJ, et al. (2022).  Evidence to decision frameworks enabled structured and explicit development of healthcare recommendations. J Clin Epidemiol 150:51-62. Manuscript available at publisher's website here.









Tuesday, September 27, 2022

Only One-Third of a Sample of RCTs Had Made Protocols Publicly Available, New Report Finds

Earlier this year, a study in PloS Medicine found that nearly one-third (30%) of a sample of randomized controlled trials (RCTs) had been discontinued prematurely, a number that had not improved over the previous decade. Furthermore, for every 10% increase in adherence to SPIRIT protocol reporting guidelines, RCTs were 29% less likely to go unpublished (OR: 0.71; 95% confidence interval: from 0.55 to 0.92), and only about 1 in every 5 unpublished trials had been registered.

Now, in this month's issue of Journal of Clinical Epidemiology, Schönenberger and colleagues have released a study of the availability of RCT protocols from a sample of published works.

Public availability of study protocols, the authors argue, improves research quality by promoting thoughtfulness in methodological design, reducing selective outcomes reporting or "cherry-picking," and reducing the misreporting of results while promoting ethical compliance. This is especially the case when trial protocols are made available before the publication of study results. 

From a random sample of RCTs approved by ethics committees in Switzerland, Germany, Canada, and the United Kingdom in 2012, the authors examined the proportion of studies that had publicly available protocols and the nature of how the protocols were cited and disseminated. Of the resulting 326 RCTs, 118 (36.2%) had publicly available protocols. Of the protocols, nearly half (47.5%) were available as standalone peer-reviewed publications while 40.7% were available as supplementary material with the published results. A smaller proportion (10.2%) of protocols were available on a trial registry. 

Studies with a sample size of >500 or that were investigator- (non-industry)-sponsored were more likely to have publicly available protocols. The nature of the intervention (drug versus non-drug) did not appear to affect protocol availability, nor did whether the trial was conducted in a multicenter or single-center setting. The majority (91.8%) of protocols were made available after the enrollment of the first patient, and just 2.7% were made available after publication of trial results. Protocols were commonly published shortly before the trial results, at a median of 90% of the time between the start of the trial and its publication.

As this sample comprised only RCTs published in 2012 and by relatively high-income countries, it is unclear whether public protocol availability has improved over time or may be different in other global regions. However, the authors argue, these numbers lend credence to the need for efforts to improve the public availability of RCT protocols, such as through trial registries or requirements by publishing or funding bodies.

Schönenberger, C.M., Griessbach, A., Heravi, A.T., et al. (2022). A meta-research study of randomized controlled trials found infrequent and delayed availability of protocols. J Clin Epidemiol 149:45-52. Manuscript available at publisher's website here










Wednesday, September 21, 2022

Standardized Mean Difference Estimates Can Vary Widely Depending on the Methods Used, New Review Finds

In meta-analyses of outcomes that utilize multiple scales of measurement, a standardized mean difference (SMD) may be used. Randomized controlled trials may also use SMDs to help interpret the effect size for readers. Most commonly, the SMD reports the effect size with Cohen's d, a metric of how many standard deviations are contained in the mean difference within or between groups (e.g., an intervention caused the outcome to increase or decrease by x number of standard deviations, or the two groups were x number of standard deviations different from one another with regards to the outcome). This is typically done by dividing the difference between groups, or from pretest to posttest in a single group, by some form of standard deviation (e.g., pooled standard deviation at baseline, posttest, or the standard deviation of change scores). Cohen's d is often utilized because a general rule of interpretation has been suggested: 0.2 is a small effect, 0.5 is a medium-sized effect, and 0.8 is large.

However, there are multiple ways to approach the calculation of SMDs, and these may result in varying interpretations of the size of the effect. To further investigate this, Luo and colleagues recently published a review of 161 articles using SMDs and the way they can be calculated. Of the 161 randomized controlled trials published since 2000 and reporting outcomes with some form of SMD, the authors calculated potential between-group SMDs using reported data and up to seven different methodological approaches.

Some studies reported more than one type of SMD, meaning that 171 total SMD approaches were reported across the 161 studies. Of these, 34 (19.9%) did not describe the chosen method at all, 84 (49.1%) reported but in insufficient detail, and 53 (31%) reported the approach in sufficient detail. The confidence interval was only reported for 52 (30.4%) of SMDs. Of the 161 individual articles, the rule for interpretation was clearly stated in only 28 (17.4%). 

The most common method of calculating SMD was using a standard deviation of baseline scores, seen in 70 (40.9%) of studies. Meanwhile, 30 (17.5%) used posttest standard deviations and 43 (25.1%) used the standard deviation of change scores.

Figure displaying the variability of SMD estimates across 161 included studies. Click to enlarge.

Of all the potential ways to calculate SMD, the median article varied by 0.3 - which could potentially be the difference between a "small" and "moderate" or "between a "moderate" and "large effect size for Cohen's d using Cohen's suggested rule of thumbThe studies with the largest variation tended to have smaller sample sizes and greater reported effect sizes.

This work raises an important point, which is that while no one method for the calculation of SMDs is considered superior to another, if calculation approaches are not prespecified by researchers, different methods could be tried until the most impressive effect size is reached. To help prevent these issues, the authors suggest prespecifying the analytical approach and reporting SMDs together with raw mean differences and standard deviations to further aid interpretation and provide context. 

Luo, Y., Funada, S., Yoshida, K., et al. (2022). Large variation existed in standardized mean difference estimates using different calculation methods in clinical trials. J Clin Epidemiol 149: 89-97. Manuscript available at the publisher's website here.










Thursday, September 8, 2022

New Study Sheds Light on the Perks of Copublication of Cochrane Systematic Reviews

Since 1994, Cochrane has allowed the publication of certain systematic reviews to extend beyond the eponymous Database of Systematic Reviews and into the pages of a medical specialty journal with the hopes of increasing dissemination. Often, these copublished reviews will include an abridged version of the Cochrane review as well as commentary or other additional features explaining the review and its findings.



A new retrospective cohort study published in this month's issue of the Journal of Clinical Epidemiology highlights the benefits of such an approach. In brief, Zhu and colleagues investigated the rate of citation of Cochrane systematic reviews that had been copublished in a second journal versus those that were published only in the Database. Using a ratio of 2:1, the authors matched two randomly selected noncopublished reviews for each copublished review that came up in an indexed journal that had copublication agreements with Cochrane.

Out of the resulting sample of 101 copublished and 202 noncopublished reviews, the median number of citations over the first five years since publication was higher for reviews that had been copublished (71 versus 32.5, or approximately 118% higher in the copublished reviews). The median as higher for copublished reviews for the first, second, third, and fifth years specifically as well. In 19% of journals, the copublication of a review led to a subsequent increase in impact factor over the following year; this was true for 27.3% of journals during the second year. There was no clear trend over time for the rate of copublication across journals, though the total number of Cochrane reviews published during the same period generally increased. 

Zhu, L., Zhang, Y., Yang, R., et al. (2022). Copublication improved the dissemination of Cochrane reviews and benefited copublishing journals: A retrospective cohort study. J Clin Epidemiol 149:110-117. Manuscript available at publisher's website here. 









Wednesday, May 18, 2022

New Study of Randomized Trial Protocols Highlights Prevalence of Early Discontinuation, Importance of Reporting Adherence

Registration of clinical trials was first introduced as a way to combat publication bias and avoid the duplication of efforts in medical research. Since then, the registration of clinical trials has been deemed a requirement of publication by the International Committee of Journal Editors (ICJE) and various federal laws. However, the mere registration of a clinical trial does not guarantee its ultimate completion, nor publication.

In a new article published last month in PLoS Medicine, Speich and colleagues set out to better understand the prevalence and impact of non-registration, early trial discontinuation, and non-publication within the current landscape of medical research.

An update of previous findings published in 2014, the study examined 360 protocols for randomized controlled trials (RCTs) approved by Research Ethics Committees (RECs) based in Switzerland, the United Kingdom, Germany, and Canada. Of these, 326 were eligible for further analysis. The team collected data on whether the trials had been registered, whether sample sizes were planned and achieved, and whether study results were available or published. The team also assessed each of the RCTs for protocol reporting quality based on the Standard Protocol Items: Recommendations for Intervention Trials (SPIRIT) guidelines. 

Overall, the included RCTs had a median planned sample size of 250 participants and met a median of 69% of SPIRIT guidelines. Just over half (55%) were sponsored by industry. The large majority (94%) were registered, though 10% were registered retrospectively. About half (53%) reported results in a registry, whereas most (79%) had published results in a peer-reviewed journal. However, adherence to reporting guidelines did not appear to the rate of trial discontinuation.

About one-third (30%) of RCTs had been prematurely discontinued, which indicated no change in this regard since the previous investigation of RCT protocols approved between 2000-2003 (28%). Most commonly, trials in the current investigation were discontinued due to preventable reasons such as poor recruitment (37%), organizational/strategic issues (6%), and limited resources (1%). A smaller proportion were discontinued due to preventable reasons such as futility (16%), harm (6%), benefit (3%), or external evidence (3%).

For every 10% increase in adherence to SPIRIT protocol reporting guidelines, RCTs were 29% less likely to go unpublished (OR: 0.71; 95% confidence interval: from 0.55 to 0.92), and only about 1 in every 5 unpublished trials had been registered.

The authors suggest that these findings implicate the need for investigators to report results of trials in registries as well as in peer-reviewed publications. Furthermore, future research may assess the utility of feasibility or pilot studies in reducing the rate of trial discontinuation due to recruitment issues. Journals can require trial registration as a requirement to publish. 

Speich, B., Gryaznov, D., Busse, J.W., lohner, S., Klatte, K., Heravi, A.T., ... & Briel, M. (2022). Nonregistration, discontinuation, and nonpublication of randomized trials: A repeated research meta-analysis. PLoS Medicine 19(4), e1003980. https://doi.org/10.1371/journal.pmed.1003980. Manuscript available at the publisher's website here.



Monday, April 11, 2022

It's Alive! Pt. IV: Results from a Trainee Living Systematic Review Experience

Living systematic reviews (LSRs) continue to be a topic of interest among systematic review and guideline developers, as evidenced by our history posts on the topic here, here, and here. While automation and machine learning have begun to help facilitate what is a generally time- and resource-intensive process to evidence syntheses perpetually up-to-date, some aspects of LSR development still require the human touch. Now, a recently published mixed-methods study discusses the successes and challenges of utilizing a crowdsourcing approach to keep the LSR wheels turning.

The article describes the process of involving trainees in the development of a living systematic review and network meta-analysis (NMA) on drug therapy for rheumatoid arthritis. In their report, the authors posit that evidence-based medicine is a key pillar of learning for trainees, but that they may learn better through an experiential rather than a purely didactic approach; providing the opportunity to participate in a real-life systematic review may provide this experiential learning. 

In short, the team first applied machine learning to sort through an initial database to filter out randomized controlled trials, which was then further assessed by a crowdsourcing platform, Cochrane Crowd. Next, trainees ranging from undergraduate students to practicing rheumatologists and researchers recruited through Canadian and Australian rheumatology mailing lists further assessed articles for eligibility and extracted data from included articles. 

Training included a mix of online webinars, one-on-one trainings, and handbook provisions. Conflicting results were further assessed by an expert member of the team. The authors then elicited both quantitative and qualitative feedback about the trainees' experiences of taking part in the project through a combination of electronic survey and one-on-one interviews. 

Overall, the 21 trainees surveyed rated their training as adequate and experience generally positive. Respondents specifically listed better understanding of PICO criteria, familiarity with outcome measures used in rheumatology, and the assessment of studies' risk of bias as the greatest learning benefits obtained. 

Of the 16 who participated in follow-up interviews, the majority (94%) described a practical and enjoyable experience. Of particular positive regard was the use of task segmentation throughout the project, during which specific tasks (i.e., eligibility assessment versus data extraction) could be "batch-processed," allowing trainees to match the specific time and focus demands to the selected task at hand. Trainees also communicated an appreciation for the international collaboration involved in the review as well as the feeling of meaningfully contributing to the project. 

Notable challenges included issues related to the clarity of communication regarding deadlines and expectations, as well as technical glitches experienced through the platforms used for screening and extraction. Though task segmentation was seen as a benefit, it also included drawbacks: namely, the risk of more repetitive tasks such as eligibility assessment becoming tedious while others that require more focus (i.e., data extraction) may be difficult to integrate into an already-busy daily schedule. To address these issues, the authors suggest improving communications to include regular, frequent updates and deadline reminders, working through technological glitches, and carefully matching tasks to the specific skillsets and availabilities of each trainee.

Lee, C., Thomas, M., Ejaredar, M., Kassam, A., Whittle, S.L., Buchbinder, R., ... & Hazlewood, G.S. (2022). Crowdsourcing trainees in a living systematic review provided valuable experiential learning opportunities: A mixed-methods study. J Clin Epidemiol (in-press). Manuscript available at the publisher's website here.











Wednesday, September 2, 2020

A New Tool for Assessing the Credibility of Effect Modification Cometh: Introducing the ICEMAN

Effect modification goes by many other names: “subgroup effect,” “statistical interaction,” and “moderation,” to name a few. Regardless of what it’s called, the existence of effect modification in the context of an individual study means that the effect of an intervention varies between individuals based on an attribute such as age, sex, or severity of underlying disease. Similarly, a systematic review may aim to identify effect modification between individual studies based on their setting, year of publication, or methodological differences (often called a “subgroup analysis”).

As many as one-quarter of randomized controlled trials (RCTs) and meta-analyses examine their findings for potential evidence of effect modification, according to a paper by Schandelmaier and colleagues published in the latest edition of CMAJ. However, it is not uncommon for claims of effect modification to be later proved spurious, which may negatively affect the quality of care in those subgroups of patients. Potential sources of these claims range from simple random chance to issues with selective reporting and misguided application of statistical analyses.


Click to enlarge.


In “Development of the Instrument to assess the Credibility of Effect Modification in Analyses (ICEMAN) in randomized controlled trials and meta-analyses,” the authors present a novel tool for evaluating the presence of a potential modifier. While several sets of criteria have been developed in the past for this purpose, the ICEMAN is the first to be based on a rigorous development process and refined with formal user testing.

 

First, the authors conducted a systematic survey of the literature to ensure a comprehensive understanding of the previously proposed criteria for evaluating effect modification. Thirty sets were identified, none of which adequately reflected the authors’ conceptual framework. Second, an expert panel of 15 members was identified randomly from a list of 40 identified through the systematic survey. These experts then pared down the initial list of 36 candidate criteria to 20 required and eight optional items. After developing a manual for its use, the authors tested the instrument among a diverse group of 17 potential users, including authors of Cochrane reviews and RCTs and journal editors using a semi-structured interview technique.


Schandelmaier, S., Briel, M., Varadhan, R., Schmid, C.H., Devasenapathy, N., Hayward, R.A., Gagnier, J., ... & Guyatt, G.H. 2020. Development of the Instrument to assess the Credibility of Effect Modification Analyses (ICEMAN) in randomized controlled trials and meta-analyses. CMAJ 192:E901-906.


Manuscript available at the publisher's website here

Friday, July 10, 2020

Room for Improvement: Use of Cochrane RoB tool in non-Cochrane Systematic Reviews is Largely Incomplete

The Cochrane Risk of Bias (RoB) tool for randomized controlled trials (RCTs) is commonly used in both Cochrane and non-Cochrane systematic reviews as a standardized way to assess and report the risk of bias within a study or a body of evidence. The tool comprises seven domains, each representing a potential source of bias within the design or execution of an RCT. Judgments for each domain (for instance, allocation concealment, or selective outcome reporting) are made between whether the study possessed a low, high, or unclear risk of bias from that source.

A new review of non-Cochrane systematic reviews (NCSRs) published in this month’s edition of the Journal of Clinical Epidemiology reports that the use of the Cochrane RoB tool in these reviews is incomplete or inadequate in most cases. Within 508 eligible systematic reviews that used the original (2011) Cochrane RoB tool published through 3 July 2018, the majority (85%) reported the analysis of risk of bias; within these papers, about half (53%) used the Cochrane tool specifically, leaving a total of 269 reviews for further analysis.

A non-negligible minority of studies included in the review by Puljak et al. either did not include certain domains of the Cochrane RoB tool, or did not report which domains were used. Only 40% of the reviews analyzed RoB through all seven domains. Click to enlarge.

Less than two-thirds (60%) of the 269 included reviews used all seven domains of the Cochrane tool, report Puljak and colleagues, and only 16 of the included reviews (5.9%) reported both a judgment and a comment explaining each judgment either within the manuscript or in a supplementary file. Within these 15 reviews, the proportion of inadequate judgments (either those in which the comment was not in line with the judgment or in which there was no supporting comment) ranged from 25% (Other Bias domain) to 65% (Selective Reporting Bias domain). The reviews “rarely” included full tables illustrating the RoB judgments for the different domains.

The authors’ findings highlight that both a judgment (low/high/unclear risk of bias) as well as a comment explaining the judgment within each domain should be included in systematic reviews that report use of the Cochrane RoB tool.

Puljak, L., Ramic, I., Naharro, C.A., Brezova, J., Lin, Y.C., Surdila, A.A,.... & Salvado, M.S. Cochrane risk of bias tool was used inadequately in the majority of non-Cochrane systematic reviews. J Clin Epidemiol, 2020; 123: 114-119.

Manuscript available from publisher's website here. 

Tuesday, July 7, 2020

Adventures in Protocol Publication Pt. II: Survey of PROSPERO Registrants Finds Peer-Reviewed Protocol Publication is a Mixed Bag

As we discussed in Part I of this series, the posting of a systematic review – for instance, in an online registry such as PROSPERO -  is a well-established practice that helps prevent results-driven biases from being introduced into a systematic review as well as reduces unintentional duplication of efforts. The additional publication of such a protocol in a peer-reviewed journal also has its benefits – such as additional citations, saving room in the methods portion of the final manuscript, and the facilitation of the actual systematic review’s publication. However, it also has its drawbacks, including the potential cost of publishing and the time required from submission to acceptance, which may result in a delayed SR timeline.

In the latest issue of the Journal of Clinical Epidemiology, Rombey and colleagues report the findings of their survey examining the practices and attitudes related to peer-reviewed protocol publication among over 4,000 authors of non-Cochrane systematic review protocols published in PROSPERO in 2018.

In order to identify potential “inhibiting factors” of peer-reviewed protocol publication, respondents who reported publishing their protocol in peer-reviewed journal in addition to PROSPERO were asked questions related to publication costs, while respondents who had not pursued this option were asked for their reasoning behind this decision.

Nearly half (44.7%) of the 4,054 respondents answered that they had published or plan to publish their protocol in a peer-reviewed journal, while the remaining 55.3% chose not to pursue this option. Of these respondents, the most common reasons given were that publishing the protocol in PROSPERO was deemed sufficient; that it was not an aim/priority of the authors; and that the time required to publish in a peer-reviewed journal would delay the review itself.

A figure from Rombey et al. shows level of agreement with statements related to peer-reviewed protocol publication among >4,000 respondents. Click to enlarge.

Of those who did report publishing their protocol, about one-third (67.9%) indicated that there were no costs associated with publication, while one-quarter (25.4%) indicated that there were costs, and the remaining 6.7% were not sure.

Respondents from Africa, Asia, and South America were twice as likely to report having published a protocol in a peer-reviewed journal than those from Europe, North America, or Oceania; however, likelihood of publication did not appear to vary by gender or experience level. Qualitative analysis of free-response text revealed that some respondents were not aware that publication of protocols in peer-reviewed journals was done at all.

Overall, while cost of publishing a protocol did not appear to be a major inhibiting factor for most respondents, issues related to time from submission to publication as well as opinions regarding the additional value of publishing beyond PROSPERO were the most commonly cited reasons for not pursuing peer-reviewed publication of a systematic review protocol.

Rombey, T., Puljak, L., Allers, K., Ruano, J., & Pieper, D. Inconsistent views among systematic review authors toward publishing protocols as peer-reviewed articles: An international survey. J Clin Epidemiol, 2020; 123:9-17.

Manuscript available from the publisher's website here.

Wednesday, July 1, 2020

A Not-So-Non-Event?: New Systematic Review Finds Exclusion of Studies with No Events from a Meta-Analysis Can Affect Direction and Statistical Significance of Findings

Studies with no events in either arm have been considered non-informative within a meta-analytical context, and thus have been left out of these analyses. A new systematic review of 442 such meta-analyses, however, reports that this practice may actually affect the resulting conclusions.

In the July 2020 issue of the Journal of Clinical Epidemiology, Xu and colleagues report their study of meta-analyses of binary outcomes in which at least one included study had no events in either arm. The authors then reanalyzed the data from 442 included papers taken from the Cochrane Database of Systematic Reviews, using modeling to determine the effect of reincorporating the excluded study.

The authors found that in 8 (1.8%) of the 442 meta-analyses, inclusion of the previously excluded studies changed the direction of the pooled odds ratio (“direction flipping”). In 12 (2.72%) of the meta-analyses, the pooled odds ratio (OR) changed by more than the predetermined threshold of 0.2. Additionally, in 41 (9.28%) of these studies, the statistical significance of that findings changed when assuming a p = 0.05 threshold (“significance flipping”). In most of these 41 meta-analyses, excluded (“non-event”) studies made up between 5 and 30% of the total sample size. About half of these alterations led to an expansion of the confidence interval; while in the other half, the incorporation of non-events reduced the confidence interval.

The figure above from Xu et al. shows the proportion of studies reporting no events within the meta-analyses that showed a substantial change in p value when these studies were included. The proportion of the total sample tended to cluster between 5 and 30%. Click to enlarge.

Post hoc simulation studies confirmed the robustness of these findings, and also found that exclusion of studies with no events preferentially affected the pooled ORs of studies that found no effect (OR = 1), whereas a large magnitude of effect was protective against these changes. The opposite was found for the effect of excluding studies with no events on the resulting p values (i.e., large magnitudes of effects were more likely to be affected whereas conclusions of no effect were protected).

In sum, though a common practice in meta-analysis, the exclusion of studies with no events in either arm may affect the direction, magnitude, or statistical significance of the resulting conclusions in a small but non-negligible number of analyses.

Xu, C., Li, L, Lin, L., Chu, H., Thabane, L., Zou, K., & Sun, X. Exclusion of studies with no events in both arms in meta-analysis impacted the conclusions. J Clin Epidemiol, 2020; 123: 91.99.

Manuscript available from the publisher's website here. 

Tuesday, June 23, 2020

Need for Speed: Documenting the Two-Week Systematic Review

In a recent post, we summarized a 2017 article describing the ways in which automation, machine learning, and crowdsourcing can be used to increase the efficiency of systematic reviews, with a specific focus on making living systematic reviews more feasible.

In a new publication in the May 2020 edition of the Journal of Clinical Epidemiology, Clark and colleagues incorporated automation in order to attempt systematic review that took no longer than two weeks from search design to manuscript submission for a moderately-sized search yielding 1,381 deduplicated records and eight ultimately included studies.

Spoiler alert: they did it. (In just 12 calendar days, to be exact).

Systematic Review, but Make it Streamlined

Clark et al. utilized some form of computer-assisted automation at almost every point in the project, including:
  • Using SRA word frequency analyzer to identify key terms that would be most helpful inclusions in a search strategy
  • Using hotkeys (custom keystroke shortcuts) within SRA Helper tool to more quickly screen items and search pre-specified databases for full texts
  • Using RobotReviewer to assist in risk of bias evaluation by searching for certain key phrases within each document
However, machines were only part of the solution. The authors also note the decidedly more human-based solutions that allowed them to proceed at an efficient clip, such as:
  • Daily, focused meetings between team members
  • Blocking off “protected time” for each team member to devote to the project
  • Planning for deliberation periods, such as decisions on screening conflicts, to occur immediately after screening so as to reduce the amount of time and energy devoted to “mental reload” and review of one’s previous decisions for context
Time Distribution of 12-Day Systematic Review by Task. Click to enlarge.

All told, the final accepted version of the manuscript took 71 person-hours to complete – a far cry from a recently published average of 881 person-hours among conventionally conducted reviews.

Clark and colleagues discuss key facilitators and barriers to their approach as well as provide suggestions for technological tools to further improve the efficiency of SR production.

Clark, J., Glasziou, P., Del Mar, C., Bannach-Brown, A., Stehlik, P., & Scott, A.M. A full systematic review was completed in 2 weeks using automation tools: A case study. J Clin Epidemiol, 2020; 121: 81-90.

Manuscript avaliable from the publisher's website here.

Thursday, June 18, 2020

It’s Alive!: Pt. III: From Living Review to Living Recommendations

In recent posts, we’ve discussed how living systematic reviews (LSRs) can help improve the currency of our understanding of the evidence, as well as the efficiency with which the evidence is identified and synthesized through novel crowdsourcing and machine learning techniques. In the fourth and final installment of the 2017 series on LSRs, Akl and colleagues apply the LSR approach to the concept of a living clinical practice guideline.

As the figure below from the paper demonstrates, while simply updating an entire guideline more frequently (Panel B) reduces the number of out-of-date recommendations (symbolized by red stars) at any given time, it comes with a serious trade-off: namely, the high amount of effort and time required to continuously update the entire guideline. Turning certain recommendations into "living" models helps solve this dilemma between currency and efficiency. Click to enlarge.













Rather than a full update of an entire guideline and all of the recommendations therein, a living guideline uses each recommendation as a separate unit of update. Recommendations that are eligible to make the transition from “traditionally updated” to “living” include those that are a current priority for healthcare decision-making, for which the emergence of new evidence may change clinical practice, and for which new evidence is being generated at a quick rate.

The Living Guideline Starter Pack

Each step of a recommendation’s formation must make the transition to “living,” including:
  • A living systematic review
  • Living summary tables, such as Evidence Profiles and Evidence-to-Decision tables
  • Online collaborative table-generating software such as GRADEpro can be used to keep these up-to-date with the emergence of newly relevant evidence
  • A living guideline panel who can remain “on-call” to contribute to updates of recommendations with relatively short notice when warranted
  • A living pool of peer-reviewers who can review and provide feedback on updates with a quick turnaround time
  • A living publication platform, such as an online version that links back to archived versions, as well as “pushes” new versions to practice tools at the point of care.
Additional Resources
Further information and support for the development of LSRs, including updated official guidance, is provided on the Cochrane website.

Akl, E.A., Meerpohl, J. J., Elliott, J., Kahale, L. A., Schünemann, H.J., and the Living Sysematic Review Network. Living systematic reviews: 4. Living guideline recommendations. J Clin Epidemiol, 2017; 91: 47-53.

Manuscript available from the publisher's website here. 

Monday, June 15, 2020

It’s Alive! Pt. II: Combining Human and Machine Effort in Living Systematic Reviews

Systematic review development is known to be a labor-intensive endeavor that require a team of researchers dedicated to the task. The development of a living systematic review (LSR) that is continually updated as newly relevant evidence becomes available presents additional challenges. However, as Thomas and colleagues write in the second installment of the 2017 series on LSRs in the Journal of Clinical Epidemiology, we can make the process quicker, easier, and more efficient by harnessing the power of machine learning and “microtasks.”

Suggestions for improvements in efficiency can be categorized as either automation (incorporation of machine learning/replacement of human effort) or crowdsourcing (distribution of human effort across a broader base of individuals).

A diagram from Thomas et al. (2017) describes the "push" model of evidence identification that can help keep Living Systematic Reviews current without the need for repeated human-led searches. Click to enlarage.

From soup to nuts, opportunities for the incorporation of machine learning into the LSR development process include:


  • Continuous, automatic searches that “push” new potentially relevant studies out to human reviewers
  • Exclusion of ineligible citations through automatic text classification, reducing the number of items that require human screening with over 99% sensitivity
  • Crowdsourcing of study identification and "microtask" screening efforts such as Cochrane Crowd, which at the time of this blog’s writing had resulted in over 4 million screening decisions from over 17,000 contributors 
  • Automated retrieval of full text versions of included documents
  • Machine-based extraction of relevant data, graphs and tables from included documents
  • Machine-assisted risk of bias assessment
  • Template-based reporting of important items
  • Statistical thresholds that flag when a change of conclusions may be warranted
As technology in this field progresses, the traditionally duplicated stages of screening and data extraction may even be taken on by a computer-human pair, combining the ease and efficiency of automation with the “human touch” and high-level discernment that algorithms still lack.

Thomas, J.,  Noel-Storr, A., Marshall, I., Wallace, B., McDonald, S., Mavergames, C... & the Living Systematic Review Network. Living systematic reviews: 2. Combining human and machine effort. J Clin Epidemiol, 2017; 91: 31-37. 

Manuscript available from publisher's website here. 

Wednesday, June 10, 2020

It’s Alive! Pt. I: An Introduction to Living Systematic Reviews

As research output continues to rise, the systematic reviews charged with comprehensively identifying and synthesizing the evidence within them are becoming more quickly out-of-date. In addition, the formation of a systematic review team can be a lengthy process, and institutional memory of the project is lost when teams are disbanded after publication.

One solution to this problem is the concept of a living systematic review, or LSR. In the first installment of a 2017 series in the Journal of Clinical Epidemiology, Elliott and colleagues introduce the concept of an LSR and provide general guidance on their format and production.

What is a Living Systematic Review (LSR)?
An LSR has a few key components:
  • Based on a regularly updated search run with an explicit and pre-established frequency (at least once every six months) to identify any potentially relevant recent publications. 
  • Utilize standard systematic review methodology (different from a rapid review)
  • Most useful for specific topics:
    • that are of high importance to decision-making,
    • for which the certainty of evidence is low or very low (meaning our certainty of the effect may likely change with the incorporation of new evidence), and
    • for which new evidence is being generated often.


A figure from Elliott et al. (2017) provides an overview of the LSR development process, from protocol to regular searching and screening and incorporation and publication of new evidence.


LSRs from End to End
An LSR can either be started from scratch with the intention of regular screening and updating of evidence – in which case the protocol should specify these planned methods – or based upon an existing up-to-date systematic review, in which case the protocol should be amended to reflect these changes.

Due to their nature, the publication of LSRs requires the use of an online platform with linking mechanisms (such as CrossRef) or with explicit versions (such as the Cochrane database) that can be updated as soon as new evidence is incorporated.

When the certainty of evidence reaches a higher level, or if the generation of new evidence substantially slows, an LSR may be discontinued in favor of traditional approaches to updating.

Additional Resources
Further information and support for the development of LSRs, including updated official guidance, is provided on the Cochrane website.

Elliott, J.H., Synnot, A., Turner, T., Simmonds, M., Akl, E.A., McDonald, S... & Thomas, J. Living systematic review: 1. Introduction - the why, what, when, and how. J Clin Epidemiol, 2017; 91:23-30.

Manuscript available from the publisher's website here.


Friday, June 5, 2020

Research Revisited: 2014’s “Guidelines 2.0: Systematic Development of a Comprehensive Checklist for a Successful Guideline Enterprise”

While several checklists for the development and appraisal of specific guidelines had been developed by 2014, there had yet to be published a thorough and systematic resource for organizations to inform the actual day-to-day operations of a guideline development program. Noticing this need, Schünemann and colleagues pooled their professional experiences and contacts in the field in addition to conducting a systematic search for self-styled “guidelines for guidelines” and other guideline development handbooks, manuals, and protocols. The reviewers, in duplicate, extracted the key stages and processes of guideline development from each of these documents, compiling them together.

The result was the G-I-N/McMaster Guideline Development Checklist: an 18-topic, 146-item soup-to-nuts comprehensive manual spanning each part and process of a guideline development program, from budgeting and planning for a program to the development of actual guidelines to their dissemination, implementation, evaluation, and updating.
An overview of the steps and parties involved in the G-I-N/McMaster guideline development checklist. Click to enlarge.

















The checklist also provides hyperlinks to tried-and-true online resources for many of these aspects, such as tips for funding a guideline program, tools for project management, topic selection criteria, and guides for patient and caregiver representatives.

Schünemann HJ, Wiercioch W, Etxeandia I, Falavigna M, Santesso N, Mustafa R, Ventresca M et al. Guidelines 2.0: Systematic development of a comprehensive checklist for a successful guideline enterprise. CMAJ 186(3): E123-E142.

Manuscript available for free here.

Tuesday, June 2, 2020

Research Shorts: Use of GRADE for the Assessment of Evidence about Prognostic Factors

In addition to questions of interventions and diagnostic tests, GRADE can also be used to assess the certainty of evidence when it comes to prognostic factors. In part 28 of the Journal of Clinical Epidemiology’s GRADE series published earlier this year, Foroutan and colleagues provide guidance for applying GRADE to a body of evidence of prognostic factors.

The Purpose of Prognostic Studies

GRADE may be applied to a body of evidence, separated by individual prognostic factors instead of outcomes, for one of two reasons. The first is a non-contextualized setting, such as when the certainty of evidence surrounding prognostic factors is being evaluated for application within research planning and analysis (e.g., determining which factors are best to use when stratifying for randomization). The second is a contextualized setting, when the certainty of evidence surrounding prognostic factors is used to help inform clinical decisions.

Establishing the Certainty of Evidence

Unlike when grading the certainty of evidence of an intervention, when assessing prognostic evidence, the overall certainty for observational studies starts out as HIGH. This is because the patient population is likely to be more representative studies than in RCTs, when eligibility criteria may place artificial restrictions on the characteristics of patients. Certainty may then be rated down based on the five traditional domains:
  • Risk of bias tools and instruments such as QUality In Prognosis Studies (QUIPS) and Prediction model Risk Of Bias ASsessment Tool (PROBAST) may be helpful here. When teasing out the effect of each potential factor, consider utilizing some form of multivariate analysis that accounts for dependence between several different prognostic factors.
  • Inconsistency can be examined via visual tests of the variability between individual point estimates and the overlap of confidence intervals; statistical tests such as i2 are likely to be less helpful, as they can often be inflated when large studies lead to particularly narrow Cis. As always, potential explanations for any observed heterogeneity should be considered a priori.
  • Imprecision will depend on whether the setting is contextualized, in which case it will depend on the relationship between the confidence interval and the previously set clinical decision threshold, or non-contextualized, in which case the threshold will most likely represent the line of no effect.
  • Indirectness should be based on a comparison of the PICOs for the clinical question at hand, and those addressed in the meta-analyzed studies.
  • Publication bias can be assessed via visually exploring a funnel plot or the use of appropriately applied statistical tests.
Foroutan F, Guyatt G, Zuk V, Vandvik PO, Alba AC, Mustafa R, Vernooij R et al. GRADE guidelines 28: Use of GRADE for the assessment of evidence about prognostic factors: Rating certainty in identification of groups of patients with different absolute risks. J Clin Epidemiol 121; 62-70.

Manuscript available from the publisher's website here.