Wednesday, March 13, 2013


The Catcher’s MITT and Questions about Research on Off-Label Drug Use:
Inadequate CONSORT Guidelines or Weak Peer Review?
 

My oldest son, a former high school catcher and avid sabermetrician, once taught me that the subtle side-to-side body movements I observed just before he caught some pitches were both common and purposeful. When a pitch deviates a bit from its intended target, he explained, it is standard practice for a catcher to move his mitt slightly to make it appear to the umpire that the ball landed exactly as expected—a technique known as “framing.” Umpires tend to overlook framing in calling “strikes,” as long as the deviation from the target is not too large.
A recent thought-provoking analysis by Vedula and colleagues[1] may provide us with an example of framing in published health care research—but, in an interesting twist, leaves us wondering if the “umpires” (peer reviewers and editors) were out in left field or watching the game closely from home plate as they were supposed to, perhaps even doing a little helpful coaching to nudge the pitches in the right direction. Vedula et al. compared internal Pfizer/Parke-Davis study reports with final published journal articles regarding four off-label uses of gabapentin—migraine prophylaxis and treatment of neuropathic pain, nociceptive pain, and bipolar disorder.  Relying on documents obtained through litigation for which one of the authors was an expert witness, the analysis found that the counts of “participants randomized and analyzed for efficacy” differed from internal report to publication in three of ten trials analyzed. More remarkably, Vedula et al. noted the use of six different definitions of “intention-to-treat” (ITT) analysis—a technique known as “modified intention-to-treat” (MITT).
Notably, none of the definitions of ITT was consistent with standard ITT, in which all cases are analyzed in the group to which they were originally randomized. And seven different types of analyses for efficacy were used, including not only ITT and MITT but also some nonstandard and creatively named techniques, such as “efficacy evaluable.”
The study by Vedula et al. contributed to a growing body of evidence about the way that studies are translated (and sometimes mis-translated) from protocol to execution to publicly reported information.[2] Analyses such as that done by Vedula et al. involve painstaking document extraction and verification, and those who conduct them should be commended for the level of effort that they require. They also provide important fodder for serious dialogue about research ethics and the publication process. However, this particular study report took a couple of unusual and, I think, mistaken twists, both of which erroneously minimized the role of journal editorial and peer review in improving the quality of published research.
First, the authors interpreted the discrepancies between the internal reports and the final publications as evidence that publications were not transparent or accurate, “presuming that the [internal] research report truly describes the facts.” This presumption is the problem, because the process of editorial and peer review should and often does result in corrections of errors in originally submitted manuscripts. Specifically, peer reviewers and editors commonly notice discrepancies (e.g., the methods section says that Group A was removed, but the tables show patients in Group A); use of inappropriate cohort definitions or statistical techniques; or even mistakes in mathematical calculations. In a less focused, perhaps less expert, review done by internal decision makers within a company, these details (often minor—for example, one of the three sample size discrepancies noted by Vidula et al. involved only a single study case) are much less likely to be noticed and corrected. Thus, authors who submit their work for peer review should be commended for participation in a process that results in the publication of more accurate findings, not criticized for making necessary corrections to a manuscript when peer review does its job.
Second and more importantly, the authors concluded based on their findings that the primary reporting standard for randomized trials, CONSORT (CONsolidated Standards Of Reporting Trials) should be enhanced to, among other changes, standardize “the definitions of various types of analyses.” This recommendation seems to reflect a misunderstanding of the purpose of CONSORT and similar guidelines:
The objective of CONSORT is to provide guidance to authors about how to improve the reporting of their trials. Trial reports need be clear, complete, and transparent. Readers, peer reviewers, and editors can also use CONSORT to help them critically appraise and interpret reports of RCTs. However, CONSORT was not meant to be used as a quality assessment instrument.[CONSORT explanation and elaboration, 3]
In other words, to the extent that they succeed in encouraging accurate reporting, CONSORT and similar guidelines enable peer reviewers and editors to do their jobs, but they do not—and were never intended to—do those jobs for them. To understand this point, consider the definition of MITT provided in one published report of the efficacy of gabapentin in migraine prophylaxis, by Mathew et al.:
This population included any patient who was randomized, took at least one dose of study medication during SP [stabilization period] 2, maintained a stable dose of 2400 mg/day during SP2, had baseline migraine headache data, and at least 1 day of migraine headache evaluations during SP2.[4]
To call this analytic strategy “MITT” is a catcher-style framing stretch that the “umpires” probably should have “called.” (UMITT—über-modified intention-to-treat—might be more accurate.) The study’s report was admirably clear and transparent but provides plenty of reason to be concerned about the accuracy of its conclusion that “gabapentin is an effective prophylactic agent for patients with migraine.” According to the report’s sample selection flowchart, only 57% of 98 gabapentin-randomized patients, compared with 69% of 45 placebo patients, met the MITT criteria. Additionally, the proportions of originally assigned patients who discontinued treatment for adverse events were 16% for gabapentin and 9% for placebo. Despite these discrepancies between the selection processes for the study groups, the report included no true ITT analysis, and apparently the journal, Headache, did not require one.
So, would enhancements to CONSORT result in a reduced use of MITT (or UMITT)? CONSORT recommendations already include admonitions that using the term “ITT” is inappropriate when cases have been removed from the analysis, and that MITT “may be appropriate in some settings” but “should be properly labeled as a non-randomised, observational comparison.” So, if the report by Mathew et al. “hit” when it should have “struck out” (we don’t know for certain without an ITT analysis), it is hard to see inadequate CONSORT guidance as the cause.
Practical Takeaway Points—How to Keep Your Eye on the Ball in the MITT
First, Mathew et al.’s report on migraine prophylaxis is a great illustration of the importance of a sample selection flowchart as recommended by CONSORT and other EQUATOR (Enhancing the QUality And Transparency Of health Research) guidelines. Without it, I would not have been able to determine that use of MITT instead of ITT was potentially problematic. So, if a report does not have either a sample selection flowchart or a very clear and quantitative description of the effect of each sampling criterion, it would be reasonable to view it skeptically.
Second, if MITT instead of ITT is used, it is wise to check the sample selection flowchart for the proportion of initially randomized patients who were included in the MITT. If the proportion is low, or if the proportions differ by study groups, skepticism about the study findings is appropriate.
Third, CONSORT guidelines recognize that modifications to initial study protocols are sometimes necessary but encourage authors to explain them. When MITT instead of ITT is used, a report that explains the decision should be viewed with more trust than a report that fails to do so.
Finally, although Vedula et al. concluded from their work that “current US legal requirements for reporting study findings are inadequate both in scope and detail,” the patterns they observed may reflect a more satisfactory picture. After all, to the extent that the peer review process prompts authors to make appropriate changes, and reporting guidelines encourage authors to clarify potential methodological problems (such as those observed by Vedula et al.), these tools have achieved their objectives.

 [2] Fairman KA, Curtiss FR. What should be done about bias and misconduct in clinical trials? J Manag Care Pharm. 2009 Mar;15(2):154-60
[4] Mathew NT, Rapoport A, Saper J, et al. Efficacy of gabapentin in migraine prophylaxis. Headache. 2001;41:119-128.

 

No comments:

Post a Comment