The Catcher’s MITT and Questions about
Research on Off-Label Drug Use:
Inadequate CONSORT Guidelines or Weak Peer Review?
Inadequate CONSORT Guidelines or Weak Peer Review?
My oldest son, a former high school catcher and avid
sabermetrician, once taught me that the subtle side-to-side
body movements I observed just before he caught some pitches were both common
and purposeful. When a pitch deviates a bit from its intended target, he
explained, it is standard practice for a catcher to move his mitt slightly to
make it appear to the umpire that the ball landed exactly as expected—a
technique known as “framing.” Umpires tend to overlook framing in calling
“strikes,” as long as the deviation from the target is not too large.
A recent thought-provoking analysis
by Vedula and colleagues[1] may provide us with an example of
framing in published health care research—but, in an interesting twist, leaves
us wondering if the “umpires” (peer reviewers and editors) were out in left
field or watching the game closely from home plate as they were supposed to, perhaps
even doing a little helpful coaching to nudge the pitches in the right
direction. Vedula et al. compared internal Pfizer/Parke-Davis study reports
with final published journal articles regarding four off-label uses of gabapentin—migraine
prophylaxis and treatment of neuropathic pain, nociceptive pain, and bipolar
disorder. Relying on documents obtained
through litigation for which one of the authors was an expert witness, the
analysis found that the counts of “participants randomized and analyzed for
efficacy” differed from internal report to publication in three of ten trials
analyzed. More remarkably, Vedula et al. noted the use of six different
definitions of “intention-to-treat” (ITT) analysis—a technique known as
“modified intention-to-treat” (MITT).
Notably, none of the definitions of ITT was consistent
with standard ITT, in which all cases
are analyzed in the group to which they were originally randomized. And seven different types of analyses for
efficacy were used, including not only ITT and MITT but also some nonstandard
and creatively named techniques, such as “efficacy evaluable.”
The study by Vedula et al. contributed to a growing body of evidence about the way that studies
are translated (and sometimes mis-translated) from protocol to execution to
publicly reported information.[2] Analyses such as that done by Vedula et al.
involve painstaking document extraction and verification, and those who conduct
them should be commended for the level of effort that they require. They also
provide important fodder for serious dialogue about research ethics and the
publication process. However, this particular study report took a couple of
unusual and, I think, mistaken twists, both of which erroneously minimized the
role of journal editorial and peer review in improving the quality of published
research.
First, the authors interpreted the discrepancies between
the internal reports and the final publications as evidence that publications
were not transparent or accurate, “presuming that the [internal] research
report truly describes the facts.” This presumption is the problem, because the
process of editorial and peer review should and often does result in
corrections of errors in originally submitted manuscripts. Specifically, peer
reviewers and editors commonly notice discrepancies (e.g., the methods section
says that Group A was removed, but the tables show patients in Group A); use of
inappropriate cohort definitions or statistical techniques; or even mistakes in
mathematical calculations. In a less focused, perhaps less expert, review done
by internal decision makers within a company, these details (often minor—for
example, one of the three sample size discrepancies noted by Vidula et al.
involved only a single study case) are much less likely to be noticed and
corrected. Thus, authors who submit their work for peer review should be
commended for participation in a process that results in the publication of more
accurate findings, not criticized for making necessary corrections to a
manuscript when peer review does its job.
Second and more importantly, the authors concluded based
on their findings that the primary reporting standard for randomized trials,
CONSORT (CONsolidated Standards Of Reporting Trials) should be enhanced to,
among other changes, standardize “the definitions of various types of analyses.”
This recommendation seems to reflect a misunderstanding of the purpose of
CONSORT and similar guidelines:
The
objective of CONSORT is to provide guidance to authors about how to improve the
reporting of their trials. Trial reports need be clear, complete, and
transparent. Readers, peer reviewers, and editors can also use CONSORT to help
them critically appraise and interpret reports of RCTs. However, CONSORT was
not meant to be used as a quality assessment instrument.[CONSORT
explanation and elaboration, 3]
In other words, to the extent that they succeed in encouraging
accurate reporting, CONSORT and similar guidelines enable peer reviewers and editors to do their jobs, but they do
not—and were never intended to—do those jobs for them. To understand this
point, consider the definition of MITT provided in one published report of the
efficacy of gabapentin in migraine prophylaxis, by Mathew et al.:
This
population included any patient who was randomized, took at least one dose of
study medication during SP [stabilization period] 2, maintained a stable dose
of 2400 mg/day during SP2, had baseline migraine headache data, and at least 1
day of migraine headache evaluations during SP2.[4]
To call this analytic strategy “MITT” is a catcher-style framing
stretch that the “umpires” probably should have “called.” (UMITT—über-modified
intention-to-treat—might be more accurate.) The study’s report was admirably clear
and transparent but provides plenty of reason to be concerned about the
accuracy of its conclusion that “gabapentin is an effective prophylactic agent
for patients with migraine.” According to the report’s sample selection
flowchart, only 57% of 98 gabapentin-randomized patients, compared with 69% of 45
placebo patients, met the MITT criteria. Additionally, the proportions of
originally assigned patients who discontinued treatment for adverse events were
16% for gabapentin and 9% for placebo. Despite these discrepancies between the selection
processes for the study groups, the report included no true ITT analysis, and
apparently the journal, Headache, did
not require one.
So, would enhancements to CONSORT result in a reduced use
of MITT (or UMITT)? CONSORT recommendations already include admonitions that
using the term “ITT” is inappropriate when cases have been removed from the
analysis, and that MITT “may be appropriate in some settings” but “should be
properly labeled as a non-randomised, observational comparison.” So, if the
report by Mathew et al. “hit” when it should have “struck out” (we don’t know for
certain without an ITT analysis), it is hard to see inadequate CONSORT guidance
as the cause.
Practical
Takeaway Points—How to Keep Your Eye on the Ball in the MITT
First, Mathew
et al.’s report on migraine prophylaxis is a great illustration of the
importance of a sample selection flowchart as recommended by CONSORT and other EQUATOR
(Enhancing the QUality And Transparency Of health Research) guidelines. Without
it, I would not have been able to determine that use of MITT instead of ITT was
potentially problematic. So, if a report does not have either a sample
selection flowchart or a very clear
and quantitative description of the
effect of each sampling criterion, it would be reasonable to view it
skeptically.
Second, if
MITT instead of ITT is used, it is wise to check the sample selection flowchart
for the proportion of initially randomized patients who were included in the
MITT. If the proportion is low, or if the proportions differ by study groups,
skepticism about the study findings is appropriate.
Third, CONSORT
guidelines recognize that modifications to initial study protocols are
sometimes necessary but encourage authors to explain them. When MITT instead
of ITT is used, a report that explains the decision should be viewed with more
trust than a report that fails to do so.
Finally,
although Vedula et al. concluded from their work that “current US legal
requirements for reporting study findings are inadequate both in scope and
detail,” the patterns they observed may reflect a more satisfactory picture.
After all, to the extent that the peer review process prompts authors to make
appropriate changes, and reporting guidelines encourage authors to clarify potential methodological problems (such as those observed by Vedula et al.), these tools have achieved their objectives.
[1] Vedula SS, Li T, Dickersin K. Differences
in reporting of analyses in internal company documents versus published trial
reports: comparisons in industry-sponsored trials in off-label uses of gabapentin. PLOS Medicine. 2013;10(1).
[2] Fairman KA,
Curtiss FR. What should be
done about bias and misconduct in clinical trials? J Manag Care Pharm. 2009
Mar;15(2):154-60
[3] Moher D, Hopewell S, Schulz KF, et al. CONSORT 2010 Explanation and
Elaboration: updated guidelines for reporting parallel group randomized trials. BMJ. March 24, 2010.
[4] Mathew NT, Rapoport A, Saper J, et al. Efficacy of
gabapentin in migraine prophylaxis. Headache.
2001;41:119-128.
No comments:
Post a Comment