Thursday, March 21, 2013


Seeing RED After Making Decisions Based on Observational Analysis:
What Randomization Teaches Us About
the Difference between Not Having Anemia and Treating Anemia

It has been known for many years that among patients with serious illness, those who have anemia have poorer outcomes and functioning than those who don’t. But does that association between higher hemoglobin (Hb) levels and better outcomes mean that treating patients with erythropoiesis-stimulating agents (ESAs) to increase their Hb levels will improve their health? The medical community thought so, for many years. All that has changed, though—thanks to a shift in the quality of research methodology.
Perhaps the most widely publicized change was to kidney disease treatment, in part because a key study, TREAT (Trial to Reduce Cardiovascular Events with Aranesp Therapy), was released shortly before Halloween in 2009, prompting a flood of predictable “Trick or TREAT” headline puns.[1] (Truth be told, Fred Curtiss and I were responsible for one of them in a 2009 JMCP editorial.)[2]

The remarkable sequence of events that “turned the world of anemia management upside down”[3] in kidney disease management was described in a 2008 American Journal of Kidney Diseases commentary by Marc Pfeffer, Principal Investigator of TREAT.[4] Typically, Pfeffer pointed out, drug therapy advances take place in a predictable sequence. First, placebo-controlled trials test the drug in the highest-risk, most severely ill subset of patients with a disorder. For example, the first tests of antihypertensives were made in patients with diastolic (I said diastolic, not systolic) blood pressures averaging 115 to 129 millimeters mercury (mm Hg).[5] Then, additional placebo-controlled trials of the drug are conducted in less severely ill patients. As Pfeffer observed, this sequence was followed in studies of antihypertensive combination therapies, ACE (angiotensin-converting enzyme) inhibitors, and statins.
For ESAs, the path was quite different. Following initial testing of ESAs in a sample of severely anemic dialyzed patients with no control group in 1989,[6] strong associations between higher Hb (or hematocrit) levels and positive patient outcomes in numerous observational studies[7] led to a widespread presumption that treatment with ESAs caused improvement in patient health and quality of life—with no placebo-controlled trials to determine whether the presumption was actually correct. (The placebo-controlled TREAT, initiated in 2004, was the first exception.)

In a discussion that today should serve as a reminder of the potentially serious hazards of standardizing clinical practice without a rigorous evidence base, one of the observational study reports even attributed better outcomes in dialyzed patients to ESA use in accordance with guidelines that had been promulgated in 1999:
The country with the highest median haemoglobin value, Spain …  had relatively high values in the three major categories of anaemia management practice: [recombinant human erythropoietin (rHuEpo)] use (91%), mean rHuEpo dose (114 IU/kg body weight/week) and prevalence of i.v. iron use (66%) … . The survey responses from … medical directors indicated that the majority of Spanish dialysis units had a policy of initiating rHuEpo therapy at relatively high haemoglobin threshold values. The high mean haemoglobin concentration observed for Spanish haemodialysis patients would appear to be largely a consequence of a country-wide practice of initiating rHuEpo use at a relatively high haemoglobin threshold value in conjunction with high use of rHuEpo, at a moderately high dose and maintaining sufficiently high iron levels for haemodialysis patients.  … It is apparent from the findings of the present study that a patient's haemoglobin concentration affects morbidity and mortality. … Although observational studies cannot prove causality, this study suggests that, if the [European Best Practice Guidelines (EBPG)] are followed, a trend for improved outcomes in anaemic CKD patients may be expected. The EBPG on anaemia, based on expert review of evidence, thus receives additional supportive evidence from the present study. Despite the dissemination of the EPBG, many European haemodialysis patients still have haemoglobin concentrations below the minimum recommended level of 11 g/dl. This suggests a major opportunity for improved anaemia management for haemodialysis patients.

Widespread use of ESAs expanded to patients with cancer and nondialyzed kidney disease—again, Pfeffer reported, without supporting evidence from placebo-controlled trials. Remarkably, the benefits of ESAs for patients with anemia in kidney disease were so widely accepted that the ethics of having a placebo group were questioned in the planning stages of TREAT.[4]

The tide began to turn in November 2006 with the publication of the CHOIR (Correction of Hemoglobin and Outcomes In Renal insufficiency) and CREATE (Cardiovascular Risk Reduction by Early Anemia Treatment with Epoetin beta) trials, which randomized nondialyzed patients with kidney disease to higher versus lower Hb targets using ESA treatment (but still no placebo group!).[8] CREATE found that complete anemia correction did not reduce the rate of cardiovascular events but did improve quality-of-life measures of physical functioning and general health. CHOIR found that use of higher Hb targets increased the risk of cardiovascular events without improving quality of life.[8] Shortly thereafter, a series of eight studies provided to the FDA found “more rapid tumor growth or shortened survival when patients with breast, non-small cell lung, head and neck, lymphoid or cervical cancers received ESAs compared to patients who did not receive this treatment,” resulting in “black box warning” additions to the product labels.[9] TREAT, which was conducted in a sample of patients with chronic kidney disease and type 2 diabetes, found no statistical difference between ESA and placebo in the primary outcomes (death or a nonfatal cardiac event; death or end-stage renal disease), but ESA-treated patients were twice as likely to have a stroke.[10] Finally and most recently, the RED-HF (Reduction of Events by Darbopoetin Alfa in Heart Failure) trial, reported in The New England Journal of Medicine in March 2013, found that ESA treatment did not improve the primary study outcome (death or hospitalization from heart failure) or any secondary outcomes in a sample of patients with heart failure and mild-to-moderate anemia, but increased the risk of thromboembolic events.[11]
Practical Takeaway Point: Quality, Not Just Quantity, of Evidence

What can we learn from the nearly 20-year history of treatment with ESAs that, as Pfeffer described it, "greatly outpaced the data" prior to CHOIR, CREATE, TREAT, and now RED-HF? The principal moral of the sequence of events in ESA treatment and research—and so many others like it in health care—is this: much as we wish that our observations told the whole story, they often don’t. Associations are often elucidating but sometimes misleading, and interventions based on them should be subjected to rigorous testing, preferably using experimental designs, before being put into practice.
So, if you read or hear a claim that a proposal or treatment is “evidence-based,” try to find out what type of evidence supports the claim. Most online press articles contain links to the original study report including the abstract, which provides a quick overview of the design and analysis. If the study investigators did not use a control group, or if the control group was not randomized using “allocation concealment” (meaning that the treatment assignment is “concealed” or hidden from investigators or staff who can influence the study results), the quality of evidence should be viewed as relatively weak. It takes a little extra time to make a determination of quality of evidence—but it is time well spent.

[1] Husten L. Halloween trick: don’t TREAT diabetes with ESAs. Cardiobrief. October 30, 2009.
[2] Curtiss FR, Fairman KA. No TREATment with darbepoetin dosed to hemoglobin 13 grams per deciliter in type 2 diabetes with pre-dialysis chronic kidney disease—safety warnings for erythropoiesis-stimulating agents. J Manag Care Pharm. 2009;15(9):759-765.

[3] Singh AK. Does TREAT give the boot to ESAs in the treatment of CKD anemia? J Am Soc Nephrol. 2010;28.
[4] Pfeffer MA. Critical missing data on erythropoiesis-stimulating agents in CKD: first beat placebo. Am J Kidney Dis. 2008;51(3):366-369.

[5] Veterans Administration Cooperative Study Group on Antihypertensive Agents. Effects of treatment on morbidity in hypertension (Results in patients with diastolic blood pressures averaging 115 through 129 mm Hg). JAMA. 1967;202:1028-1034.
[6] Eschbach JW, Abdulhadi MH, Browne JK, et al. Recombinant human erythropoietin in anemic patients with end-stage renal disease. Results of a phase III multicenter clinical trial. Ann Intern Med. 1989;111:992-1000.

[7] Foley RN, Parfrey PS, Harnett JD, Kent GM, Murray DC, Barre PE. The impact of anemia on cardiomyopathy, morbidity, and mortality in end-stage renal disease. Am J Kidney Dis. 1996;28(1):53-61; Tong PC, Kong AP, So WY, et al. Hematocrit, independent of chronic kidney disease, predicts adverse cardiovascular outcomes in Chinese patients with type 2 diabetes. Diabetes Care. 2006;29(11):2439-2444; Locatelli F, Pisoni RL, Combe C, et al. Anaemia in haemodialysis patients of five European countries: association with morbidity and mortality in the Dialysis Outcomes and Practice Patterns Study (DOPPS). Nephrol Dial Transplan. 2004;19(1):121-132; Thorp M, Johnson ES, Yang X, Petrik AF, Platt R, Smith DH. Effect of anaemia on mortality, cardiovascular hospitalizations and end-stage renal disease among patients with chronic kidney disease. Nephrology (Carlton). 2009;14(2):240-246.
[8] Drueke TB, Locatelli F, Clyne N, et al.; CREATE Investigators. Normalization of hemoglobin level in patients with chronic kidney disease and anemia. N Engl J Med. 2006;355(20):2071-2084; Singh AK, Szczech L, Tang KL, et al.; CHOIR Investigators. Correction of anemia with epoetin alfa in chronic kidney disease. N Engl J Med. 2006;355(20:2085-2098.

[10] Pfeffer MA, Burdmann EA, Chen CY, et al; TREAT investigators. A trial of darbepoetin alfa in type 2 diabetes and chronic kidney disease. N Engl J Med. 2009;361:2019-2032.

[11] Swedberg K, Young JB, Anand IS, et al.; RED-HF Committees and Investigators. Treatment of anemia with darbepoetin alfa in systolic heart failure. N Engl J Med. 2013 [EPub ahead of print]

Wednesday, March 13, 2013


The Catcher’s MITT and Questions about Research on Off-Label Drug Use:
Inadequate CONSORT Guidelines or Weak Peer Review?
 

My oldest son, a former high school catcher and avid sabermetrician, once taught me that the subtle side-to-side body movements I observed just before he caught some pitches were both common and purposeful. When a pitch deviates a bit from its intended target, he explained, it is standard practice for a catcher to move his mitt slightly to make it appear to the umpire that the ball landed exactly as expected—a technique known as “framing.” Umpires tend to overlook framing in calling “strikes,” as long as the deviation from the target is not too large.
A recent thought-provoking analysis by Vedula and colleagues[1] may provide us with an example of framing in published health care research—but, in an interesting twist, leaves us wondering if the “umpires” (peer reviewers and editors) were out in left field or watching the game closely from home plate as they were supposed to, perhaps even doing a little helpful coaching to nudge the pitches in the right direction. Vedula et al. compared internal Pfizer/Parke-Davis study reports with final published journal articles regarding four off-label uses of gabapentin—migraine prophylaxis and treatment of neuropathic pain, nociceptive pain, and bipolar disorder.  Relying on documents obtained through litigation for which one of the authors was an expert witness, the analysis found that the counts of “participants randomized and analyzed for efficacy” differed from internal report to publication in three of ten trials analyzed. More remarkably, Vedula et al. noted the use of six different definitions of “intention-to-treat” (ITT) analysis—a technique known as “modified intention-to-treat” (MITT).
Notably, none of the definitions of ITT was consistent with standard ITT, in which all cases are analyzed in the group to which they were originally randomized. And seven different types of analyses for efficacy were used, including not only ITT and MITT but also some nonstandard and creatively named techniques, such as “efficacy evaluable.”
The study by Vedula et al. contributed to a growing body of evidence about the way that studies are translated (and sometimes mis-translated) from protocol to execution to publicly reported information.[2] Analyses such as that done by Vedula et al. involve painstaking document extraction and verification, and those who conduct them should be commended for the level of effort that they require. They also provide important fodder for serious dialogue about research ethics and the publication process. However, this particular study report took a couple of unusual and, I think, mistaken twists, both of which erroneously minimized the role of journal editorial and peer review in improving the quality of published research.
First, the authors interpreted the discrepancies between the internal reports and the final publications as evidence that publications were not transparent or accurate, “presuming that the [internal] research report truly describes the facts.” This presumption is the problem, because the process of editorial and peer review should and often does result in corrections of errors in originally submitted manuscripts. Specifically, peer reviewers and editors commonly notice discrepancies (e.g., the methods section says that Group A was removed, but the tables show patients in Group A); use of inappropriate cohort definitions or statistical techniques; or even mistakes in mathematical calculations. In a less focused, perhaps less expert, review done by internal decision makers within a company, these details (often minor—for example, one of the three sample size discrepancies noted by Vidula et al. involved only a single study case) are much less likely to be noticed and corrected. Thus, authors who submit their work for peer review should be commended for participation in a process that results in the publication of more accurate findings, not criticized for making necessary corrections to a manuscript when peer review does its job.
Second and more importantly, the authors concluded based on their findings that the primary reporting standard for randomized trials, CONSORT (CONsolidated Standards Of Reporting Trials) should be enhanced to, among other changes, standardize “the definitions of various types of analyses.” This recommendation seems to reflect a misunderstanding of the purpose of CONSORT and similar guidelines:
The objective of CONSORT is to provide guidance to authors about how to improve the reporting of their trials. Trial reports need be clear, complete, and transparent. Readers, peer reviewers, and editors can also use CONSORT to help them critically appraise and interpret reports of RCTs. However, CONSORT was not meant to be used as a quality assessment instrument.[CONSORT explanation and elaboration, 3]
In other words, to the extent that they succeed in encouraging accurate reporting, CONSORT and similar guidelines enable peer reviewers and editors to do their jobs, but they do not—and were never intended to—do those jobs for them. To understand this point, consider the definition of MITT provided in one published report of the efficacy of gabapentin in migraine prophylaxis, by Mathew et al.:
This population included any patient who was randomized, took at least one dose of study medication during SP [stabilization period] 2, maintained a stable dose of 2400 mg/day during SP2, had baseline migraine headache data, and at least 1 day of migraine headache evaluations during SP2.[4]
To call this analytic strategy “MITT” is a catcher-style framing stretch that the “umpires” probably should have “called.” (UMITT—ΓΌber-modified intention-to-treat—might be more accurate.) The study’s report was admirably clear and transparent but provides plenty of reason to be concerned about the accuracy of its conclusion that “gabapentin is an effective prophylactic agent for patients with migraine.” According to the report’s sample selection flowchart, only 57% of 98 gabapentin-randomized patients, compared with 69% of 45 placebo patients, met the MITT criteria. Additionally, the proportions of originally assigned patients who discontinued treatment for adverse events were 16% for gabapentin and 9% for placebo. Despite these discrepancies between the selection processes for the study groups, the report included no true ITT analysis, and apparently the journal, Headache, did not require one.
So, would enhancements to CONSORT result in a reduced use of MITT (or UMITT)? CONSORT recommendations already include admonitions that using the term “ITT” is inappropriate when cases have been removed from the analysis, and that MITT “may be appropriate in some settings” but “should be properly labeled as a non-randomised, observational comparison.” So, if the report by Mathew et al. “hit” when it should have “struck out” (we don’t know for certain without an ITT analysis), it is hard to see inadequate CONSORT guidance as the cause.
Practical Takeaway Points—How to Keep Your Eye on the Ball in the MITT
First, Mathew et al.’s report on migraine prophylaxis is a great illustration of the importance of a sample selection flowchart as recommended by CONSORT and other EQUATOR (Enhancing the QUality And Transparency Of health Research) guidelines. Without it, I would not have been able to determine that use of MITT instead of ITT was potentially problematic. So, if a report does not have either a sample selection flowchart or a very clear and quantitative description of the effect of each sampling criterion, it would be reasonable to view it skeptically.
Second, if MITT instead of ITT is used, it is wise to check the sample selection flowchart for the proportion of initially randomized patients who were included in the MITT. If the proportion is low, or if the proportions differ by study groups, skepticism about the study findings is appropriate.
Third, CONSORT guidelines recognize that modifications to initial study protocols are sometimes necessary but encourage authors to explain them. When MITT instead of ITT is used, a report that explains the decision should be viewed with more trust than a report that fails to do so.
Finally, although Vedula et al. concluded from their work that “current US legal requirements for reporting study findings are inadequate both in scope and detail,” the patterns they observed may reflect a more satisfactory picture. After all, to the extent that the peer review process prompts authors to make appropriate changes, and reporting guidelines encourage authors to clarify potential methodological problems (such as those observed by Vedula et al.), these tools have achieved their objectives.

 [2] Fairman KA, Curtiss FR. What should be done about bias and misconduct in clinical trials? J Manag Care Pharm. 2009 Mar;15(2):154-60
[4] Mathew NT, Rapoport A, Saper J, et al. Efficacy of gabapentin in migraine prophylaxis. Headache. 2001;41:119-128.

 

Wednesday, March 6, 2013

Are We All Happy about the End of Flu Season?
A Surprising Source Says “No”


The Centers for Disease Control
announced recently[1] that the 2012-2013 influenza outbreak,[2] which has to date killed 81 children and prompted the declaration of public health emergencies in multiple U.S. cities, is finally waning. Since the proportion of deaths attributable to pneumonia and influenza has been at epidemic levels since January, the announcement of the end of one of the most dangerous flu seasons in years comes as welcome news to all of us.

Well … maybe to most of us. Because if an unintentionally shocking
January 2013 Drug Store News article is to be believed (hint: I don’t think that it is), your local pharmacist may actually be upset by the decline in flu activity. Why? Because “the sickest winter we will have had in years” is “big business for retail pharmacy.”[3]

Yes, for those who were searching for a reason to be pleased with an influenza season that has been characterized by only “moderate” vaccine effectiveness and markedly heightened risk for senior citizens, in whom an estimated 90% of flu-related deaths have occurred—you now have your answer. The severity of this year’s flu season is a good thing, says the Drug Store News article, because it has generated a lot of revenue for pharmacies, making last year’s mild flu season a “distant, bad memory.” After all, consider the positive aspects of the 2013 outbreak as highlighted in the article:

For the four weeks ended Dec. 30 (and before this year's flu incidence peak), sales of hand sanitizers were up 15.4% to $14.4 million; sales of personal thermometers were up 35.8% to $17.9 million; sales of cold and allergy liquid formulations were up 27.4% to $133 million; and sales of cold and allergy tablets were up 7.9% to $352 million. … And the Food and Drug Administration took measures to ensure adequate supply of Tamiflu for both adults and children. … And there are still two months of flu season to go.

Feeling encouraged by all this good news?

Aside from the obvious point that the article was tasteless and should have been edited to remove what I can only assume was an inadvertently callous and commercial tone not representative of the true values of the author or Drug Store News, there is a critically important lesson for research methods here: we serve people.

It is an easy lesson to forget. Many of us in
health care research work only with computer files; anonymized, HIPAA-compliant identifiers; and alpha-numeric diagnosis and procedure codes that become clinically meaningful only with the use of esoteric and complex coding systems. It’s not that easy to see all those numbers as human beings. That is one of several reasons that I have chosen to continue working in nursing homes for at least a few hours each month, as a sort of counterweight to analyzing and writing about anonymous study “cases:” it helps me to remember that they are people. Without that regular contact with patients, I might forget.

So I repeat here an observation that I intentionally call to mind on a regular basis. When I analyze a dataset documenting the medical and pharmacy services provided to cases, I am receiving the informational benefits of a glimpse into events—sometimes life-changing, sometimes tragic, sometimes merely inconvenient—affecting people. They are not revenues, median survival times, units of service, dispensed days supply, QALYs, or even “covered lives.” They are people. They are our moms and dads, our kids, our friends, our coworkers, and sometimes—as I have sadly observed in my work with nursing home patients—caregivers to spouses who struggle to go on after their primary source of medical and social support becomes incapacitated or dies. A flu epidemic is no small thing.

To receive and benefit from information about people is both a privilege and an immense responsibility. We fulfill that responsibility when we choose excellence in research ethics and methods; select research topics that meet a true therapeutic or informational need rather than a commercial purpose alone; and report findings transparently, thoroughly, and accurately. As I wrote in
Health Care Research Done Right,[4] we all fail, to some extent, in that endeavor. But we are all ethically obligated to try to succeed—and, when necessary, to call out those who find it acceptable to view people and their health care needs as revenue generators while ignoring the humanistic costs of the medical conditions that our health care system is intended to prevent and treat.


[1] Weise E.
Flu no longer widespread in the U.S. USAToday. March 1, 2013.
[2] Centers for Disease Control.
What you should know for the 2012-2013 influenza season. February 21, 2013.
[3] Johnsen M.
Monster flu season makes last year’s non existent [sic] flu season a distant, bad memory. Drug Store News.
[4] Fairman KA.
Health Care Research Done Right: A Journal Editor Shares Practical Tips and Techniques for High Quality and Efficiency. Outskirts Press; 2012.