Saturday, August 10, 2013

BD, ICD, and GIGO:
Why “Big Data” May Be Less than Meets the Eye

With both funding and pressure piling onto the business of comparative effectiveness research, the old mantra that you cannot improve what you do not measure has seemingly never been more critically important in health care than it is today. Many would argue that measuring health care outcomes has also never been easier.

And, in a sense, they’d be right. Health care researchers today have sophisticated equipment with supercomputing capabilities that many of us couldn’t even dream of ten or fifteen years ago. We also have access to tens of millions of data points—medical and pharmacy claims, mortality records, survey data, and even detailed genetic data, depending on the organization in which we are working. So much information, available so conveniently, at such a low cost, and analyzable at such speed. It’s enough to make a policy analyst positively giddy. Perhaps it is not surprising then, that one commentator, a Health Policy and Life Sciences Group manager at Intel, assessed Big Data’s potential to “revolutionize health care” in this way:

Big Data provides us an opportunity to transition to a personal care system. Rather than making assumptions based on what has worked for other people, this personal view would allow us to take data about a patient’s genomes, medical history and behaviors to construct a virtual model that would help predict which treatments will be most effective and customize them to an individual — improving quality of life for the patient and saving the delivery system money.
With all this excitement about the potential of large datasets to unlock the secrets of greater longevity at lower cost, it’s easy to forget a crucial pitfall encountered by researchers who use them unawares: those beautifully packaged data were collected by human beings. And, because many of those human beings did not have research on their mind when they collected and recorded the data, they may have been motivated to treat our valuable information in ways that we did not expect or want. One example, which I discussed in Chapter 7 of Health Care Research Done Right, is billing (claims) data. Because the main purpose of billing data is to generate payment for health care services, claims coding is subject to “upcoding” and deliberate miscoding to enhance reimbursement.

This example is known to many claims database researchers. However, not all similar threats to study validity are as widely recognized or publicized, despite their potential effect on the well-being of patients whose plans or health care providers forget that real-world data collection may affect ideal-world research in unexpected ways. Several examples have been highlighted in recent press articles. I’ll talk about two in this post.
First is updated information about the long-awaited transition to ICD-10, which will increase the number of available diagnosis codes from about 14,000 to more than 68,000, the number of procedure codes from about 4,000 to more than 72,000, and the number of pages in the American Family Practice Association “superbill” (a standard form intended to list most of the diagnosis codes encountered in a typical practice, used for the physician’s convenience) from 2 to a whopping 9 pages. If used properly, the coding system has the potential to improve the accuracy of electronic diagnostic record-keeping, reimbursement, and, ultimately, quality of care, because—in theory—physicians who are given more accurate feedback will have greater incentive to use evidence-based procedures and treatment protocols.

But I said if used properly, and that big “if” is looking iffier all the time. The deadline for compliance with ICD-10 coding, originally scheduled by the Centers for Medicare & Medicaid Services (CMS) for October 2011, has been pushed back several times and is now slated for October 2014. Except that a survey of “providers, payers and health information technology vendors” conducted in February 2013 found that about one-half of participating vendors reported less than 50% progress toward ICD-10 readiness, and more than 40% did not know when they would begin detailed steps toward final implementation, with about one-quarter reporting being nearly finished.

And those are just the data processing problems. Industry insiders are reporting that experienced diagnosis coders are choosing to retire rather than learn the new system, leaving an unknown proportion of the implementation of ICD-10 in the hands of newbies. How long it will take for coders to become familiar with the new system is unknown. Which will bring researchers to a critically important question: when analyzing claims data coded with ICD-10, how accurate are the diagnoses?

A second example involves a completely different data source but a similar problem. An anonymous, Internet-based survey of New York City hospital residents found that 49% had knowingly reported cause of death inaccurately when completing a death certificate.[1] Of residents who had completed at least eleven death certificates in the previous three years, nearly six in ten reported deliberate inaccuracy. About three-quarters said that the computer system “would not accept the correct cause” of death, 41% said that they were told to “put something else” by the hospital admitting office, and 31% said that the medical examiner told them to report the diagnosis incorrectly. Among the more common actual causes of death associated with deliberately inaccurate reporting was septic shock, management of which is a quality-of-care indicator.[2] Noting that death certificates “contain critical information for epidemiology, public health research, disease surveillance, and community health programs,” the researchers noted that the routine reporting of inaccurate causes of death “may have lasting effects on the public health priorities of the community.”

All of which should give us pause as we consider the current level of enthusiasm for “big data.” Are automated data a silver bullet for all that ails the American health care system, or are they GIGO (garbage in, garbage out)? Savvy researchers should recognize that either possibility exists and should know how to investigate the data prior to using them.

[1] Wexelman BA, Eden E, Rose KM. Survey of New York City resident physicians on cause-of-death reporting, 2010. Prev Chronic Dis. 2013;10:E76.

[2] NQF #0500 Severe Sepsis and Septic Shock: Management Bundle, Last Updated Date: Oct 05, 2012.

Tuesday, June 11, 2013


Value-Based Insurance Design Defined as Copayment Reduction:
Where Have All the Studies Gone? Part 2


In Part 1 of this posting, I explained the recent shift in the meaning of the term “value-based,” hinted at missing research evidence, and made a passing reference to an award-winning song included on one of the few folk albums in U.S. history to attain #1 status for a full month—Peter, Paul, and Mary, released in 1962 by the artists of the same name. Today, on the fifty-first anniversary of John F. Kennedy’s commencement address at Yale, also in 1962, I draw on his key theme and ask: when it comes to VBID, have we “enjoyed the comfort of opinion without the discomfort of thought”?
A characterization that no “discomfort” associated with hard work has been experienced in VBID research would be unfair because extensive analysis has been done—much (but not the majority) of high quality.[1] However, some key information has gone unreported, as shown in the table below:

Table 1. Copayment Reduction Studies
First Author (Publication Year)—Description
Current Status
Chernew (2008). Quasi-experimental (difference-in-difference times series analysis with nonequivalent comparison group)[2]
 
Copayment of $0 for generic, 50% reduction for brand
 
Outcome measure was MPR change.
Reported. Copayment reductions were associated with a 2.6 to 4.0 MPR point difference (i.e., an added 7 to 14 days of therapy per year).
 
Key information missing from publication included:
*Pre-intervention cost-sharing amounts in the comparison plan
*Comparability of relevant benefit features in the study plans
*Industry sectors of employees in study plans
Original analytic design of planned follow-up to Chernew 2008 (study described in previous table row). Outcomes included hospitalizations and ER visits, which had been predicted by the investigators to decline as a result of the medication adherence improvements observed.[2]
 
Results not reported. Instead, after initial analysis found “considerable uncertainty surrounding estimates of the impact of the [VBID] intervention on aggregate spending,” the authors developed a pharmacoeconomic model to estimate outcomes, describing it as “evidence that value-based insurance can be effective.”[3]
Spaulding (2009). MHealthy intervention for University of Michigan employees and dependents with diabetes; prospective trial with nonrandomized comparison group. Primary outcomes were medication utilization and adherence. Secondary outcomes were health care costs and utilization rates for hospitalizations, ER visits, and outpatient visits.[4]
Results not reported, although the study was initiated in 2006 and follow-up was scheduled to end in January 2009. Only exception was a poster described in a review article that reported “a 7 percent increase in adherence to blood pressure lowering medication and a nonsignificant 4 percent increase in adherence to statins.”[5]
Observational studies of copayment reductions conducted between 2010 and 2012[1]
Reported. Methodologically strongest of these studies indicated average MPR effect sizes of 1.4 to 4.0 percentage points, or about 5 to 15 added days of therapy per year.
 
Common problems identified in an editorial by Fairman and Curtiss included:
 
*No information about generic drug utilization
*No information about net payer cost
*Incomplete reporting of key study details (e.g., sample composition method)
*Failure to control for or report mail order utilization rates or 90-day fill policies in all but one study
Choudhry (2011). Post-MI FREEE study of providing evidence-based medications to heart attack survivors free of charge after hospital discharge; block-randomized controlled trial (health plans, rather than patients, were randomized).[6] Nonequivalent communications to study groups (see below).
 
Primary outcome was occurrence of revascularization procedure or major vascular event (heart attack, stroke, unstable angina, CHF, or in-hospital death from CVD)
 
 
 
 
 
 
 
 
 
 
Primary outcome reported.
*No significant difference between the study groups.
*No pre-intervention values reported for either group.
Secondary outcome: Rate of major events as defined above
 
 
One secondary outcome reported. *Slightly and significantly lower for free medication group – 21.5 versus 23.3 events per 100 person-years
 
*No pre-intervention values reported for either group
Secondary outcome: Published analysis plan indicated that utilization outcomes would include annual rates of doctor visits, ER visits, and hospital admissions.
*Secondary outcome results not reported, although follow-up ended in November 2010
Secondary outcome: Published analysis plan indicated that investigators would report the outcome including out-of-hospital death data from the CDC.
*Secondary outcome results not yet reported; out-of-hospital death data are subject to a 1.5 to 2-year time lag and therefore should have been available for all patients by November 2012
Results for originally planned minimum follow-up time of one year
 
*Results not reported; results reflect minimum three-month follow-up. 
CHF=congestive heart failure; CVD=cardiovascular disease; MPR=medication possession ratio, a commonly used measure of medication adherence, defined as days supply of dispensed medication divided by calendar time in days.

As Fred Curtiss and I pointed out in an editorial in April 2012, the most important of these studies is Post-MI FREEE, which, as its name implies, was an assessment of the effect of providing free medications to patients after hospital discharge for a heart attack. Despite Post-MI FREEE’s unique status as the first randomized study of copayment reductions, its value was greatly diminished by three critically important limitations.[7]
First was that the utilization outcomes pre-specified in the analysis plan have never been reported, although patient follow-up ended about two and one-half years ago. Nonreporting of pre-specified outcomes is always a cause for concern but was especially problematic in Post-MI FREEE because of the block-randomized design that assigned entire plans (rather than individual members) to free medication or usual coverage (copayments).  Specifically, because there was no control on medical reimbursement levels in the block randomization process, it is possible that the study groups did not have equivalent payment rates for the same medical services at baseline. Thus, it is possible that the post-intervention cost outcomes give an incomplete or even misleading picture of the effect of the intervention on service utilization. Compounding the problem, no baseline (pre-intervention) costs were reported for either study group.

Second patients in only one group—the group that did not receive free medications—were told that they would be on medications for a long time and would have to give up a lot of fun stuff. For both groups, letters advising patients about the study contained a list of medications that are recommended “to keep your heart strong,” but the usual coverage (control group) letters advised patients that “You may be on some medications for many years …” The usual coverage group letters also, unlike the free coverage letter, advised patients of American Heart Association recommendations for diet, exercise, and limitations on alcohol intake.
Third was a change in the study analysis plan to reduce the minimum length of follow-up from twelve months to three months because of lower-than-expected participation. A typical and sensible approach to handling this unavoidable development would have been a sensitivity analysis limited to study patients who had the originally intended twelve month minimum follow-up. However, although 65% of the study sample met that criterion, meaning that plenty of cases were available to perform a sensitivity analysis, none was performed.[7] As a result, we don’t know if the intervention’s effects would have persisted beyond the short follow-up period. Previous research suggests cause for concern on this point because of the large amount of noncompliance that takes place after the first three months of therapy but before the end of the first year. This decision was especially unfortunate because previous research in a similar patient population suggests that the survival benefits of statins are not observed until about 24 months of therapy.[8]
And—even with all of those potential biases toward finding optimal results for the free-medication group—the reported findings showed that only 12% of patients given free drugs, compared with 9% of those who paid copayments, were fully adherent to their drugs.[6]

Key Takeaway Points for Health Care Payers:
What To Do When the Studies Have “Gone to Flowers, Every One …” 
 
The history of VBID (copayment reduction) promotion is a notable but certainly not isolated example of noise outstripping scientific evidence in modern culture.  Much of the promised data have never been reported, and those that have been published suggest that the effects of copayment reductions are modest and may not be worth the added cost. It’s been said before, but it bears repeating: the strongest voices do not necessarily have the most important or valuable message.

So, when you read or hear about a proposed policy change in health care, whether it is VBID or a different proposal, consider the following questions:

  • What was the follow-up time for the study being cited as evidence for the policy change? Does it represent a time frame during which one can reasonably expect meaningful benefits to occur?  If not, the results may represent confounding or the effects of another factor not measured in the study and/or presented in the report.
        
  • Does the study follow-up period represent the time frame during which you will have to pay for the recommended intervention? (The investigators may have “lived with” this intervention for just a few short months; if you implement it, you may have its costs with you for a long time—perhaps long after intervention benefits stop accruing).
        
  • To what degree do the characteristics reported by the authors match to those of the group to which you might apply the proposed intervention? For example, an intervention tested in a group of university professors may not have the same effect if implemented in an auto-workers’ union.
        
  • Do the outcome measures represent your cost? Some copayment reduction studies report only total cost, which does not reflect the actual cost of the program to the payer. It is essential to report the net cost after taking loss of copayment revenues (offsets) from patients into account.
        
  • If a pre-intervention versus post-intervention study was performed, did the investigators report the pre-intervention values for the outcome measures? Without this information, it’s hard for a decision maker to understand the actual effects of an intervention relative to the baseline status of the study sample.
        
  • How was the intervention—such as “value-based”—defined? For example, the results of a “value-based” intervention that promoted a high-efficiency provider network provide little or no information about a “value-based” program that uses copayment reduction to incentivize medication use.
        
  • Does the program described in the report include multiple components that you might or might not want to implement simultaneously? For example, if the program combined copayment reductions with other strategies, such as gym club memberships or educational interventions, it is hard to tell which program components were responsible for the outcomes. (And, if you cannot tell from reading the reporting what features were included in the program, it is reasonable to be skeptical about its findings.)
        
  • Have relevant details for both study groups been presented in the study report? For example, for an intervention targeted to improve drug adherence, have the authors reported whether formularies, copayment levels, and utilization management policies (e.g., step therapy and prior authorization) were reasonably equivalent at baseline or otherwise explained the comparability of the groups (e.g., both groups were drawn from the same health plan or employer group with a uniform benefit design)?
        
  • Have all the pre-specified outcomes been reported? (Note: This one takes a little detective work but is extremely helpful.) Use PubMed or, for randomized trials, www.clinicaltrials.gov to find the investigators’ pre-specified analysis plan.[9] If an analysis plan is available, compare it to the study report. If outcomes are missing from the report, query the investigators. If the outcomes are still not reported a considerable length of time after the end of follow-up, it is reasonable to consider the possibility of publication bias—that is, the outcomes were not reported because they did not fit the predispositions of the investigators, study sponsors, journal peer reviewers, or editors.
        
  • Have previous studies of the same topic gone unreported, as has been the case in VBID? If so, it is reasonable to be skeptical about even the most promising research findings. For those of you familiar with probability theory—the underlying logic of statistical significance testing is that a certain number of results within a sampling distribution (a hypothetical sequence of repeated samples) will be statistically significant based on chance (sampling error) alone. If the results for actual samples have gone unreported, information necessary to interpret the results of statistical tests for any one sample is missing.


For health care payers, who are routinely bombarded today with assertions about "evidence-based" policies or services, these considerations should be viewed as critically important factors in decision making about value-based designs—or any proposed intervention, for that matter. When making decisions that affect not only pocketbooks, but also patient well-being, the best policy is circumspection: review the evidence that is reported, take into account the absence of evidence that has gone unreported, carefully examine the applicability of the available information for your population and setting, and don’t be afraid to ask questions.

 
[1] Fairman KA, Curtiss FR. What do we really know about VBID? Quality of the evidence and ethical considerations for plan sponsors. J Manag Care Pharm. 2011;17(2):156-174.
[2] Chernew ME, Shah MR, Wegh A, et al. Impact of decreasing copayments
on medication adherence within a disease management environment. Health
Aff (Millwood). 2008;27(1):103-12; Fairman KA, Curtiss FR. Making the world safe for evidence-based  policy: let’s slay the biases in research on value-based insurance design. J Manag Care Pharm. 2008;14(2):198-204.
[3] Chernew ME, Juster IA, Shah M, et al. Evidence that value-based insurance can be effective. Health Aff (Millwood). 2010;29(3):530-36.
[4] Spaulding A, Fendrick AM, Herman WH, et al. A controlled trial of value-based insurance design - the MHealthy: Focus on Diabetes (FOD) trial. Implement Sci. April 2009.
[5] Choudhry NK, Rosenthal MB, Milstein A. Assessing the evidence for
value-based insurance design. Health Aff (Millwood). 2010;29(11):1988-94.
[6] Choudhry NK, Avorn J, Glynn RJ, et al.; Post-Myocardial Infarction Free Rx Event and Economic Evaluation (MI FREEE) trial. Full coverage for preventive medications after myocardial infarction. N Engl J Med. 2011;365(22):2088-2097; Choudhry NK, Brennan T, Toscano M, et al. Rationale and design of the Post-MI FREEE trial: a randomized evaluation of first-dollar drug coverage for post-myocardial infarction secondary preventive therapies. Am Heart J. 2008;156(1):31-37.
[7] Fairman KA, Curtiss FR. VBID, the PPACA, and FREEE medications: Did politics trump the evidence about cost sharing? J Manag Care Pharm. 2012;18(2):146-156.
[8] Bavry AA, Mood GR, Kumbhani DJ, Borek PP, Askari AT, Bhatt DL.
Long-term benefit of statin therapy initiated during hospitalization for an
acute coronary syndrome: a systematic review of randomized trials. Am J
Cardiovasc Drugs. 2007;7(2):135-41.
[9] At www.clinicaltrials.gov, you can search by name of a medical condition, drug, or investigator.

Saturday, June 8, 2013


Value-Based Insurance Design Defined as Copayment Reduction:
Where Have All the Studies Gone? Part 1

“The great enemy of the truth is very often not the lie—deliberate, contrived, and dishonest—but the myth: persistent, persuasive, and unrealistic. Too often we hold fast to the clichés of our forebears. We subject all facts to a prefabricated set of interpretations. We enjoy the comfort of opinion without the discomfort of thought.

President John F. Kennedy, speaking at Yale University on June 11, 1962

In a February posting on the wisdom of proposals to invest Medicare funds in value-based insurance designs (VBID) without testing them first, I promised to provide an updated look at the current quality of evidence regarding VBID. Since then, there have been three interesting new developments on the VBID front, all in April 2013:
(1) A report from the Partnership for Sustainable Health Care (PSHC) recommended that cost-sharing structures should include “differentiation to encourage the use of high-value services and providers” as a way to achieve “savings from improved adherence to preventive measures and evidence-based care, lower utilization of unnecessary services, and the use of more efficient, higher-quality providers.”[1] PSHC’s assessment echoed a previous description of VBID by the National Coalition on Health Care as a “game changer.”[2]

(2) The Chairman of the Medicare Payment Advisory Commission (MEDPAC) testified before the U.S. House Subcommittee on Health, Committee on Energy and Commerce, recommending that VBID be used in Medicare.[3]
(3) The Center for Value-Based Insurance Design (CVBID) at the University of Michigan issued an interesting policy brief on the use of VBID in health plans that are “grandfathered” (allowed to continue in their present form so long as they do not make major benefit design cuts) under the terms of the Affordable Care Act, or PPACA.[4]

These recent developments make it all the more important to ask now: what is driving all this attention?

What Is VBID, Exactly? Depends on When You Asked
Before addressing the current state of research regarding VBID, it is helpful to clarify the subtle but important shift in the meaning of the term “value-based” that has taken place over the past several years. For example, the PSHC report refers to “value-based payment approaches” using “a range of models that include incentives for patient safety, bundled payments, accountable care organizations, and global payments.”[1] Also defined as VBID are financial incentives for patients “to obtain care from providers with a demonstrated ability to deliver quality, efficient health care,” as well as incentives to quit smoking, lose weight, or join diabetes prevention programs.[1]

These definitional shifts have considerably expanded the original concept of VBID (called “benefit-based copay” for prescription drugs in 2001), which was “a system of cost sharing that tailors copayments at the point of service to the evidence-based value of specific services for targeted groups of patients.”[5] In other words, a definition of “value-based” that initially referred to a novel concept—reduced copayments for “high-value” medications—has now been expanded to include bundled payment methods, provider network management, and wellness promotion. The change is notable, since all of these “VBID” features have been basic (albeit somewhat inconsistently used) mainstays of managed care for the past several decades.
This conceptual expansion is perhaps not surprising. More than a decade after first being proposed, copayment reductions targeted to "high-value" drugs have generally had a low adoption rate by commercial insurers and employers, hovering at around the 20% range for some years now.[6] Only 24% of employers in 2012 reported using “reduced copay for specific drug classes/health conditions” in the Pharmacy Benefit Management Institute’s annual prescription drug benefit cost and plan design report,[7] and only 11% of respondents to the Towers-Watson annual Employer Survey on Purchasing Value in Health Care said that they were using “value-based benefit designs (e.g., different levels of coverage based on value or cost of services)” in 2013.[8]

Quality of the Evidence for Copayment Reductions
But what of the quality of evidence regarding copayment reductions, the original linchpin of the “benefit-based copay”? This point is becoming increasingly important for plan sponsors now because, as the CVBID piece on the PPACA correctly observed, plans can lower copayments without losing “grandfather” status, but they cannot substantially increase them. Within a grandfathered plan, there is little opportunity to offset cost-sharing decreases in one therapy class with increases in another.

Unfortunately—in a pattern of reporting (and nonreporting) of research results that raises important questions about publication bias in health policy research—the history of utilization and cost outcomes for copayment reductions is much more notable for what was not said than for what was. I’ll post more on that topic later this week, on the fifty-first anniversary of President Kennedy’s 1962 commencement address at Yale. We’ll see if VBID studies have “gone to flowers, every one.” And meanwhile, for those of you who have no idea what the song lyric references in this title or text mean, help is available here, in an article about a smash hit that was also released in 1962.


[1] Partnership for Sustainable Health Care. Strengthening affordability and quality in America’s health care system. April 2013.
[2] Center for Value-Based Insurance Design. Press release. For immediate release: key stakeholders support V-BID. April 17, 2013.
[3] Medicare Payment Advisory Commission. Reforming Medicare’s benefit design. Statement of Glenn M. Hackbarth, JD, before the Subcommittee on Health, Committee on Energy and Commerce, U.S. House of Representatives. April 11, 2013.
[4] Center for Value-Based Insurance Design. V-BID and grandfathered health plans: promoting high-value services and controlling costs.
[5] Fendrick AM, Smith DG, Chernew ME, Shah SN. A benefit-based copay
for prescription drugs: patient contribution based on total benefits, not drug
acquisition cost. Am J Manag Care. 2001;7(9):861-67.
[6] Fairman KA, Curtiss FR. What do we really know about VBID? Quality of the evidence and ethical considerations for plan sponsors. J Manag Care Pharm. 2011;17(2):156-174.
[7] Pharmacy Benefit Management Institute. 2012-2013 prescription drug benefit cost and plan design report. 2012.
[8] Towers Watson. Reshaping health care: performers leading the way. 18th Annual Towers Watson/National Business Group on Health Employer Survey on Purchasing Value in Health Care. 2013.

Tuesday, April 23, 2013


Oops, They Did It Again:
Smackdown Shows Why “Small Stuff” Actually Matters. A Lot.

With a new job starting this week, I thought that I wouldn’t have time to write a blog post.  That is, until my husband noticed one of our favorite Wall Street Journal columns by Carl Bialik, the “numbers guy,” this weekend.[1] Bialik highlighted the rapid-fire debate that ensued after students in an Applied Econometrics class at the University of Massachusetts were assigned the task of replicating a published study. Twenty-eight-year-old doctoral student Thomas Herndon chose “Growth in a Time of Debt,” authored by Carmen Reinhart and Kenneth Rogoff—two Harvard economists with sterling reputations—and published in the American Economic Review in 2010. The study, described by the New York Magazine “Daily Intelligencer” as “massively influential,”[2] found a strong association between increasing government debt and declining economic growth (e.g., average annual growth of only 1.7% for countries with debt exceeding 90% of gross domestic product versus 3.7% when debt was less than 30%).[1,3] Political debate today being what it is, the study has been cited repeatedly by leading conservatives to justify cutbacks in government spending.[2]

Much to his credit, Herndon wrote to the authors and requested a copy of their dataset and, after not receiving a reply initially, he repeated his query. Much to their credit, Reinhart and Rogoff sent the data, in an Excel spreadsheet, along with a note telling him to “feel free to publish whatever results” he found, according to Herndon.[2] That’s when things started to get interesting:

“I clicked on cell L51,” Herndon said, describing the incident in a recent interview, “and saw that [the authors] had only averaged rows 30 through 44, instead of rows 30 through 49.”[2] In other words, Reinhart and Rogoff had inadvertently excluded five data points (countries) from their calculations. In a reanalysis including the omitted nations, high debt was associated with 0.2% annual growth.[1] And, again with political debate today being what it is, the reanalysis has been seized upon by progressives to argue that Reinhart and Rogoff’s central thesis of a link between debt and suppression of economic growth “has been substantially weakened”[4] and "there’s no question that the austerity movement has been dealt a major blow” by Herndon.[2]

Yet, in all the hubbub over these findings—and the justifiable pride of U Mass in the good detective work done by one of its own—several critically important points were missed. And, although the desire to capitalize politically on a mistake made by someone with a different viewpoint is perhaps predictable, none of the most important aspects of this story has anything whatsoever to do with liberal or conservative politics. They don’t even have anything to do with academic competition, although that aspect of the story is interesting. They are all, in fact, about research methods. (C’mon, you had to know I’d say that—this is a research methods blog, after all.)

First, although news coverage of this debacle described it as remarkable, situations of this type are commonplace. As a research consultant and then journal editor, I met many highly talented professionals with advanced degrees—almost all of whom were knowledgeable in their fields, and none of whom had received any training in managing a dataset and data analysis to prevent calculation errors. One might believe that this training deficit is of no importance—until looking at the history of publication missteps and discovering that what happened to Reinhart and Rogoff happens all the time.

Not convinced? Peek “inside the statistical black box,” as I reported in 2006 and in Health Care Research Done Right:

·  A highly influential study on the economic effects of divorce, published by a respected Stanford University sociologist, was in the subsequent decade cited in 175 popular press stories, 250 law review articles, 348 social science articles, numerous court cases, and President Clinton’s 1996 budget. The study’s author, Lenore Weitzman, later recalled that 14 new divorce laws were passed in the state of California (the study’s setting) as a result of her finding that no-fault divorce had “disastrous” effects on women and children. But when initially reported, the study’s findings were so at variance with previously published analyses of the same research question that other social scientists began requesting access to the data. Unfortunately, Weitzman refused to provide them for many years, and no reanalysis took place until 1996—more than ten years after the original publication. Reanalysis showed that Weitzman’s widely used calculations were wrong—the result of keying errors and calculation mistakes made by one or more graduate students working under her supervision.[5,6]

·   Because of a miscommunication, a data analytics consulting firm working for Arizona’s Independent Redistricting Commission inadvertently tallied both active and inactive voters in calculating district sizes. The error was discovered three months after the commission thought its work was completed, and a local newspaper reported in April 2002 that candidates didn’t “know where to collect the signatures and donations they need to run for office.”[5]
 
·   The University of East Anglia’s Climatic Research Unit (CRU) reported in 2009 that the raw data necessary to replicate the “hockey stick” graph documenting the existence of human-caused global warming were no longer available to anyone—including the CRU. For years, the CRU had failed to respond to Freedom of Information Act requests for the data because after “cleaning” the data (using undocumented methods), CRU threw them away when it moved to a new building. Independent investigations conducted by a team of scientists and by a United Kingdom House of Commons panel reached the same conclusion: the CRU scientists were dedicated, but disorganized. The original, raw data had been irretrievably lost and could not be recovered.[6]

 If there is anything remarkable here, it is that Rogoff and Reinhart—unlike Weitzman and the CRU—provided their dataset in what was apparently (given that Herndon was able to complete his assignment) a reasonable period of time. That shouldn’t be remarkable behavior, but the history described above suggests that it is.

Thus, the sequence of events surrounding the reanalysis of Rogoff and Reinhart’s work is just one more reminder of the need to incorporate systematic and structured data quality control measures in conducting research projects. As described in Health Care Research Done Right, five relatively simple procedures will, if practiced consistently, prevent small errors from sidetracking otherwise important research.[6]

Second, as the “numbers guy” reminds us, both of these analyses of the relationship between debt and growth have limited value, because they measure association. Only. Not causation. “The findings show a link between debt and GDP,” Bialik correctly observed, “not which causes which.”[1] To assess causation, directly addressing the policy question of whether government austerity or liberal spending policies produce growth, would require an experiment in which reasonably comparable states or countries implement private-sector versus government-growth policies and measure the results. (Although outside the scope of this blog, one could argue that today’s between-state differences in economic policies are, to some extent, meeting this need, and we should measure those outcomes.)

And this brings me to my final point. For anyone to argue that any single research finding proves anything is hubris, especially when previously reported results point in a different direction. (One possible exception is when a study is truly groundbreaking, exceptionally well done, and nationally representative, such as the Rand Health Insurance Experiment.) Herndon did well, and I hope that he benefits from his willingness to peek inside the statistical black box. However, given the methodological limitations of measuring association instead of causation, neither his work nor the work that he replicated provides actionable information. There was no "major blow" here, just as there was no compelling finding in the original analysis. From a policy perspective, this was much ado about (almost) nothing.

So, for nonresearchers, the moral of this story is to pay attention to research design, remembering again the important lesson that association does not prove causation. And we researchers should hope that Herndon’s instructors encourage him to do what we all should be doing—interpreting our research findings with humility and in the context of work performed by others. Even Harvard economists.

 
[1] Bialik C. The numbers guy. Spreadsheet slips not economists’ only problem. WSJ. April 20-21, 2013.
[2] Roose K. Meet the 28-year-old grad student who just shook the global austerity movement. April 18, 2013.
[3] Reinhart CM, Rogoff KS. Growth in a time of debt. NBER working paper no. 15639. January 2010.
[4] How Thomas Herndon, a student, took on Harvard economists and won. April 18, 2013.
[5] Fairman KA. Peeking inside the statistical black box: how to analyze quantitative information and get it right the first time. J Manag Care Pharm. 2006;13(1):70-74.
[6] Fairman KA. Health Care Research Done Right: A Journal Editor Shares Practical Tips and Techniques for High Qualityand Efficiency. Denver, CO: Outskirts Press; 2012.

 

Wednesday, April 10, 2013


Percentages Are Funny:
Why Absolutes are Absolute in Research Reporting

A member of our family recently picked up a prescription at our local pharmacy, finding that our out-of-pocket cost for the medication had increased by 367% over a six-month period of time. Assuming that the drug is medically necessary and appropriate for this patient, which of the following is an accurate assessment of the situation?

(a) Sure sign of price gouging by pharmaceutical manufacturers—we need cost controls, and we need ‘em soon
(b) Sure sign of the need for copayment relief—it is “penny wise and pound foolish” to discourage members from purchasing appropriate medication
(c) Means nothing at all

As the title of this posting suggests, the correct answer is (c). The out-of-pocket cost change in question was a mere $2.46—from $0.67 to $3.13—for a one-month supply, about eight cents per day. That’s not enough to merit even a passing glance at the price, let alone “cost-related nonadherence.” (Okay, we noticed, but I think that’s just because my husband and I are a little . . . well, geeky.)[1]

This simple example provides a good illustration of the reason that research-reporting guidelines recommend presentation of absolute numbers—not just relative measures, such as hazard ratios, odds ratios, or percentages—in describing quantitative findings. Percentages (and other relative measures) are funny. They show us how study groups or time periods compare with one another, in relative terms, but they tell us little or nothing about what those differences mean in practical terms.

For example, a mortality odds ratio of 3.23 for Drug A, with Drug B as the reference category, could mean that a patient has a 52% probability of death using Drug A compared with a 25% probability using Drug B—at an additional 27,000 deaths per 100,000 treated patients, clearly a risk worth paying attention to. Or, the same odds ratio could mean that the probability of death is 0.00004% with Drug B and 0.000129% with Drug A—1.29 per million, the approximate probability of getting struck by lightning in any given year. (If you are not familiar with these calculations, see the note below for an explanation.)[2]

In this context, the rationale for the following CONSORT (CONsolidated Standards Of Reporting Trials) guidance, as described in its “explanation and elaboration” document, should be clear:

For each outcome, study results should be reported as a summary of the outcome in each group (for example, the number of participants with or without the event and the denominators, or the mean and standard deviation of measurements), together with the contrast between the groups, known as the effect size. For binary outcomes, the effect size could be the risk ratio (relative risk), odds ratio, or risk difference; for survival time data, it could be the hazard ratio or difference in median survival time; and for continuous data, it is usually the difference in means. Confidence intervals should be presented for the contrast between groups. … For binary outcomes, presentation of both absolute and relative effect sizes is recommended.[3]
In other words, relative measures (along with estimates of uncertainty, usually confidence intervals) are necessary—but not sufficient—to inform the reader of a study’s results. The CONSORT authors provided two tables from previously reported research as helpful examples; for illustration, I show an adapted version of just the first row of each table below:

Table 1. Example of Reporting Binary Outcomes
 
Number (%)
 
Endpoint
Etanercept (n=30)
Placebo (n=30)
Risk Difference
(95% CI)
Achieved PsARC at 12 weeks
26 (87)
7 (23)
63% (44 to 83)

CI=confidence interval; PsARC=psoriatic arthritis response criteria.[3]

Table 2. Example of Reporting Continuous Outcomes
 
Exercise Therapy (n=65)
Control (n=66)
 
 
Baseline Mean [SD]
12 Months Mean [SD]
Baseline Mean [SD]
12 Months Mean [SD]
Adjusted Difference (95% CI) at 12 Months
Function score (0-100)
64.4 (13.9)
83.2 (14.8)
65.9 (15.2)
79.8 (17.5)
4.52 (-0.73-9.76)

CI=confidence interval; SD=standard deviation.[3]

Note also that in the second example shown, baseline (pre-intervention) as well as follow-up values are shown to enable the reader to assess the clinical significance of the change amounts in light of the group characteristics prior to the intervention.
The practice of reporting outcomes measured at baseline is recommended by CONSORT “so that readers can assess how similar [the study groups] were” but is unfortunately not always followed even in observational (nonrandomized cohort) studies of interventions, where baseline comparability of the study groups is a critically important issue.[4] For example, observational assessments of therapy outcomes for employer groups that implemented step therapy programs, compared with groups that had no step therapy, have failed to report baseline values on even basic key outcome measures including utilization of the target drug classes and health care costs.[5]

Practical Take-Away Points: Insist on Absolutes. Absolutely.
Without information about both the absolute and relative effects of study variables of interest, it is impossible to determine whether results represent practically/clinically meaningful outcomes or statistical artifact, often due to the enormous sample sizes that are commonplace in health care databases today. (With a sufficiently large number of study subjects, even completely meaningless changes can be statistically significant). The most informative reports indicate baseline values, follow-up values, and absolute change amounts (follow-up minus baseline), in addition to measures of relative difference (e.g., odds ratios) and uncertainty (e.g., confidence intervals).

So if the report of an intervention study with an observational design fails to provide baseline characteristics of the study subjects, including baseline values of the outcome measures, or if it fails to report absolute post-intervention change amounts, its worth is limited. Without this information, there is no way to determine the comparability of the groups prior to the intervention or to get a sense of the practical/clinical effect of the intervention on the outcome.
If a randomized study report fails to provide baseline values on the outcome measures, the report is less informative than it could or should be; however, the problem is usually not a fatal flaw because the randomization process should produce comparable groups. A possible exception is block randomization (randomization of groups instead of individual subjects, such as randomizing all patients treated by a particular physician instead of randomizing individual patients), because the blocks may differ in ways that affect response to the intervention.
And don’t be shy. If you don’t see the information you need in a study report, remember that journals provide contact information for the first author for a good reason—so that you can write to him or her if you have a question. It’s appropriate to ask the author to provide missing information and to ask follow-up questions (nicely) if you have them.

But in a broader sense, a good general rule is this: the less the investigators conformed to reporting guidelines, the more cautious you should be about the validity of the study findings. For that reason, if you have some research training and use research results in your work, it is a good idea to read through the CONSORT or STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) Explanation and Elaboration documents.[3,6] Either will provide a good general sense of the purpose and spirit of reporting guidelines, and knowing what to expect from a high-quality research report will prove invaluable.

[1] Fairman KA, Rucker ML. Fractal mathematics in managed care? How a simple and revealing analysis could improve the forecasting and management of medical costs and events. J Manag Care Pharm. 2009;15(4):351-358.

[2] Calculation note: odds=probability/(1–probability)—in other words, the odds of an event are defined as the probability divided by the probability of the alternative. The odds ratio for A versus B=odds[A]÷odds[B]. For the sake of providing a simplified example, the results shown in this posting are slightly affected by rounding error.


[4] Des Jarlais DC, Lyles C, Crepaz N; TREND Group. Improving the quality of nonrandomized evaluations of behavioral and public health interventions: the TREND statement. Am J Public Health. 2004;94(3):361-366.

[5] Motheral BR. Pharmaceutical step-therapy interventions: a critical review of the literature. J Manag Care Pharm. 2011;17(2):143-155.

[6] Vandenbroucke JP, von Elm E, Altman DG, et al. Strengthening the Reporting of Observational Studies in Epidemiology (STROBE): explanation and elaboration. PLoS Med. 2007;4(10):e296.