Tuesday, April 23, 2013


Oops, They Did It Again:
Smackdown Shows Why “Small Stuff” Actually Matters. A Lot.

With a new job starting this week, I thought that I wouldn’t have time to write a blog post.  That is, until my husband noticed one of our favorite Wall Street Journal columns by Carl Bialik, the “numbers guy,” this weekend.[1] Bialik highlighted the rapid-fire debate that ensued after students in an Applied Econometrics class at the University of Massachusetts were assigned the task of replicating a published study. Twenty-eight-year-old doctoral student Thomas Herndon chose “Growth in a Time of Debt,” authored by Carmen Reinhart and Kenneth Rogoff—two Harvard economists with sterling reputations—and published in the American Economic Review in 2010. The study, described by the New York Magazine “Daily Intelligencer” as “massively influential,”[2] found a strong association between increasing government debt and declining economic growth (e.g., average annual growth of only 1.7% for countries with debt exceeding 90% of gross domestic product versus 3.7% when debt was less than 30%).[1,3] Political debate today being what it is, the study has been cited repeatedly by leading conservatives to justify cutbacks in government spending.[2]

Much to his credit, Herndon wrote to the authors and requested a copy of their dataset and, after not receiving a reply initially, he repeated his query. Much to their credit, Reinhart and Rogoff sent the data, in an Excel spreadsheet, along with a note telling him to “feel free to publish whatever results” he found, according to Herndon.[2] That’s when things started to get interesting:

“I clicked on cell L51,” Herndon said, describing the incident in a recent interview, “and saw that [the authors] had only averaged rows 30 through 44, instead of rows 30 through 49.”[2] In other words, Reinhart and Rogoff had inadvertently excluded five data points (countries) from their calculations. In a reanalysis including the omitted nations, high debt was associated with 0.2% annual growth.[1] And, again with political debate today being what it is, the reanalysis has been seized upon by progressives to argue that Reinhart and Rogoff’s central thesis of a link between debt and suppression of economic growth “has been substantially weakened”[4] and "there’s no question that the austerity movement has been dealt a major blow” by Herndon.[2]

Yet, in all the hubbub over these findings—and the justifiable pride of U Mass in the good detective work done by one of its own—several critically important points were missed. And, although the desire to capitalize politically on a mistake made by someone with a different viewpoint is perhaps predictable, none of the most important aspects of this story has anything whatsoever to do with liberal or conservative politics. They don’t even have anything to do with academic competition, although that aspect of the story is interesting. They are all, in fact, about research methods. (C’mon, you had to know I’d say that—this is a research methods blog, after all.)

First, although news coverage of this debacle described it as remarkable, situations of this type are commonplace. As a research consultant and then journal editor, I met many highly talented professionals with advanced degrees—almost all of whom were knowledgeable in their fields, and none of whom had received any training in managing a dataset and data analysis to prevent calculation errors. One might believe that this training deficit is of no importance—until looking at the history of publication missteps and discovering that what happened to Reinhart and Rogoff happens all the time.

Not convinced? Peek “inside the statistical black box,” as I reported in 2006 and in Health Care Research Done Right:

·  A highly influential study on the economic effects of divorce, published by a respected Stanford University sociologist, was in the subsequent decade cited in 175 popular press stories, 250 law review articles, 348 social science articles, numerous court cases, and President Clinton’s 1996 budget. The study’s author, Lenore Weitzman, later recalled that 14 new divorce laws were passed in the state of California (the study’s setting) as a result of her finding that no-fault divorce had “disastrous” effects on women and children. But when initially reported, the study’s findings were so at variance with previously published analyses of the same research question that other social scientists began requesting access to the data. Unfortunately, Weitzman refused to provide them for many years, and no reanalysis took place until 1996—more than ten years after the original publication. Reanalysis showed that Weitzman’s widely used calculations were wrong—the result of keying errors and calculation mistakes made by one or more graduate students working under her supervision.[5,6]

·   Because of a miscommunication, a data analytics consulting firm working for Arizona’s Independent Redistricting Commission inadvertently tallied both active and inactive voters in calculating district sizes. The error was discovered three months after the commission thought its work was completed, and a local newspaper reported in April 2002 that candidates didn’t “know where to collect the signatures and donations they need to run for office.”[5]
 
·   The University of East Anglia’s Climatic Research Unit (CRU) reported in 2009 that the raw data necessary to replicate the “hockey stick” graph documenting the existence of human-caused global warming were no longer available to anyone—including the CRU. For years, the CRU had failed to respond to Freedom of Information Act requests for the data because after “cleaning” the data (using undocumented methods), CRU threw them away when it moved to a new building. Independent investigations conducted by a team of scientists and by a United Kingdom House of Commons panel reached the same conclusion: the CRU scientists were dedicated, but disorganized. The original, raw data had been irretrievably lost and could not be recovered.[6]

 If there is anything remarkable here, it is that Rogoff and Reinhart—unlike Weitzman and the CRU—provided their dataset in what was apparently (given that Herndon was able to complete his assignment) a reasonable period of time. That shouldn’t be remarkable behavior, but the history described above suggests that it is.

Thus, the sequence of events surrounding the reanalysis of Rogoff and Reinhart’s work is just one more reminder of the need to incorporate systematic and structured data quality control measures in conducting research projects. As described in Health Care Research Done Right, five relatively simple procedures will, if practiced consistently, prevent small errors from sidetracking otherwise important research.[6]

Second, as the “numbers guy” reminds us, both of these analyses of the relationship between debt and growth have limited value, because they measure association. Only. Not causation. “The findings show a link between debt and GDP,” Bialik correctly observed, “not which causes which.”[1] To assess causation, directly addressing the policy question of whether government austerity or liberal spending policies produce growth, would require an experiment in which reasonably comparable states or countries implement private-sector versus government-growth policies and measure the results. (Although outside the scope of this blog, one could argue that today’s between-state differences in economic policies are, to some extent, meeting this need, and we should measure those outcomes.)

And this brings me to my final point. For anyone to argue that any single research finding proves anything is hubris, especially when previously reported results point in a different direction. (One possible exception is when a study is truly groundbreaking, exceptionally well done, and nationally representative, such as the Rand Health Insurance Experiment.) Herndon did well, and I hope that he benefits from his willingness to peek inside the statistical black box. However, given the methodological limitations of measuring association instead of causation, neither his work nor the work that he replicated provides actionable information. There was no "major blow" here, just as there was no compelling finding in the original analysis. From a policy perspective, this was much ado about (almost) nothing.

So, for nonresearchers, the moral of this story is to pay attention to research design, remembering again the important lesson that association does not prove causation. And we researchers should hope that Herndon’s instructors encourage him to do what we all should be doing—interpreting our research findings with humility and in the context of work performed by others. Even Harvard economists.

 
[1] Bialik C. The numbers guy. Spreadsheet slips not economists’ only problem. WSJ. April 20-21, 2013.
[2] Roose K. Meet the 28-year-old grad student who just shook the global austerity movement. April 18, 2013.
[3] Reinhart CM, Rogoff KS. Growth in a time of debt. NBER working paper no. 15639. January 2010.
[4] How Thomas Herndon, a student, took on Harvard economists and won. April 18, 2013.
[5] Fairman KA. Peeking inside the statistical black box: how to analyze quantitative information and get it right the first time. J Manag Care Pharm. 2006;13(1):70-74.
[6] Fairman KA. Health Care Research Done Right: A Journal Editor Shares Practical Tips and Techniques for High Qualityand Efficiency. Denver, CO: Outskirts Press; 2012.

 

Wednesday, April 10, 2013


Percentages Are Funny:
Why Absolutes are Absolute in Research Reporting

A member of our family recently picked up a prescription at our local pharmacy, finding that our out-of-pocket cost for the medication had increased by 367% over a six-month period of time. Assuming that the drug is medically necessary and appropriate for this patient, which of the following is an accurate assessment of the situation?

(a) Sure sign of price gouging by pharmaceutical manufacturers—we need cost controls, and we need ‘em soon
(b) Sure sign of the need for copayment relief—it is “penny wise and pound foolish” to discourage members from purchasing appropriate medication
(c) Means nothing at all

As the title of this posting suggests, the correct answer is (c). The out-of-pocket cost change in question was a mere $2.46—from $0.67 to $3.13—for a one-month supply, about eight cents per day. That’s not enough to merit even a passing glance at the price, let alone “cost-related nonadherence.” (Okay, we noticed, but I think that’s just because my husband and I are a little . . . well, geeky.)[1]

This simple example provides a good illustration of the reason that research-reporting guidelines recommend presentation of absolute numbers—not just relative measures, such as hazard ratios, odds ratios, or percentages—in describing quantitative findings. Percentages (and other relative measures) are funny. They show us how study groups or time periods compare with one another, in relative terms, but they tell us little or nothing about what those differences mean in practical terms.

For example, a mortality odds ratio of 3.23 for Drug A, with Drug B as the reference category, could mean that a patient has a 52% probability of death using Drug A compared with a 25% probability using Drug B—at an additional 27,000 deaths per 100,000 treated patients, clearly a risk worth paying attention to. Or, the same odds ratio could mean that the probability of death is 0.00004% with Drug B and 0.000129% with Drug A—1.29 per million, the approximate probability of getting struck by lightning in any given year. (If you are not familiar with these calculations, see the note below for an explanation.)[2]

In this context, the rationale for the following CONSORT (CONsolidated Standards Of Reporting Trials) guidance, as described in its “explanation and elaboration” document, should be clear:

For each outcome, study results should be reported as a summary of the outcome in each group (for example, the number of participants with or without the event and the denominators, or the mean and standard deviation of measurements), together with the contrast between the groups, known as the effect size. For binary outcomes, the effect size could be the risk ratio (relative risk), odds ratio, or risk difference; for survival time data, it could be the hazard ratio or difference in median survival time; and for continuous data, it is usually the difference in means. Confidence intervals should be presented for the contrast between groups. … For binary outcomes, presentation of both absolute and relative effect sizes is recommended.[3]
In other words, relative measures (along with estimates of uncertainty, usually confidence intervals) are necessary—but not sufficient—to inform the reader of a study’s results. The CONSORT authors provided two tables from previously reported research as helpful examples; for illustration, I show an adapted version of just the first row of each table below:

Table 1. Example of Reporting Binary Outcomes
 
Number (%)
 
Endpoint
Etanercept (n=30)
Placebo (n=30)
Risk Difference
(95% CI)
Achieved PsARC at 12 weeks
26 (87)
7 (23)
63% (44 to 83)

CI=confidence interval; PsARC=psoriatic arthritis response criteria.[3]

Table 2. Example of Reporting Continuous Outcomes
 
Exercise Therapy (n=65)
Control (n=66)
 
 
Baseline Mean [SD]
12 Months Mean [SD]
Baseline Mean [SD]
12 Months Mean [SD]
Adjusted Difference (95% CI) at 12 Months
Function score (0-100)
64.4 (13.9)
83.2 (14.8)
65.9 (15.2)
79.8 (17.5)
4.52 (-0.73-9.76)

CI=confidence interval; SD=standard deviation.[3]

Note also that in the second example shown, baseline (pre-intervention) as well as follow-up values are shown to enable the reader to assess the clinical significance of the change amounts in light of the group characteristics prior to the intervention.
The practice of reporting outcomes measured at baseline is recommended by CONSORT “so that readers can assess how similar [the study groups] were” but is unfortunately not always followed even in observational (nonrandomized cohort) studies of interventions, where baseline comparability of the study groups is a critically important issue.[4] For example, observational assessments of therapy outcomes for employer groups that implemented step therapy programs, compared with groups that had no step therapy, have failed to report baseline values on even basic key outcome measures including utilization of the target drug classes and health care costs.[5]

Practical Take-Away Points: Insist on Absolutes. Absolutely.
Without information about both the absolute and relative effects of study variables of interest, it is impossible to determine whether results represent practically/clinically meaningful outcomes or statistical artifact, often due to the enormous sample sizes that are commonplace in health care databases today. (With a sufficiently large number of study subjects, even completely meaningless changes can be statistically significant). The most informative reports indicate baseline values, follow-up values, and absolute change amounts (follow-up minus baseline), in addition to measures of relative difference (e.g., odds ratios) and uncertainty (e.g., confidence intervals).

So if the report of an intervention study with an observational design fails to provide baseline characteristics of the study subjects, including baseline values of the outcome measures, or if it fails to report absolute post-intervention change amounts, its worth is limited. Without this information, there is no way to determine the comparability of the groups prior to the intervention or to get a sense of the practical/clinical effect of the intervention on the outcome.
If a randomized study report fails to provide baseline values on the outcome measures, the report is less informative than it could or should be; however, the problem is usually not a fatal flaw because the randomization process should produce comparable groups. A possible exception is block randomization (randomization of groups instead of individual subjects, such as randomizing all patients treated by a particular physician instead of randomizing individual patients), because the blocks may differ in ways that affect response to the intervention.
And don’t be shy. If you don’t see the information you need in a study report, remember that journals provide contact information for the first author for a good reason—so that you can write to him or her if you have a question. It’s appropriate to ask the author to provide missing information and to ask follow-up questions (nicely) if you have them.

But in a broader sense, a good general rule is this: the less the investigators conformed to reporting guidelines, the more cautious you should be about the validity of the study findings. For that reason, if you have some research training and use research results in your work, it is a good idea to read through the CONSORT or STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) Explanation and Elaboration documents.[3,6] Either will provide a good general sense of the purpose and spirit of reporting guidelines, and knowing what to expect from a high-quality research report will prove invaluable.

[1] Fairman KA, Rucker ML. Fractal mathematics in managed care? How a simple and revealing analysis could improve the forecasting and management of medical costs and events. J Manag Care Pharm. 2009;15(4):351-358.

[2] Calculation note: odds=probability/(1–probability)—in other words, the odds of an event are defined as the probability divided by the probability of the alternative. The odds ratio for A versus B=odds[A]÷odds[B]. For the sake of providing a simplified example, the results shown in this posting are slightly affected by rounding error.


[4] Des Jarlais DC, Lyles C, Crepaz N; TREND Group. Improving the quality of nonrandomized evaluations of behavioral and public health interventions: the TREND statement. Am J Public Health. 2004;94(3):361-366.

[5] Motheral BR. Pharmaceutical step-therapy interventions: a critical review of the literature. J Manag Care Pharm. 2011;17(2):143-155.

[6] Vandenbroucke JP, von Elm E, Altman DG, et al. Strengthening the Reporting of Observational Studies in Epidemiology (STROBE): explanation and elaboration. PLoS Med. 2007;4(10):e296.