Wednesday, April 10, 2013


Percentages Are Funny:
Why Absolutes are Absolute in Research Reporting

A member of our family recently picked up a prescription at our local pharmacy, finding that our out-of-pocket cost for the medication had increased by 367% over a six-month period of time. Assuming that the drug is medically necessary and appropriate for this patient, which of the following is an accurate assessment of the situation?

(a) Sure sign of price gouging by pharmaceutical manufacturers—we need cost controls, and we need ‘em soon
(b) Sure sign of the need for copayment relief—it is “penny wise and pound foolish” to discourage members from purchasing appropriate medication
(c) Means nothing at all

As the title of this posting suggests, the correct answer is (c). The out-of-pocket cost change in question was a mere $2.46—from $0.67 to $3.13—for a one-month supply, about eight cents per day. That’s not enough to merit even a passing glance at the price, let alone “cost-related nonadherence.” (Okay, we noticed, but I think that’s just because my husband and I are a little . . . well, geeky.)[1]

This simple example provides a good illustration of the reason that research-reporting guidelines recommend presentation of absolute numbers—not just relative measures, such as hazard ratios, odds ratios, or percentages—in describing quantitative findings. Percentages (and other relative measures) are funny. They show us how study groups or time periods compare with one another, in relative terms, but they tell us little or nothing about what those differences mean in practical terms.

For example, a mortality odds ratio of 3.23 for Drug A, with Drug B as the reference category, could mean that a patient has a 52% probability of death using Drug A compared with a 25% probability using Drug B—at an additional 27,000 deaths per 100,000 treated patients, clearly a risk worth paying attention to. Or, the same odds ratio could mean that the probability of death is 0.00004% with Drug B and 0.000129% with Drug A—1.29 per million, the approximate probability of getting struck by lightning in any given year. (If you are not familiar with these calculations, see the note below for an explanation.)[2]

In this context, the rationale for the following CONSORT (CONsolidated Standards Of Reporting Trials) guidance, as described in its “explanation and elaboration” document, should be clear:

For each outcome, study results should be reported as a summary of the outcome in each group (for example, the number of participants with or without the event and the denominators, or the mean and standard deviation of measurements), together with the contrast between the groups, known as the effect size. For binary outcomes, the effect size could be the risk ratio (relative risk), odds ratio, or risk difference; for survival time data, it could be the hazard ratio or difference in median survival time; and for continuous data, it is usually the difference in means. Confidence intervals should be presented for the contrast between groups. … For binary outcomes, presentation of both absolute and relative effect sizes is recommended.[3]
In other words, relative measures (along with estimates of uncertainty, usually confidence intervals) are necessary—but not sufficient—to inform the reader of a study’s results. The CONSORT authors provided two tables from previously reported research as helpful examples; for illustration, I show an adapted version of just the first row of each table below:

Table 1. Example of Reporting Binary Outcomes
 
Number (%)
 
Endpoint
Etanercept (n=30)
Placebo (n=30)
Risk Difference
(95% CI)
Achieved PsARC at 12 weeks
26 (87)
7 (23)
63% (44 to 83)

CI=confidence interval; PsARC=psoriatic arthritis response criteria.[3]

Table 2. Example of Reporting Continuous Outcomes
 
Exercise Therapy (n=65)
Control (n=66)
 
 
Baseline Mean [SD]
12 Months Mean [SD]
Baseline Mean [SD]
12 Months Mean [SD]
Adjusted Difference (95% CI) at 12 Months
Function score (0-100)
64.4 (13.9)
83.2 (14.8)
65.9 (15.2)
79.8 (17.5)
4.52 (-0.73-9.76)

CI=confidence interval; SD=standard deviation.[3]

Note also that in the second example shown, baseline (pre-intervention) as well as follow-up values are shown to enable the reader to assess the clinical significance of the change amounts in light of the group characteristics prior to the intervention.
The practice of reporting outcomes measured at baseline is recommended by CONSORT “so that readers can assess how similar [the study groups] were” but is unfortunately not always followed even in observational (nonrandomized cohort) studies of interventions, where baseline comparability of the study groups is a critically important issue.[4] For example, observational assessments of therapy outcomes for employer groups that implemented step therapy programs, compared with groups that had no step therapy, have failed to report baseline values on even basic key outcome measures including utilization of the target drug classes and health care costs.[5]

Practical Take-Away Points: Insist on Absolutes. Absolutely.
Without information about both the absolute and relative effects of study variables of interest, it is impossible to determine whether results represent practically/clinically meaningful outcomes or statistical artifact, often due to the enormous sample sizes that are commonplace in health care databases today. (With a sufficiently large number of study subjects, even completely meaningless changes can be statistically significant). The most informative reports indicate baseline values, follow-up values, and absolute change amounts (follow-up minus baseline), in addition to measures of relative difference (e.g., odds ratios) and uncertainty (e.g., confidence intervals).

So if the report of an intervention study with an observational design fails to provide baseline characteristics of the study subjects, including baseline values of the outcome measures, or if it fails to report absolute post-intervention change amounts, its worth is limited. Without this information, there is no way to determine the comparability of the groups prior to the intervention or to get a sense of the practical/clinical effect of the intervention on the outcome.
If a randomized study report fails to provide baseline values on the outcome measures, the report is less informative than it could or should be; however, the problem is usually not a fatal flaw because the randomization process should produce comparable groups. A possible exception is block randomization (randomization of groups instead of individual subjects, such as randomizing all patients treated by a particular physician instead of randomizing individual patients), because the blocks may differ in ways that affect response to the intervention.
And don’t be shy. If you don’t see the information you need in a study report, remember that journals provide contact information for the first author for a good reason—so that you can write to him or her if you have a question. It’s appropriate to ask the author to provide missing information and to ask follow-up questions (nicely) if you have them.

But in a broader sense, a good general rule is this: the less the investigators conformed to reporting guidelines, the more cautious you should be about the validity of the study findings. For that reason, if you have some research training and use research results in your work, it is a good idea to read through the CONSORT or STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) Explanation and Elaboration documents.[3,6] Either will provide a good general sense of the purpose and spirit of reporting guidelines, and knowing what to expect from a high-quality research report will prove invaluable.

[1] Fairman KA, Rucker ML. Fractal mathematics in managed care? How a simple and revealing analysis could improve the forecasting and management of medical costs and events. J Manag Care Pharm. 2009;15(4):351-358.

[2] Calculation note: odds=probability/(1–probability)—in other words, the odds of an event are defined as the probability divided by the probability of the alternative. The odds ratio for A versus B=odds[A]÷odds[B]. For the sake of providing a simplified example, the results shown in this posting are slightly affected by rounding error.


[4] Des Jarlais DC, Lyles C, Crepaz N; TREND Group. Improving the quality of nonrandomized evaluations of behavioral and public health interventions: the TREND statement. Am J Public Health. 2004;94(3):361-366.

[5] Motheral BR. Pharmaceutical step-therapy interventions: a critical review of the literature. J Manag Care Pharm. 2011;17(2):143-155.

[6] Vandenbroucke JP, von Elm E, Altman DG, et al. Strengthening the Reporting of Observational Studies in Epidemiology (STROBE): explanation and elaboration. PLoS Med. 2007;4(10):e296.

No comments:

Post a Comment