Saturday, August 10, 2013

BD, ICD, and GIGO:
Why “Big Data” May Be Less than Meets the Eye

With both funding and pressure piling onto the business of comparative effectiveness research, the old mantra that you cannot improve what you do not measure has seemingly never been more critically important in health care than it is today. Many would argue that measuring health care outcomes has also never been easier.

And, in a sense, they’d be right. Health care researchers today have sophisticated equipment with supercomputing capabilities that many of us couldn’t even dream of ten or fifteen years ago. We also have access to tens of millions of data points—medical and pharmacy claims, mortality records, survey data, and even detailed genetic data, depending on the organization in which we are working. So much information, available so conveniently, at such a low cost, and analyzable at such speed. It’s enough to make a policy analyst positively giddy. Perhaps it is not surprising then, that one commentator, a Health Policy and Life Sciences Group manager at Intel, assessed Big Data’s potential to “revolutionize health care” in this way:

Big Data provides us an opportunity to transition to a personal care system. Rather than making assumptions based on what has worked for other people, this personal view would allow us to take data about a patient’s genomes, medical history and behaviors to construct a virtual model that would help predict which treatments will be most effective and customize them to an individual — improving quality of life for the patient and saving the delivery system money.
With all this excitement about the potential of large datasets to unlock the secrets of greater longevity at lower cost, it’s easy to forget a crucial pitfall encountered by researchers who use them unawares: those beautifully packaged data were collected by human beings. And, because many of those human beings did not have research on their mind when they collected and recorded the data, they may have been motivated to treat our valuable information in ways that we did not expect or want. One example, which I discussed in Chapter 7 of Health Care Research Done Right, is billing (claims) data. Because the main purpose of billing data is to generate payment for health care services, claims coding is subject to “upcoding” and deliberate miscoding to enhance reimbursement.

This example is known to many claims database researchers. However, not all similar threats to study validity are as widely recognized or publicized, despite their potential effect on the well-being of patients whose plans or health care providers forget that real-world data collection may affect ideal-world research in unexpected ways. Several examples have been highlighted in recent press articles. I’ll talk about two in this post.
First is updated information about the long-awaited transition to ICD-10, which will increase the number of available diagnosis codes from about 14,000 to more than 68,000, the number of procedure codes from about 4,000 to more than 72,000, and the number of pages in the American Family Practice Association “superbill” (a standard form intended to list most of the diagnosis codes encountered in a typical practice, used for the physician’s convenience) from 2 to a whopping 9 pages. If used properly, the coding system has the potential to improve the accuracy of electronic diagnostic record-keeping, reimbursement, and, ultimately, quality of care, because—in theory—physicians who are given more accurate feedback will have greater incentive to use evidence-based procedures and treatment protocols.

But I said if used properly, and that big “if” is looking iffier all the time. The deadline for compliance with ICD-10 coding, originally scheduled by the Centers for Medicare & Medicaid Services (CMS) for October 2011, has been pushed back several times and is now slated for October 2014. Except that a survey of “providers, payers and health information technology vendors” conducted in February 2013 found that about one-half of participating vendors reported less than 50% progress toward ICD-10 readiness, and more than 40% did not know when they would begin detailed steps toward final implementation, with about one-quarter reporting being nearly finished.

And those are just the data processing problems. Industry insiders are reporting that experienced diagnosis coders are choosing to retire rather than learn the new system, leaving an unknown proportion of the implementation of ICD-10 in the hands of newbies. How long it will take for coders to become familiar with the new system is unknown. Which will bring researchers to a critically important question: when analyzing claims data coded with ICD-10, how accurate are the diagnoses?

A second example involves a completely different data source but a similar problem. An anonymous, Internet-based survey of New York City hospital residents found that 49% had knowingly reported cause of death inaccurately when completing a death certificate.[1] Of residents who had completed at least eleven death certificates in the previous three years, nearly six in ten reported deliberate inaccuracy. About three-quarters said that the computer system “would not accept the correct cause” of death, 41% said that they were told to “put something else” by the hospital admitting office, and 31% said that the medical examiner told them to report the diagnosis incorrectly. Among the more common actual causes of death associated with deliberately inaccurate reporting was septic shock, management of which is a quality-of-care indicator.[2] Noting that death certificates “contain critical information for epidemiology, public health research, disease surveillance, and community health programs,” the researchers noted that the routine reporting of inaccurate causes of death “may have lasting effects on the public health priorities of the community.”

All of which should give us pause as we consider the current level of enthusiasm for “big data.” Are automated data a silver bullet for all that ails the American health care system, or are they GIGO (garbage in, garbage out)? Savvy researchers should recognize that either possibility exists and should know how to investigate the data prior to using them.

[1] Wexelman BA, Eden E, Rose KM. Survey of New York City resident physicians on cause-of-death reporting, 2010. Prev Chronic Dis. 2013;10:E76.

[2] NQF #0500 Severe Sepsis and Septic Shock: Management Bundle, Last Updated Date: Oct 05, 2012.

No comments:

Post a Comment