BD, ICD, and GIGO:
Why “Big Data” May Be Less than Meets the
Eye
With both funding and pressure piling onto the business of
comparative effectiveness research, the old mantra that you cannot improve what
you do not measure has seemingly never been more critically important in health
care than it is today. Many would argue that measuring health care outcomes has
also never been easier.
And, in a sense, they’d be right. Health care researchers
today have sophisticated equipment with supercomputing capabilities that many
of us couldn’t even dream of ten or fifteen years ago. We also have access to
tens of millions of data points—medical and pharmacy claims, mortality records,
survey data, and even detailed genetic data, depending on the organization in
which we are working. So much information, available so conveniently, at such a
low cost, and analyzable at such speed. It’s enough to make a policy analyst
positively giddy. Perhaps it is not surprising then, that one commentator, a
Health Policy and Life Sciences Group manager at Intel, assessed Big
Data’s potential to “revolutionize health care” in this way:
Big Data provides us an opportunity to
transition to a personal care system. Rather than making assumptions based on
what has worked for other people, this personal view would allow us to take
data about a patient’s genomes, medical history and behaviors to construct a
virtual model that would help predict which treatments will be most effective
and customize them to an individual — improving quality of life for the patient
and saving the delivery system money.
With all this excitement about
the potential of large datasets to unlock the secrets of greater longevity at
lower cost, it’s easy to forget a crucial pitfall encountered by researchers
who use them unawares: those beautifully packaged data were collected by human
beings. And, because many of those human beings did not have research on their mind when they collected and recorded
the data, they may have been motivated to treat our valuable information in
ways that we did not expect or want. One example, which I discussed in Chapter
7 of Health Care Research Done Right, is billing (claims) data. Because
the main purpose of billing data is to generate payment for health care
services, claims coding is subject to “upcoding” and deliberate miscoding to
enhance reimbursement.
This example is known to many claims
database researchers. However, not all similar threats to study validity are as
widely recognized or publicized, despite their potential effect on the
well-being of patients whose plans or health care providers forget that real-world
data collection may affect ideal-world research in unexpected ways. Several
examples have been highlighted in recent press articles. I’ll talk about two in
this post.
First is updated information about the long-awaited
transition to ICD-10, which will increase the number of available diagnosis
codes from about 14,000 to more than 68,000, the number of procedure codes from
about 4,000 to more than 72,000, and the number of pages in the American Family
Practice Association “superbill” (a standard form intended to list most of the
diagnosis codes encountered in a typical practice, used for the physician’s
convenience) from 2 to a whopping 9 pages. If
used properly, the coding system has the potential to improve the accuracy of
electronic diagnostic record-keeping, reimbursement, and, ultimately, quality
of care, because—in theory—physicians who are given more accurate feedback will
have greater incentive to use evidence-based procedures and treatment
protocols.
But I said if used
properly, and that big “if” is looking iffier all the time. The deadline for
compliance with ICD-10 coding, originally scheduled by the Centers for Medicare
& Medicaid Services (CMS) for October 2011, has been pushed back several
times and is now slated for October 2014. Except that a survey
of “providers, payers and health information technology vendors” conducted
in February 2013 found that about one-half of participating vendors reported
less than 50% progress toward ICD-10 readiness, and more than 40% did not know
when they would begin detailed steps toward final implementation, with about one-quarter reporting being nearly finished.
And those are just the data processing problems. Industry
insiders are reporting that experienced diagnosis coders are choosing to retire
rather than learn the new system, leaving an unknown proportion of the
implementation of ICD-10 in the hands of newbies. How long it will take for
coders to become familiar with the new system is unknown. Which will bring
researchers to a critically important question: when analyzing claims data
coded with ICD-10, how accurate are the diagnoses?
A second example involves a completely different data source
but a similar problem. An anonymous, Internet-based survey of New York City
hospital residents found that 49% had knowingly reported cause of death
inaccurately when completing a death certificate.[1] Of residents who had
completed at least eleven death certificates in the previous three years,
nearly six in ten reported deliberate inaccuracy. About three-quarters said
that the computer system “would not accept the correct cause” of death, 41%
said that they were told to “put something else” by the hospital admitting
office, and 31% said that the medical examiner told them to report the
diagnosis incorrectly. Among the more common actual causes of death associated
with deliberately inaccurate reporting was septic shock, management of which is
a quality-of-care indicator.[2] Noting that death certificates “contain
critical information for epidemiology, public health research, disease
surveillance, and community health programs,” the researchers noted that the
routine reporting of inaccurate causes of death “may have lasting effects on
the public health priorities of the community.”
All of which should give us pause as we consider the current
level of enthusiasm for “big data.” Are automated data a silver bullet for all
that ails the American health care system, or are they GIGO (garbage in,
garbage out)? Savvy researchers should recognize that either possibility exists
and should know how to
investigate the data prior to using them.
[1] Wexelman BA, Eden E, Rose KM. Survey of New York City
resident physicians on cause-of-death reporting, 2010. Prev Chronic Dis. 2013;10:E76.
[2] NQF
#0500 Severe Sepsis and Septic Shock: Management Bundle, Last Updated Date: Oct
05, 2012.
No comments:
Post a Comment