Defining emerging conditions via health records

April 5th, 2024PLOS ONE Paper

Share:

Summary

The Problem: Electronic health records document diagnoses continuously and at scale, but they can only be searched for conditions that have already been named, characterised, and assigned a diagnosis code; for emerging conditions, establishing that definition takes years. This lag makes it difficult to measure how common a new condition is, who it affects, and whether its incidence is changing, precisely when that information is most valuable. This study demonstrates how we defined an emerging condition directly from health record data, using post-COVID conditions (commonly called “long COVID”) as the test case. While our focus was on long COVID, the method applies to any condition that appears in clinical records before a consensus definition exists.

Our Approach: We developed a three-step process to derive a condition definition from patient records alone.

  1. Establishing a Baseline: We matched each affected patient to comparable unaffected patients, so that symptom rates could be interpreted against how often those symptoms occur anyway.

  2. Weighting the Diagnoses: For every diagnosis category we computed an incidence ratio (its rate in the affected group divided by its rate in the comparison group); categories occurring more often in the affected group were retained, each weighted by that ratio.

  3. Labelling the Patients: For each patient we combined their newly acquired diagnoses with the corresponding weights to produce a single score, which determined whether they were labelled as having the condition.

The Impact: Our work provides a way to measure an emerging condition during the years in which no diagnosis code exists for it. Applied to 588,611 patients with COVID-19, our definition identified 6.6 times as many cases compared to the official diagnosis code. While our focus was on post-COVID conditions, the approach generalises to any condition that arrives in clinical data ahead of its formal definition.




The Problem

Health records can only be searched for conditions that have already been defined and coded, which leaves emerging conditions unmeasurable during the years before a consensus definition is reached.

Electronic health records are increasingly used to study disease incidence in real-world populations, with each new dataset covering more lives and more of the care they receive. However, that use depends on a prior definition: a condition must be named, characterised, and assigned a diagnosis code before it can be retrieved from the record at all. For emerging conditions, establishing that definition often takes years, and the records accumulated in the meantime hold the clinical evidence with no means of identifying it.

Post-COVID conditions illustrate the scale of the problem. The CDC, WHO, and NICE each published a different case definition, disagreeing on the qualifying time window, the constituent symptoms, and the severity threshold; published incidence estimates over the same period ranged from 9% to 52%, a spread that reflects definitional disagreement more than epidemiological variation. The corresponding diagnosis code (ICD-10-CM U09.9) did not become available in the United States until October 1, 2021, approximately eighteen months into the pandemic, leaving the majority of cases unlabelled in the record.

Our Approach

We derived the definition from the records themselves: identifying the diagnoses that follow a condition more often than they follow ordinary care, weighting them by that excess, and scoring each patient against the result.


1. Establishing a valid baseline for comparison: To attribute a symptom to a condition, its rate must be interpreted against how often that symptom occurs in comparable patients. We drew both groups from the same de-identified health record dataset (covering 104 million lives across 760 hospitals and 7,000 clinics), and matched up to three comparison patients to each of 588,611 patients with a confirmed COVID-19 episode, using a propensity score built on sex, race, ethnicity, insurance payor, obesity, smoking history, and comorbidity burden. Comparison patients were required to have presented to the health system in the same month for an unrelated reason, which controls for detection bias arising from differential contact with care (see Table 1).

Characteristic COVID-19 group
n = 588,611
Comparison group
n = 1,286,050
Mean age (years)48.148.3
Male38.4%40.2%
Overweight or obesity68.0%66.4%
Smoking history41.9%43.2%
Comorbidity burden (index)0.90.9
Table 1: Characteristics of the matched groups. Every measured characteristic fell within a standardised mean difference of 10%, the conventional threshold for treating two groups as comparable.

2. Quantifying which diagnoses follow the condition: For each diagnosis category we computed an incidence ratio: its rate in the COVID-19 group divided by its rate in the comparison group. Categories with a ratio above one were retained to form the definition, each carrying its ratio as a weight. The largest imbalances were respiratory and cardiovascular; viral pneumonia occurred approximately 1,360 times as often in the COVID-19 group (see Fig. 1). Constructing the definition this way means its constituent conditions were determined by what clinicians recorded, rather than selected in advance by consensus.

Fig. 1: The ten diagnosis categories most over-represented following COVID-19, shown on a logarithmic scale (the highest sits three orders of magnitude above the tenth). Hover over a bar to see the category name and its exact incidence ratio.

3. Deriving a patient-level label: For each patient we combined the retained diagnoses present in their record with the corresponding weights, producing a single score. Diagnoses documented in the year prior to infection were assigned zero, so that pre-existing conditions could not be counted as new; diagnoses appearing during the acute phase that did not recur after 30 days were likewise excluded, since these describe the infection itself rather than its sequelae. Patients scoring at or above the threshold were labelled as having the condition.

The Impact

Our appraoch identified post-COVID conditions in 6.6 times as many cases as the official diagnosis code.


The definition identified post-COVID conditions in 118,018 of 588,611 COVID-19 patients (20.1%). Incidence was stable across the study period (19.9% for patients diagnosed in 2020 against 20.3% in 2021), which is notable given that most published work over the same period reported declining incidence; it rose substantially with age, from 7.8% in patients aged 0–17 to 33.2% in those aged 65 and older. We then evaluated the definition against the official diagnosis code, the code identified 2.9% of these patients while our definition identified 19.0%. The divergence indicates that the two definitions capture different populations, and that either used alone will undercount; maximal capture requires combining them. While our focus was on post-COVID conditions, none of the three steps depend on the condition itself, requiring only an exposure, an affected population, and a comparable unexposed population. The benchmarks and specification required to reproduce this work are publicly available, providing a basis for defining other conditions that appear in clinical data before a consensus definition exists.

© 2026 Ghamut Corporation