Study assesses safety and accuracy in emergency medicine

In a recent study published in Researchers developed and evaluated the accuracy, safety, and utility of emergency medicine (EM) handoff notes generated using a big language model (LLM) to scale back the documentation burden on physicians without compromising patient safety.

The critical role of handovers in healthcare

Handoffs are critical communication points in healthcare and a known source of medical errors. In consequence, quite a few organizations resembling the Joint Commission and the Accreditation Council for Graduate Medical Education (ACGME) have advocated for standardized processes to enhance safety.

Handoffs of EM patients to inpatients (IP) present unique challenges, including medical complexity, time constraints, and diagnostic uncertainty; Nonetheless, they’re still insufficiently standardized and implemented inconsistently. Electronic health record (EHR)-based tools have attempted to beat these limitations. Nonetheless, there continues to be little research into them in emergency situations.

LLMs have emerged as potential solutions to streamline clinical documentation. Nevertheless, concerns about factual inconsistencies require further investigation to make sure safety and reliability in critical workflows.

In regards to the study

The current study was conducted in an 840-bed urban academic quaternary care hospital in Latest York City. EHR data from 1,600 EM patient encounters that resulted in acute hospitalizations between April and September 2023 were analyzed. Because of the implementation of an updated EM-to-IP handover system, only matches after April 2023 were considered.

Retrospective data were used with the waiver of informed consent to make sure minimal risk to patients. Handover notes were created using a mix of a fine-tuned LLM and rules-based heuristics, following standardized reporting guidelines.

The handover note template closely resembled the present manual structure, incorporating rule-based elements resembling laboratory tests and vital signs, in addition to LLM-generated components resembling current illness history and differential diagnoses. Computer science experts and EM physicians curated data to fine-tune the LLM to enhance its quality while excluding race-based attributes to avoid bias.

Two LLMs, Robustly Optimized BiDirectional Encoder Representations from Transformers Approach (RoBERTa) and Large Language Model Meta AI (Llama-2), were used for salient content selection and abstract summarization, respectively. Data processing included heuristic prioritization and saliency modeling to account for the potential limitations of the models.

Researchers evaluated automated metrics resembling Recall-Oriented Understudy for Gisting Evaluation (ROUGE) and Bidirectional Encoder Representations from Transformers Rating (BERTScore), in addition to a novel patient safety-focused framework. A clinical review of fifty handover notes assessed completeness, readability and security to make sure rigorous validation.

Study results

Among the many 1,600 patient cases included within the evaluation, the mean age was 59.8 years with a regular deviation of 18.9 years, and 52% of patients were female. Automated evaluation metrics revealed that summaries produced by the LLM outperformed those written by physicians in several points.

ROUGE-2 scores were significantly higher for LLM-generated summaries than for physician summaries, at 0.322 and 0.088, respectively. Likewise, BERT precision values ​​were higher at 0.859 than for medical summaries at 0.796. In contrast, the source chunking approach for large-scale inconsistency assessment (SCALE) resulted in a worth of 0.691 in comparison with 0.456. These results suggest that LLM-generated summaries had greater lexical similarities, higher accuracy to source notes, and more detailed content than their human-authored counterparts.

In clinical assessments, the standard of summaries written by LLMs was comparable to summaries written by physicians, but barely worse on several dimensions. On a Likert scale of 1 to 5, summaries produced by LLM scored lower in usefulness, completeness, curation, readability, accuracy, and patient safety. Despite these differences, automated summaries were generally considered acceptable for clinical use, with not one of the identified issues considered life-threatening to patient safety.

When assessing worst-case scenarios, clinicians identified potential level two safety risks, which included incompleteness and faulty logic at 8.7% and seven.3%, respectively, for LLM-generated summaries in comparison with physician-authored summaries there have been no associated risks. Hallucinations were rare within the summaries prepared by LLM, with five cases identified all receiving safety rankings between 4 and five, suggesting mild to negligible safety risks. Overall, LLM-written notes had a better error rate of 9.6% than physician-written notes of two%, although these inaccuracies rarely had a big impact on security.

Inter-rater reliability was calculated using intraclass correlation coefficients (ICC). The ICCs showed good agreement between the three expert rankings for completeness, curation, correctness and usefulness at 0.79, 0.70, 0.76 and 0.74, respectively. The readability achieved quite good reliability with an ICC of 0.59.

Conclusions

The present study successfully generated EM-to-IP handover notes using a refined LLM and rule-based approach inside a user-developed template.

Traditional automated assessments were related to superior LLM performance. Nonetheless, manual clinical evaluations found that although most LLM-generated notes achieved promising quality scores between 4 and five, they were generally inferior to physician-authored notes. Identified errors, including incompleteness and faulty logic, occasionally posed a moderate safety risk, with lower than 10% potentially leading to significant problems in comparison with medical notes.

Magazine reference:

Leave a Reply

Your email address will not be published. Required fields are marked *