Spencer Fox Eccles School of Medicine
60 Utilizing Artificial Intelligence For Enhanced Quality Assurance in Emergency Medical Services: Feasibility and Applications
Drew Youngquist
Faculty Mentor: Graham Brant-Zawadzki (Emergency Medicine, University of Utah)
Drew Younquist was an exceptional student and a remarkable young man. I had the privilege of mentoring Drew through multiple research projects as well as his honors thesis, during which I came to know him not only as an inquisitive and disciplined researcher but as a person of deep kindness and quiet confidence. Drew possessed a calm, thoughtful presence that conveyed a wisdom well beyond his years—an “old soul” whose insight and maturity stood out from the moment you met him.
Drew approached his work with both curiosity and care. He asked questions not just to solve problems, but to truly understand, and his approach to the subject reflected a depth of thought that always sought meaning and purpose.
Since his passing, through stories shared by his family and friends, videos of his performances, and notes he wrote in self-reflection, I have learned even more about the breadth of his life—his athleticism, his musical talent, and his tireless work ethic—but what stands out most through every story is his compassion. Drew had an uncommon ability to see and care for others, whether lifelong friends or strangers in passing.
Drew would have been an extraordinary physician, not only because of his intellect, but through his humanity. To know Drew was to be reminded of what it means to engage wholeheartedly with empathy and curiosity. His passing is an immeasurable loss—to his family, to his community, and to the countless lives he would have touched—but his legacy endures in the example he set. It is with both pride and sorrow that we present his work here, a reflection of his sharp mind, kind heart, and enduring spirit. — Graham Brant-Zawadzki
Introduction
Large Language Models (LLMs) are a type of machine learning, a subset of artificial intelligence (AI), capable of learning language and generating human-like text in response to input they receive. These models are trained on vast amounts of data, allowing them to recognize patterns, context, and semantics in language. Over the last decade, LLMs have increased in quantity and quality, showing strong performance in language generation, information retrieval, and, more recently, complex reasoning (Bathaee, 2018). Many fields have begun integrating LLMs, streamlining or completely automating many processes that require language comprehension. In healthcare and especially emergency medicine, where accuracy and efficiency are necessary for complex situations, LLMs can support healthcare professionals in several ways. As the capabilities of LLMs rapidly expand, their utility and impact are only beginning to be realized. A wide variety of AI technologies are already being integrated or tested in clinical applications.
This honors thesis explores the emerging applications of LLMs in Emergency Medical Services (EMS), in three main areas: triage, documentation, and quality assurance. I will devote the most attention to this latter application. Building on a recent publication that I coauthored exploring the potential of LLMs to perform quality assurance, I will propose updated methods that address limitations based on recent innovations. Additionally, I will address the practical challenges of implementing an AI-assisted quality assurance system in a real EMS agency. I believe that proactive consideration of the ethical and effective use of AI-generated clinical assistance in EMS is essential to safely maximize their potential impact.
Background
EMS plays a critical role in public health and safety, carrying many of the same expectations for quality and efficiency as other healthcare sectors, all while working in the high-risk and uncontrolled environment of the prehospital setting. Despite this, EMS faces funding limitations compared to other healthcare sectors, with fees for services capped by government fee structures and operations underwritten by municipal governments. Such funding limitations affect the scope of EMS quality assurance activities that an agency can carry out (Swor, 1992). AI offers an opportunity to increase the scope and efficiency of quality assurance in a tight budgetary environment. EMS contains multiple areas that have potential to be improved by AI and specifically LLMs.
Applications: Triage
Self-triage is a process where an individual researches their symptoms and determines how rapidly they should be evaluated by a physician. Currently, many of the symptom-checker websites and apps that are available often suggest care that is overly urgent (Chenais et al, 2023). This may lead to 911 calls when emergent care and transport is not required, unduly burdening EMS systems. LLMs’ language comprehension abilities allow them to pick up on context and ask specific follow up questions, crucial components of effective triage. Implementing LLM assisted triage could potentially improve the accuracy of online care recommendations, thereby reducing the strain on EMS.
Prehospital triage performed by EMS providers on scene often follows predetermined protocols to assess the severity of a patient’s condition. For example, EMS personnel assess a patient with chest pain using a combination of a history and physical examination, vital signs, and a 12-lead Electrocardiogram (ECG). A patient suspected of having an acute myocardial infarction is then preferentially taken to a hospital that is equipped to provide appropriate and timely care. While more sophisticated than patient performed self-triage, provider-performed triage also has room for improvements in speed and accuracy. A literature review investigating the causes of preventable prehospital deaths revealed that delays in treatment were most commonly identified as a leading cause. (Pfeifer et al, 2019). Improving the speed and accuracy of triage could improve outcomes. For example, AI assisted ECG interpretation is available and can improve the detection of acute myocardial infarction when compared to clinicians (Herman, 2023).
Although triage is performed by healthcare professionals and is informed by protocols, its consistency is impacted by human bias. Research on paramedics’ ability to accurately triage patients using structured guidelines found that their ultimate triage decision is influenced by subjective elements such as judgement, experience, and situational interpretation (Pointer et al, 2001). Although using mental heuristics and experience may be helpful in reducing the time it takes to make triage decisions, these heuristics are not always accurate. This can lead to variation in triage decisions and inconsistencies prioritizing care, impacting patient outcomes.
A study of using LLM assisted triage found an improvement in health care professionals’ clinical decision making and triage consistency (Chenais et al, 2023). Implementing rapid, accurate, and unbiased LLM assisted triage in an EMS system could reduce delays in treatment and improve outcomes.
Applications: Documentation
Documentation is an area of EMS that stands to benefit greatly from LLM innovations. EMS personnel, as the first line of care for many patients, have the dual role of gathering critical information, documenting their encounter, all while performing any lifesaving interventions. These records, referred to as patient care reports (PCRs) or electronic health records (EHRs), contain important clinical information from the prehospital encounter. These reports inform future providers of baselines and trends from EMS’s initial point of contact with a patient. EMS also passes on details regarding the environment the patient was found in, which may provide clues to the nature of illness or mechanism of injury. Additionally, medical history collected enroute to the hospital may later become unavailable if the patient’s condition deteriorates. EMS documentation also plays a role in legal and operational contexts, where thorough records are essential for defending care decisions during lawsuits and for accurate billing. PCRs may also be used for medical research (Anderson, 2023). For these reasons, EMS providers’ role as timely and accurate information gatherers is almost as important as their role transporting and performing lifesaving interventions.
As important as proper documentation is, EMS providers face significant time constraints and challenges. While thorough documentation is important, spending excessive time on documentation can detract from patient care and observation, potentially delaying treatments or leading to an incomplete assessment. In emergent situations this leads critical interventions to be prioritized over comprehensive documentation. While patient care is first priority, postponing documentation can affect the quality of the information and delay the time it takes for subsequent care providers to access PCRs (Angeli, 2019). Patients are not the only ones affected, as excess time spent on documentation can also negatively impact EMS providers.
Physicians encounter similar issues, spending nearly twice as much time completing documentation as they do interacting with patients, leading to dissatisfaction and a decrease in the time available to care for patients (Ball, 2021). Technologies that streamline the documentation process should promote a hands-off approach while maintaining accuracy. This would enable providers to spend more time with their patients, which is why they go into medicine.
A myriad of upcoming LLM solutions to these problems are being brought to market. Voice recognition technology, combined with the natural language processing of LLMs, offers solutions to many of the challenges of clinical documentation. Ambient listening software that can monitor a patient-provider conversation, parse out the medically relevant information, summarize it, and sort it into the appropriate fields within a PCR are being implemented currently, with great benefit to providers and patients alike. These transcription tools harness the effectiveness of a human medical scribe at a reduced cost and heightened scalability (Biswas et al, 2024).
Companies like DeepScribe offer some such technologies, claiming to reduce documentation time up to 75%, improve billing cycles, and mitigate burnout by reducing chart closure time (DeepScribe, n.d.). How this technology would be implemented in an EMS setting is yet to be seen. Additional challenges for a transcription tool in EMS include background noise, multiple speakers such as bystanders, or multiple providers/agencies caring for a patient at once.
Applications: Quality Assurance
Another major function of PCRs is for quality assurance (QA), a process where agencies systematically review performance, identify areas for improvement, and ensure that the best possible care is being delivered consistently. This type of retrospective analysis performed during the QA process is vital for providers and agencies to achieve continual improvement.
The major challenge of QA of PCRs is that it is a specialized and labor-intensive activity, requiring funded positions for specifically trained personnel in the EMS agency. Each PCR must be assessed for accuracy, completeness, and compliance with protocols and standard of care. Manually reviewing PCRs is time intensive, with each report requiring attention to detail and an unbiased evaluation. EMS agencies in major urban areas may run over 100,000 calls each year, where the sheer number of PCRs generated each day makes it impractical, if not impossible, for all reports to be reviewed by human personnel. Thus, only a small sample of reports can be checked, offering insights that may or may not be representative of the quality of care provided throughout an agency. Additionally, human QA reviewers, despite their expertise, may bring subjective perspectives into the process.
QA also serves as an opportunity for EMS providers to receive administrative feedback on their PCRs, helping them improve their clinical performance and documentation as a result. The value of performance feedback in EMS is well established; however, many providers report dissatisfaction with current systems and a desire for more robust and frequent feedback methods (Morrison, 2017). In QA, only a small sample of PCRs are reviewed, meaning that most providers receive administrative feedback on just a handful of reports per year. With the turnaround time of the manual review process, feedback is often delivered long after the original call, making it difficult for providers to recall the context needed to reflect meaningfully on it. These delays diminish the potency of the feedback, limiting its ability to promote learning. As a result, the potential for QA feedback to drive meaningful improvement is often underutilized.
LLM’s language processing capabilities offer an extremely promising avenue for streamlining EMS QA. LLMs can rapidly process and analyze large quantities of data and make comparisons with the most up to date EMS protocol requirements. These algorithms can then identify documentation that is inadequate, demonstrates improper care, is likely erroneous, and requires additional human review. This helps facilitate enhanced quality assurance, minimizing the requirement for human labor and potentially reducing bias (Khinvasara, 2024). These LLM driven insights can pinpoint areas of care and documentation in need of improvement, helping EMS agencies make well informed decisions and EMS providers improve their care and documentation.
I joined a cohort of EMS physicians at the University of Utah investigating the potential application of LLMs to provide QA to two urban EMS agencies. Our research utilized the popular LLM ChatGPT version 3 from the company OpenAI. We set out to determine if it was possible for non-AI experts to prompt an LLM to perform QA of PCRs generated during the care of patients with non-traumatic chest pain and ST-elevation myocardial infarction (STEMI) as the paramedics’ primary impression. The goal was to determine whether the LLM could analyze the documents and determine if paramedic providers followed established protocols for the assessment and care of these patients when compared to human reviewers. For example, we prompted the LLM to tell us if an ECG was performed within a certain amount of time from arriving on scene and whether aspirin was administered. This information was available in the paramedic PCRs we uploaded as PDFs, which contain a written narrative, timecodes, and treatment/assessment sections.
Before uploading these PDFs, we had to manually deidentify protected health information (PHI) from each PCR’s narrative and demographics sections. This step was necessary because the LLMs available at the time processed data on external servers, which would otherwise have posed privacy concerns. Then, we trained ChatGPT to perform quality assurance through a process called iterative prompting. We provided instructions and examples to the model and asked it to perform QA tasks, responding with feedback on each response until we reached the desired output. This process was initially laborious, but incrementally improved the performance of the LLM. Once the LLM was able to consistently respond to the instructions with the desired output and appeared sufficiently trained, we were ready. Finally, we created a timed survey for documenting answers about the key quality metrics for both human and AI review.
For the trial, we compared the performance of expert human reviewers, which served as the criterion standard for comparing automated review (our trained GPT). The expert human reviewers that we tested were board-certified EMS medical directors. After starting the timed survey, they would read the report and answer questions about quality metrics while filling out the survey. When submitted, the timer would end. For the automated review, the process was more complex. The process still required manual copying and pasting of quality assurance questions and uploading the PCR into the chat. We didn’t have a way to automatically populate the survey answers that the LLM responded with into the survey, so that was also done via copy and paste. The total time of this process was recorded for comparison.
To get a sense for how often human raters agreed with one another, we first compared the analysis of the same reports done by two different human reviewers. Their analyses had 91.2% agreement with each other (resulting in a kappa statistic of 0.782). Not perfect, but very strong agreement. The automated review was then compared to human reviews and demonstrated 76.2% agreement (kappa 0.401). For efficiency, each automated report saved on average 4 seconds, which was not statistically significant with all review times taking only 1-2 minutes in total (Brant-Zawadzki, 2025).
Through our pilot study, we showed there is potential for LLMs to assist QA in the future, however at the time of the study there were many limitations with the tools we used. Looking back, with the rate of development of AI technologies, the methods which were drafted in 2023 are already outdated. Many of the original limitations of the study could be overcome with recent advancements to LLMs. By using updated methods for a similar experiment, future research may result in better outcomes.
Limitations Revisited
One of the key limitations of the study was the model’s inability to perform batch processing without additional coding, which restricted the effort to reviewing individual PCRs sequentially rather than all at once. This led to the model experiencing some of the same limitations as human reviewers, with individual review taking more time. An effective implementation of LLMs in EMS QA would allow for the ability to process large batches of data simultaneously, providing analysis of a high volume of PCRs with timely feedback.
Since the publication of the study, OpenAI has introduced API (Application Programming Interface) access to its LLMs. An API is a tool that allows developers to integrate LLMs into their software and run these models automatically using code (OpenAI Platform, 2025). In this context, the API enables multiple deidentified PCRs to be submitted and reviewed simultaneously, with model- generated responses organized into a structured format for analysis. Implementing this kind of batch processing would directly address the major efficiency limitation of the original study. With appropriate coding, developers could create a workflow that processes large volumes of PCRs rapidly and even applies statistical tools to them, identifying trends among the documents. Rather than relying on small samples of PCRs to infer quality issues, such a system could provide real-time, system-wide statistics derived from every PCR, giving a more accurate and actionable picture of EMS performance.
Another limitation of the study was ChatGPT’s difficulty distinguishing between PCRs when tasked with analyzing multiple reports in a single chat log. This is because over time in a chat log, ChatGPT draws on previous information to provide context when answering future requests. While this memory feature is helpful for providing responses that fit within the context of the rest of the conversation, it can cause the model to draw on previous inquiries that are no longer relevant to the newest request. In our study, this led to errors analyzing a specific PCR, with the LLM often answering questions about a previous PCR uploaded to the chat. Batch processing could help also remedy these inaccuracies, processing each PCR on an individual basis.
In order for LLM-assisted QA to be a viable tool for an EMS agency, it needs to reach higher levels of agreement with expert human reviewers, approaching the benchmarks of manual review. Rapid advancements in LLM comprehension capabilities could close this gap in a future experiment. Since GPT-4, which was used in the study, newer models like GPT-4o have been released, exhibiting superior reasoning capabilities, processing speed, and cost-efficiency (OpenAI, 2024). As the performance of ChatGPT improves, so do the capabilities and number of its competitors. An increasing variety of LLMs are being developed and released across various domains. High quality, open-source LLMs are appearing with their own set of benefits. Because they can be downloaded and run locally, EMS agencies could process PCRs containing sensitive PHI while remaining HIPAA compliant with proper in-house cybersecurity measures. Additionally, open-source models can be fine- tuned on EMS-specific language and documentation styles, potentially improving their performance.
As model accuracy and customization options continue to improve, future studies may see much stronger agreement between AI-generated reviews and manual assessments—bringing automated QA closer to real-world implementation in EMS systems.
Considerations
Thoughtful implementation of LLM-assisted QA in real EMS agencies must strike a balance between automation and human oversight. One strategy that could balance these is a flagging system, where all PCRs are automatically reviewed and records with protocol deviations, documentation issues, or ambiguity are flagged for human review. This system would remain efficient while surfacing cases that merit closer inspection. Despite the increase in efficiency, it is essential that humans continue performing manual quality assurance of the outputs of such a system. Human Ai-overseers could also periodically randomly sample PCRs reviewed by the LLM, validating or adjusting the feedback. These solutions would help ensure that automated reviews remain accurate and human expertise is involved with difficult cases. Additionally, examination of edge-cases can provide meaningful insights on LLM performance, which can be used to train and fine-tune the model to improve future feedback. By reducing the burden of manual review, EMS medical directors and quality assurance staff can redirect their efforts toward developing targeted training programs based on system-wide trends identified through automated analysis and remediate poor performance.
Along with the basic analysis of protocol adherence and documentation completeness, the generative capabilities of LLM-assisted QA systems introduce several possibilities for delivering meaningful, case-specific feedback to EMS providers. While the ideal model for post-call feedback would integrate EMS electronic health records with hospital systems to provide final patient diagnoses, such access is often limited by siloed healthcare data, interoperability challenges and privacy regulations. As a practical alternative, EMS providers’ own documented primary impression can serve as a basis for tailored feedback, with a trained LLM able to evaluate whether the protocols for assessment and treatment of the suspected condition were followed. More advanced feedback could compare the documented vital signs and physical exam findings to the typical clinical presentation of the suspected condition. The model could then highlight congruencies and discrepancies, offering targeted insights to help providers improve their diagnostic reasoning. While providers may initially approach AI-generated feedback with a degree of skepticism—particularly if the system appears to challenge their clinical judgment—this approach focuses the feedback on evidence-based characteristics of the suspected condition rather than presuming to know the final diagnosis. By reframing each case as an opportunity to revisit the typical presentation, assessment, and vitals associated with common conditions, this system would allow providers to draw their own conclusions about the accuracy of their impression and identify areas for improvement. Rather than simply confirming compliance, this approach transforms the QA process into a case-based educational tool and facilitates ongoing clinical learning.
Debriefing immediately after a clinical encounter is a valuable, evidence-based tool for EMS education and system improvement. Structured post-call debriefing is especially beneficial following significant or high-acuity events. For example, a quality improvement project conducted in a UK emergency department implemented hot debriefing and found increased levels of provider satisfaction in patient care, self-care, decision-making, and teamwork (Sugarman et al., 2021). LLM-generated feedback could function as a required structured debrief, completed daily at the end of each provider’s shift while each encounter remains vivid in memory. This consistent reflection would mimic the learning benefits of hot debriefing, supporting the continuous improvement goals of an effective quality assurance program.
Despite recent technological advancements that make LLM-assisted QA increasingly viable, several logistical challenges remain for implementing this system. Developing an accurate and reliable system will require funding, expertise, and a lengthy training phase that involves human oversight and iterative quality checks. Furthermore, a fully functional system must be tailored to the diversity of EMS calls, with different feedback for trauma, medical, psychiatric, and refusal of treatment or transport scenarios. Another concern is provider trust in AI-generated feedback. Building credibility will depend on transparency in how feedback is generated and clear communication with prehospital providers on the rationale, benefits, and limitations of integrating such a system. These complexities highlight the need for robust infrastructure that can support ongoing human oversight throughout the development and deployment phases.
Discussion
Future research should evaluate current LLMs’ abilities to perform quality assurance tasks, assessing their agreement with human reviewers and their overall efficiency. To improve accuracy and efficiency, these LLMs should be trained on EMS-specific data and possess batch processing capabilities. Additional studies should explore EMS provider perceptions of AI-generated feedback and investigate strategies to increase trust, usability, and engagement with these systems. More research is also needed examining the best practices for delivering generated feedback and maintaining human oversight in these systems.
LLM-assisted QA presents a promising avenue to improve the delivery of prehospital care. With the capabilities of these models to rapidly process and analyze large volumes of PCRs, EMS agencies can expand the scope and educational value of their QA. The resulting increased administrative efficiency would allow QA staff to focus their efforts on training initiatives guided by system wide trends. As recent advancements bring opportunities to overcome previously identified limitations, the potential of this technology only grows. While challenges remain, the integration of LLMs into EMS QA offers meaningful benefits for providers and the systems that support them.
Bibliography
Anderson, M. K. (2023). Documentation and Legal Liability. JEMS. Retrieved October 31, 2024 from https://www.jems.com/documentation/documentation-legal-liability/
Angeli, E. L. (2019). Improving EMS documentation outcomes. ZOLL Data Systems Blog. Retrieved November 2, 2024, from https://www.zolldata.com/blog/improving-ems-documentation-outcomes
Ball, C. G., & McBeth, P. B. (2021). The impact of documentation burden on patient care and surgeon satisfaction. Canadian Journal of Surgery, 64(4), E457-E458. https://doi.org/10.1503/cjs.013921
Bathaee, Y. (2018). The artificial intelligence black box and the failure of intent and causation. Harvard Journal of Law & Technology, 31(2), 889–938.
Biswas, A., & Talukdar, W. (2024). Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation. International Journal of Innovative Science and Research Technology, 9(5), 994-1008. https://doi.org/10.48550/arXiv.2405.18346
Brant-Zawadzki, G., Klapthor, B., Ryba, C., Youngquist, D. C., Burton, B., Palatinus, H., & Youngquist, S. T. (2024). The Performance of ChatGPT-4 and Gemini Ultra 1.0 for Quality Assurance Review in Emergency Medical Services Chest Pain Calls. Prehospital Emergency Care, 1–8. https://doi.org/10.1080/10903127.2024.2376757
Chenais, G., Lagarde, E., & Gil-Jardiné, C. (2023). Artificial Intelligence in emergency medicine: Viewpoint of current applications and foreseeable opportunities and challenges. Journal of Medical Internet Research, 25. https://doi.org/10.2196/40031
Criss, E. A. (March 1993). EMS research: Obstacles of the past, opportunities in the present, models for the future. JEMS, 18(3). Retrieved Nov 18, 2024 from https://www.cpc.mednet.ucla.edu/sites/default/files/pcrf_attached_files/pcrfarticle1.shtml
DeepScribe. (n.d.). Medical scribe: AI-powered medical documentation. DeepScribe. Retrieved October 31, 2024, from https://www.deepscribe.ai/medical-scribe
Herman, R., Meyers, H. P., Smith, S. W., Bertolone, D. T., Leone, A., Bermpeis, K., Viscusi, M. M., Belmonte, M., Demolder, A., Boza, V., Vavrik, B., Kresnakova, V., Iring, A., Martonak, M., Bahyl, J., Kisova, T., Schelfaut, D., Vanderheyden, M., Perl, L., Aslanger, E. K., … Barbato, E. (2023). International evaluation of an artificial intelligence-powered electrocardiogram model detecting acute coronary occlusion myocardial infarction. European heart journal. Digital health, 5(2), 123–133. https://doi.org/10.1093/ehjdh/ztad074
Khinvasara, Tushar and Ness, Stephanie and Shankar, Abhishek (2024) Leveraging AI for Enhanced Quality Assurance in Medical Device Manufacturing. Asian Journal of Research in Computer Science, 17 (6). pp. 13-35. ISSN 2581-8260 http://archive.bionaturalists.in/id/eprint/2355/
Morrison, L., Cassidy, L., Welsford, M., & Chan, T. M. (2017). Clinical Performance Feedback to Paramedics: What They Receive and What They Need. AEM education and training, 1(2), 87–97. https://doi.org/10.1002/aet2.10028
OpenAI Platform. (2025). Openai.com. https://platform.openai.com/docs/quickstart?api-mode=chat
OpenAI. (2024, May 13). Hello GPT-4o. Openai.com. https://openai.com/index/hello-gpt-4o/
Pfeifer, R., Halvachizadeh, S., Schick, S., Sprengel, K., Jensen, K. O., Teuben, M., & others. (2019). Are pre-hospital trauma deaths preventable? A systematic literature review. World Journal of Surgery, 43(10), 2438–2446. https://doi.org/10.1007/s00268-019-05056-1
Pointer, J. E., Levitt, M. A., Young, J. C., Promes, S. B., Messana, B. J., & Adèr, M. E. J. (2001). Can paramedics using guidelines accurately triage patients? Annals of Emergency Medicine, 38(3), 268-273. doi:10.1067/mem.2001.117029
Sugarman, M., Graham, B., Langston, S., Nelmes, P., & Matthews, J. (2021). Implementation of the “TAKE STOCK” Hot Debrief Tool in the ED: a quality improvement project. Emergency Medicine Journal, 38(8), 579–584. https://doi.org/10.1136/emermed-2019- 208830
Swor R. A. (1992). Quality assurance in EMS systems. Emergency medicine clinics of North America, 10(3), 597–610.