Case Studies
Case studies showcase real-world applications of policies and innovations, offering insights into improving outcomes and efficiency. They enhance problem-solving skills, highlight best practices, and engage through storytelling for education and transparency.
Scaling Transplant Evaluation Without Sacrificing Oversight: An AI-Assisted Workflow for Timely, Patient-Centered Transplant Evaluation
The Challenge
Transplant programs are under increasing pressure to expand evaluation and waitlisting capacity while operating within fixed staffing, regulatory, and financial constraints. Referral volumes continue to grow, expectations around timeliness and access are rising, and programs are asked to do more with the same or fewer resources.
A persistent operational bottleneck across transplant evaluation workflows is the manual identification and confirmation of required clinical information from free-text documentation generated outside the transplant center. Radiology reports, cardiology studies, cancer screening results, and other prerequisite testing are frequently performed at external facilities and arrive as scanned documents or narrative reports in the electronic health record (EHR). Transplant coordinators must search across multiple tabs, media sections, and reports to determine whether required information exists and whether it satisfies evaluation requirements. These administrative delays are experienced directly by patients as prolonged evaluation timelines, uncertainty, and delayed progression toward waitlisting, even when required testing has been completed.
Highly trained coordinators spend substantial portions of their day locating information that already exists, rather than advancing patients through evaluation or coordinating care across teams. As evaluation volumes increase, these document-heavy processes scale poorly. Delays in locating required findings slow evaluation completion, extend time to waitlisting, and increase the risk of patient drop-off during the evaluation process. When progress appears stalled due to missing or hard-to-locate documentation, patients may disengage or abandon evaluation, not because of clinical ineligibility, but because of administrative friction. Over time, the cumulative burden of manual document search contributes to burnout, turnover, and limits program growth.
The underlying challenge is not a lack of clinical expertise, but the mismatch between the complexity of transplant workflows and the fragmented, unstructured nature of incoming clinical data. The growing volume and dispersion of documentation make it difficult to maintain consistency, speed, and throughput when evaluation workflows depend heavily on manual document search.
To address this challenge, we evaluated an AI-assisted workflow designed to support transplant evaluation by reliably identifying whether required, predefined transplant-relevant clinical information is present within free-text clinical documents. The intent was not to automate clinical judgment or decision-making, but to reduce the burden of document search and routine information abstraction so coordinators can focus on patient communication, care coordination, and timely advancement through evaluation.By reducing delays unrelated to clinical readiness, this approach supports faster, more-transparent evaluations while preserving patient-centered care as volumes grow.
The Approach
We implemented an AI-assisted workflow to support transplant coordinators by identifying required clinical information within free-text documents. The workflow was applied to 10,000+ transplant-related documents across three transplant evaluation document types: mammography, abdominal CT, and chest X-ray.
The task was to determine the presence or absence of a predefined set of 27 transplant-relevant clinical variables within each document. Findings were limited to explicit statements in the source text, without inference or clinical judgment. When a variable was identified as present, the workflow assigned a categorical value from predefined options shared across reviewers.
For evaluation, a randomly selected subset of 150 documents (50 per document type) yielded 1,350 variable-level evaluations. Of these, 659 contained an explicit statement about a predefined clinical variable, while the remainder did not. Three label sources were generated for each document-variable pair: AI predictions, clinical annotations, and adjudicated ground truth (AGT).
Clinical annotation was performed by medically-trained annotators, including post-clinical medical students and nurses. Each document was reviewed by a single annotator using standardized definitions and guidelines provided by the study team. Annotators assessed the presence or absence of predefined variables and assigned categorical values when present, without inference beyond explicit text.
AGT labels were established through independent clinician review. Concordant AI and clinical annotation labels were accepted without review, while discrepancies were resolved by two clinicians through source text review.
Detection Performance (Presence-Level Agreement): Detection performance assessed whether required transplant-relevant information was correctly identified as present or absent within each document. AI predictions and clinical annotations were each compared to adjudicated ground truth (AGT). Agreement between the AI or clinical annotation and AGT that a finding was present was considered a true positive, while agreement that a finding was absent was considered a true negative. Sensitivity was emphasized due to the operational risk of missing completed testing, while false positives were interpreted in the context of screening workflows where over-flagging is acceptable.
Conditional Extraction Accuracy: Extraction accuracy was evaluated only among cases where a clinical variable was correctly identified as present. For these cases, extracted categorical values from the AI system and from clinical annotations were compared to AGT to assess extraction accuracy at the document and variable level. This conditional analysis mirrors real-world coordinator workflows, in which relevant information must first be located before its content is verified, and avoids conflating extraction errors with failures to detect relevant documents.
The Results
Detection Performance of AI (Presence-Level Agreement):
The AI-assisted workflow demonstrated high sensitivity for identifying whether required transplant-relevant clinical information was present within free-text documents. Sensitivity ranged from 0.93–0.99 across document types, with an overall false-negative rate of 3.9% (26 of 659 cases with explicit statements), indicating that required information was rarely missed. This level of performance aligns with safety-oriented screening objectives, where missed testing can delay waitlisting.
Specificity ranged from 0.66–1.00 across document types and reflected conservative over-flagging of ambiguous, historical, or negated statements. False positives were broadly distributed across clinical variables, consistent with screening workflows that prioritize completeness and downstream human review.
Detection Performance of Clinical Annotation Relative to Adjudicated Ground Truth
Across document types, clinical annotation demonstrated detection performance broadly comparable to AI-assisted detection, with greater variability across document types. Sensitivity ranged from 0.50–1.00 and specificity from 0.19–1.00 across document types. As with AI-assisted detection, lower specificity reflected conservative over-flagging of ambiguous, historical, or negated statements.
It is important to note that annotations in this study were performed by medically-trained general clinicians rather than transplant coordinators. Given their domain-specific familiarity with transplant evaluation workflows, coordinator performance may reasonably be expected to differ and represents an area for future study.
Conditional Extraction Accuracy (Usability Once Information Is Identified)
Extraction accuracy was evaluated only among cases in which a clinical variable was correctly identified as present. For these true-positive cases, extracted categorical values from both the AI system and clinical annotation were compared to AGT to assess correctness at the document and variable level.
Across document types, AI-extracted values matched AGT in 80.6% of true-positive cases, compared with 69.3% for clinical annotation. This conditional analysis reflects coordinator workflows, in which locating relevant information precedes interpretation. Evaluating extraction accuracy in this restricted context avoids conflating errors in categorization with failures to identify relevant documents.
All AI-extracted findings included sentence-level traceability to the source document, enabling rapid human verification without re-reading full reports. In practice, this supports a shift from manual data entry toward a verification-centered workflow, in which human expertise is applied selectively and efficiently while maintaining clinical oversight.
Taken together, these results indicate that AI-assisted workflows can reliably support both identification and extraction of required transplant-relevant information at a level comparable to routine clinical annotation, while preserving clinician verification.
Insights & Lessons Learned
The findings demonstrate that AI-assisted detection reliably identifies the presence of required transplant-relevant information within free-text documents, with a low rate of missed findings. Additionally, AI results were comparable to manual clinical annotation and more consistent across document types for both identifying the presence of required clinical information and extracting the corresponding transplant-relevant findings once identified.
Importantly, this consistency does not replace clinical oversight. Instead, it supports a workflow in which AI performs first-pass screening and “below-license” routine abstraction, while clinicians retain responsibility for verification, interpretation, decision-making, and patient care.
A key contribution of this work is the explicit separation of detection from extraction accuracy. Detection failures pose the greatest operational risk because missed information can delay evaluation progression or waitlisting. Extraction accuracy, evaluated conditionally among correctly identified findings, reflects usability rather than diagnostic correctness. The inclusion of sentence-level traceability for every AI-generated finding is critical in this context, as it enables clinicians to rapidly verify, confirm, or correct extracted information without re-reading entire documents. This design preserves transparency and ensures that human judgment remains central to patient care.
This study highlights a central challenge in transplant evaluation workflows: delays and inefficiencies are driven less by gaps in clinical expertise than by the manual execution required to locate and confirm information that already exists within fragmented, unstructured documentation. As evaluation volumes increase under fixed staffing and regulatory constraints, these execution bottlenecks increasingly limit both operational performance and patient experience.
By shifting transplant coordinators from manual document search and repetitive data entry toward verification, communication, and care coordination, this AI-assisted workflow supports top-of-license practice while maintaining clinician-in-the-loop safeguards. For patients, this approach has the potential to reduce non-clinical delays, improve transparency during evaluation, and decrease drop-off driven by administrative friction rather than medical eligibility.
Together, these results suggest that AI-assisted workflows, when explicitly designed to support human verification and oversight, can help transplant programs manage growing demand more effectively while preserving patient-centered, clinician-led care.