
7 Questions to Ask Before Choosing an AI Patient Screening Solution
Before selecting an AI patient screening tool for clinical trial recruitment, ask these 7 critical questions to avoid costly mismatches and protocol failures.
Recruitment delays affect roughly 80% of clinical trials, and the problem rarely originates in the lab [1]. It begins at the front end: the slow, error-prone process of identifying which patients actually qualify for a given study. A coordinator reviews charts. Another reviews them again. Weeks pass. Eligible patients are missed, ineligible patients are advanced, and screen failure rates climb until the budget starts to buckle.
AI-powered patient screening tools promise to solve this. A growing number of vendors now offer platforms that claim to automate eligibility matching, pull from EHR data, and dramatically reduce the manual burden on site staff. Some of these claims are well-founded. Others outrun the evidence.
The difficulty is that choosing the wrong tool does not just waste money. It introduces data integrity risks, compliance exposure, and the operational chaos of integrating a system that does not actually fit the protocol. Sponsors and CROs who evaluate these platforms only on interface quality or headline enrollment speed metrics tend to regret it.
This article offers seven specific questions to ask any AI patient screening vendor before signing a contract. The questions are grounded in the clinical, regulatory, and operational realities of modern trial recruitment, not in a generic software buyer's guide.
Why AI screening vendor selection matters
Why Getting This Wrong Is Expensive
The numbers on enrollment failure are well-documented. A Tufts Center for the Study of Drug Development (Tufts CSDD) analysis of over 150 Phase II and III trials found that 11% of activated investigative sites fail to enroll a single patient, and 37% under-enroll against their targets [2]. A separate Tufts CSDD report found that approximately 53% of studies required extended enrollment timelines and that about 41% of activated sites did not reach their enrollment goals [3].
Protocol complexity compounds the recruitment burden. Tufts CSDD research published in a 2014 peer-reviewed analysis found that by 2012, the typical Phase III protocol required volunteers to meet 50 eligibility criteria before enrollment, compared to an average of 31 criteria a decade earlier [4]. Each additional criterion narrows the eligible population and extends the time a coordinator must spend reviewing each chart. The relationship between protocol complexity and screening burden has only grown more consequential as therapeutic focus has shifted toward rare diseases and molecularly defined subpopulations.
AI screening tools were designed to address exactly this pressure. But the category spans a wide range of architectures, from simple keyword-matching tools applied against structured EHR fields, to LLM-based systems that parse unstructured clinical notes in real time. Knowing which category a given vendor actually occupies, and whether that architecture fits the demands of a specific protocol, is the first skill a procurement team needs to develop.
The 7-question AI screening evaluation framework
The right AI screening platform should fit the protocol, the site network, the EHR environment, and the regulatory context.
The 7 Questions
1. Does the System Process Unstructured Data, or Only Structured EHR Fields?
This question exposes the most consequential architectural gap in the AI screening market, and most vendors do not volunteer the answer.
Structured EHR data covers discrete fields: ICD-10 diagnosis codes, lab values in numeric fields, medication lists, demographic variables. It is machine-readable and easy to query. The problem is that a large share of the clinical information required to assess trial eligibility does not live there. Prior surgical histories, comorbidity nuances, dosing decisions, tumor characteristics, and adverse event descriptions are recorded in free-text clinical notes, consultation summaries, and discharge narratives.
A 2019 study published in the International Journal of Medical Informatics assessed NLP-based eligibility surveillance across three breast cancer trials at the Medical University of South Carolina [5]. The system extracted eligibility criteria from unstructured EHR clinical notes with sensitivity of up to 95.5%, and achieved an AUC of 89.8% for matching patients to the appropriate trials. Those results came specifically from working with unstructured text. A system restricted to coded EHR fields would have missed the clinical information that makes that performance possible, because the relevant details were buried in clinical notes, not in discrete data fields.
Ask the vendor directly: does your matching algorithm read clinical notes, consultation summaries, and discharge documents, or does it work only from coded fields? A tool that cannot process unstructured text will miss a material fraction of eligible patients, no matter how polished its interface is.
2. How Does the System Handle FHIR and EHR Interoperability at Your Specific Sites?
The standard answer from vendors is that their platform "supports FHIR." What matters operationally is what that actually means at the site level.
FHIR (Fast Healthcare Interoperability Resources) is the current standard for health data exchange, and its use in clinical trial recruitment has been shown to work in practice. A 2022 study published in Studies in Health Technology and Informatics developed a FHIR R4-based recruitment support system in a cardiology department and found that the system correctly identified 52 out of 55 patients enrolled across four active clinical trials [6]. The authors concluded that using FHIR to define eligibility criteria could allow automatic screening across multiple sites from different healthcare providers, a capability that current non-interoperable systems cannot replicate at scale.
The practical question is whether the vendor's FHIR integration handles the real-world inconsistency of how individual hospitals implement FHIR. Different Epic instances, for example, encode the same clinical data differently. A well-built integration layer normalizes incoming data to standard vocabularies, SNOMED CT for conditions, RxNorm for medications, LOINC for lab results, before matching begins. Without normalization, a patient whose statin is coded differently across two sites may be matched differently at those same sites for the same trial.
Ask the vendor: which EHR systems have you deployed live integrations with, not just certified compatibility for? What is your normalization approach for coding inconsistencies? How long does integration take at a typical site, and what data access model do you use (pull vs. push vs. API query)?
The difference between a vendor who has a working FHIR API and one who has live, validated integrations with the specific EHR systems used by your trial sites is not a minor distinction. It determines whether the tool works on day one of site activation.
3. What Is the Evidence Base for the Matching Algorithm's Accuracy?
Performance claims in this market frequently cite sensitivity and specificity figures from internal validation studies. Some of these are methodologically sound. Many are not. The evaluation conditions, the patient population, the trial types, the site mix, and the clinical area all shape the numbers, and a breast cancer pre-screening tool validated in a single academic oncology center does not automatically transfer to a rare-disease protocol at a community site network.
Published literature on LLM-based patient-trial matching offers a useful frame for calibration. A 2026 retrospective, single-center pilot study (the PANCR-AI study) published in JMIR Cancer evaluated three large language models, including GPT-4.5 and Claude 3.7 Sonnet, against a double-blind human oncologist gold standard for patient screening in pancreatic cancer clinical trials [7]. Using de-identified clinical reports rather than live EHR integration, the study found meaningful variation in sensitivity and specificity across models, and the authors concluded that even in favorable conditions, LLM-based screening requires human review of borderline cases. This is consistent with the broader principle that AI screening tools function best as candidate prioritization tools for clinical staff, not as final eligibility adjudicators.
When a vendor presents accuracy figures, ask for the following: the dataset on which validation was performed (size, therapeutic area, site type), whether the validation was prospective or retrospective, what the false-positive rate was at the operating threshold used in practice, and whether any independent third-party validation has been conducted. A vendor with a genuinely well-validated tool will be able to answer each of these without difficulty. One who cannot should raise concern proportional to how central their screening accuracy claim is to the product pitch.
4. Is the Platform Compliant with 21 CFR Part 11, HIPAA, and Applicable Data Privacy Regulations?
Any software system that creates, modifies, maintains, retrieves, or transmits regulated electronic records in connection with an FDA-regulated clinical trial is subject to 21 CFR Part 11 [8]. This regulation requires documented system validation, secure computer-generated time-stamped audit trails, role-based access controls, and electronic signature integrity. It is not optional, and it applies regardless of whether the system is AI-based or conventional software.
The compliance obligation extends beyond 21 CFR Part 11. Under HIPAA, any AI system used by or on behalf of a covered entity or business associate that accesses, processes, or transmits protected health information (PHI) as part of a clinical trial workflow must satisfy obligations under both the Privacy Rule and the Security Rule. The Security Rule requires covered entities and their business associates to implement administrative, physical, and technical safeguards, including role-based access controls, to protect ePHI [9]. Vendors handling PHI on behalf of a covered entity or business associate typically require a HIPAA Business Associate Agreement. For trials enrolling patients in the EU, GDPR may add further obligations where decisions are based solely on automated processing and produce legal or similarly significant effects; in specified cases, Article 22 requires safeguards such as human intervention, the ability to express a point of view, and the ability to contest the decision [13].
ICH E6(R3), finalized January 6, 2025, introduced strengthened requirements for data governance and sponsor oversight that apply directly to computerized systems used in trial conduct [10]. The guideline requires sponsors to implement procedures for data transparency, traceability, and accountability throughout the trial lifecycle, which includes any AI or digital system connected to the recruitment workflow. Since ICH E6(R3) is a non-binding guideline, it does not carry the force of domestic law in the same way that 21 CFR Part 11 does, but regulators in ICH member jurisdictions, including the FDA and EMA, expect sponsors to demonstrate adherence to GCP principles that E6(R3) articulates.
Ask the vendor for their 21 CFR Part 11 compliance documentation, their HIPAA Business Associate Agreement (BAA) template, their SOC 2 Type II audit report, and their computer system validation package. For platforms deployed in multi-regional trials, ask whether they have conducted a Data Protection Impact Assessment and how they handle GDPR's requirements for automated processing of personal data.
A vendor who cannot produce this documentation, or who frames it as "in progress," represents a compliance risk that sits with the sponsor, not with the vendor.
5. How Does the Platform Manage False Positives, and Who Makes the Final Eligibility Call?
This question matters more than most buyers realize at procurement time. It surfaces when a coordinator discovers that a patient the system flagged as eligible has a contraindicated medication in a note the algorithm did not parse, or when a site principal investigator is asked to sign off on screening documentation generated by a tool they do not fully understand.
No AI screening system eliminates false positives. A 2026 study published in the Journal of the American Medical Informatics Association (JAMIA) evaluated an adapted TrialGPT system for real-world clinical trial eligibility screening against an expert-adjudicated gold standard across 149 screened patients [11]. The system achieved sensitivity of 81.8% and a positive predictive value of 75.0%, meaning roughly one in four system-positive flags received a different determination upon expert review. The same study also found that the tool identified more than twice as many truly eligible patients as the existing manual screening process, which illustrates where the tool's value actually lies: expanding candidate identification, not eliminating review. A well-designed system surfaces more of the right patients while making it faster to resolve each flag, it does not hand investigators a final answer.
Where vendors diverge is in how their platform structures the human review step. A well-designed system makes it transparent which criteria the algorithm matched, which it could not assess because the data was absent or ambiguous, and which it flagged for human judgment. A poorly designed one outputs a ranked list of "likely eligible" patients with minimal explanation of the reasoning behind each flag.
From a regulatory standpoint, this matters directly. FDA's January 2026 final guidance on clinical decision support software distinguishes between tools that support clinician judgment and those that in effect replace it [12]. Under the four statutory Non-Device CDS criteria in Section 520(o)(1)(E) of the FD&C Act, a screening tool intended for displaying, analyzing, or printing medical information about a patient, and designed so that the clinician can independently review the basis for any recommendation before acting, may be excluded from device regulation if all applicable criteria are met. A tool that presents eligibility determinations as final rather than as flags for investigator review moves closer to the regulated end of that spectrum and introduces oversight requirements that sponsors need to account for during vendor selection.
Ask the vendor: what information does the platform display alongside each eligibility flag? Can a coordinator see which specific criteria were matched, which were unresolvable, and which were not assessed? Is there a structured workflow for investigator sign-off on pre-screened candidates?
6. Can the Platform Scale Across Your Full Site Network Without Degrading Data Quality?
A tool that works at three academic medical centers with well-maintained Epic instances is not automatically the right tool for a 60-site trial that includes community practices, rural health systems, and sites in multiple regulatory jurisdictions.
Site diversity introduces data quality variability that amplifies the limitations of any AI screening system. Community sites often have less complete EHR records, lower rates of structured data capture for relevant clinical events, and fewer dedicated research coordinators to handle the volume of pre-screened candidates the tool produces. A platform that performs well in high-data-quality environments may generate a higher false-positive rate at data-sparse sites, increasing rather than reducing coordinator workload.
The Meystre et al. study on NLP-based eligibility surveillance, which produced sensitivity of 95.5% in its primary validation, was conducted at an academic medical center with structured research data infrastructure [5]. Many published validation studies for AI screening tools have similar settings, and their results should be interpreted with that context in mind. Performance at a community-based oncology practice or a rural cardiology site cannot be assumed from academic-center performance figures.
Ask the vendor: have you deployed at community sites and data-sparse environments, not only at academic medical centers? What does your validation look like across the specific site types in our network? What is your approach to site-level configuration and threshold tuning when data completeness varies?
7. What Does Implementation, Ongoing Support, and Protocol Amendment Handling Look Like?
The last question is operational rather than scientific, but it determines whether the tool delivers value across the life of a trial rather than just at the kickoff demo.
AI screening platforms require configuration against each protocol's specific inclusion and exclusion criteria. That configuration is non-trivial, and the time required to complete it has direct consequences for site activation timelines. Vendors who minimize the configuration step during the sales process often surface real implementation timelines only after contract signature, at which point sponsors have limited leverage.
Protocol amendments are a particularly important consideration. Trials routinely amend eligibility criteria mid-enrollment, sometimes multiple times. A platform that requires extensive re-configuration with each amendment, or that cannot maintain version-controlled documentation of which criteria were in effect when each candidate was screened, creates audit trail problems that surface later in inspection.
Ask the vendor: what is the typical implementation timeline at a single site, and at a multi-site network? Who owns the configuration work, your team or ours? How does the platform handle protocol amendments mid-enrollment, and what documentation does it produce to support inspection? What is your support model for sites encountering issues during active enrollment?
Vendors with mature implementation processes will have detailed answers. Those without will offer reassurances that should be tested against reference customers operating at similar scale and complexity.
Compliance questions before go-live
Compliance documentation should be available before contract signature, not discovered during trial startup.
The Regulatory Context for 2025 and Beyond
The regulatory environment for AI in clinical trials has clarified substantially in the past 18 months, and sponsors evaluating screening platforms should understand what has shifted.
ICH E6(R3), finalized in January 2025, sets the current expectation for data governance, sponsor oversight, and computerized system validation in GCP-regulated trials [10]. Its provisions on traceability and accountability apply directly to AI screening tools, even though the guideline does not address AI by name.
FDA's final guidance on Clinical Decision Support Software, issued in January 2026, clarifies which categories of decision-support tools fall under medical device regulation and which do not [12]. The distinction turns on whether the software supports clinician judgment, in which case it may be excluded from device regulation if all applicable statutory criteria are met, or in effect replaces it, in which case device regulation typically applies. AI screening tools that present eligibility determinations as final rather than as candidate flags for investigator review may cross into device-regulated territory.
For trials with European arms, GDPR's Article 22 provisions on automated decision-making may apply where AI screening produces decisions that have legal or similarly significant effects on individuals [13]. In those cases, the regulation requires safeguards including human intervention, the ability for the data subject to express a point of view, and the ability to contest the decision.
The practical implication for vendor selection is that compliance documentation is no longer a back-office concern. It is part of the operational due-diligence package, and platforms without it create exposure that flows directly to the sponsor.
How Kitsa Fits Into This Workflow
KScreener, Kitsa's FHIR-based patient pre-screening platform, is designed to address the integration and data-completeness challenges described above. The platform is built on FHIR-native architecture and is designed to surface candidate patients for investigator review rather than render independent eligibility determinations. For sponsors and CROs evaluating AI screening options, Kitsa describes KScreener as a tool intended to reduce coordinator screening burden while preserving the site investigator as the final decision-maker in the eligibility process. For current compliance certifications and documentation, contact Kitsa directly at kitsa.ai.
Key Takeaways
- Recruitment delays affect roughly 80% of clinical trials, and screening inefficiency is a structural contributor to that outcome [1].
- Published NLP studies show that unstructured clinical notes can support high-sensitivity eligibility surveillance, capturing signals that structured-field-only systems may miss [5].
- FHIR compatibility is necessary but not sufficient. Live validated integrations with specific EHR systems at specific site types determine whether the tool works at trial activation [6].
- 21 CFR Part 11, HIPAA BAA documentation where applicable, and ICH E6(R3) data governance requirements are non-negotiable compliance baselines for any AI screening tool in a regulated trial [8] [10].
- False positives are inherent to AI screening systems. The key evaluation criterion is not whether they exist, but whether the platform makes them transparent and structures an appropriate investigator review workflow [11].
- Performance metrics should be assessed at the site level. Accuracy in academic medical centers does not predict performance at data-sparse community sites, where most published validation studies have not been conducted [5].
- Integration timelines, protocol amendment support, and audit trail completeness are operational factors that determine whether a tool helps or hinders trial startup and inspection readiness.
FAQ
What is AI patient screening in clinical trials?+
AI patient screening refers to automated systems that analyze electronic health record data, including both structured fields and unstructured clinical notes, to identify patients who may meet a trial's eligibility criteria. These systems use NLP and machine learning techniques to match patient profiles against protocol-defined inclusion and exclusion criteria, surfacing candidates for review by site coordinators and investigators.
How accurate are AI patient screening tools?+
Accuracy varies significantly by system architecture and clinical area. NLP-based systems that process unstructured EHR data have achieved sensitivity of up to 95.5% and AUC values of approximately 89.8% for patient-trial matching in published studies [5]. A 2026 JAMIA study of an adapted TrialGPT system found sensitivity of 81.8% and a positive predictive value of 75.0% against an expert gold standard, and the system identified more than twice as many eligible patients as manual screening [11]. These results reflect favorable academic center conditions; performance at community sites with less complete EHR data remains less well-characterized in the published literature.
Do AI patient screening tools need to comply with 21 CFR Part 11?+
Yes, when the system creates, modifies, maintains, retrieves, or transmits regulated electronic records in an FDA-regulated trial. 21 CFR Part 11 requires system validation, audit trails, access controls, and electronic signature integrity [8]. Sponsors retain responsibility for ensuring that any vendor system used in their trials satisfies these requirements.
What is the role of FHIR in AI patient screening?+
FHIR (Fast Healthcare Interoperability Resources) is the current standard for exchanging structured health data between systems. FHIR-based screening tools can query EHR data against trial eligibility criteria in a standardized way, supporting screening at multiple sites across different healthcare providers. A 2022 study demonstrated that a FHIR R4-based system correctly identified 52 of 55 trial-enrolled patients across four concurrent cardiology trials [6].
Can AI patient screening tools completely automate eligibility determination?+
No, and platforms framing themselves that way should be scrutinized carefully. FDA's January 2026 final guidance on clinical decision support software distinguishes between tools that inform clinician judgment and those that in effect replace it, with the latter potentially requiring regulatory review as medical device software [12]. Published evidence consistently shows that current AI screening tools function most appropriately as candidate identification systems, with investigators retaining responsibility for final eligibility decisions.
How does ICH E6(R3) affect AI screening tool selection?+
ICH E6(R3), finalized January 2025, introduces strengthened requirements for data governance, sponsor oversight, and traceability of computerized systems used in clinical trials [10]. For AI screening platforms, sponsors must confirm the system is validated, that roles and responsibilities for its use are documented, and that data flowing through the system is traceable and auditable. Platforms with documented 21 CFR Part 11 compliance are better positioned to satisfy these requirements; SOC 2 Type II certification supports vendor security assurance but does not by itself constitute evidence of Part 11 or GCP compliance. Sponsors should request separate validation documentation for each framework.
References
- [1] Olawade DB, Fidelis SC, Marinze S, Egbon E, Osunmakinde A, Osborne A. "Artificial Intelligence in Clinical Trials: A Comprehensive Review of Opportunities, Challenges, and Future Directions." International Journal of Medical Informatics, 2026 Feb; 206:106141. DOI: 10.1016/j.ijmedinf.2025.106141. PubMed PMID: 41075423. https://www.sciencedirect.com/science/article/pii/S1386505625003582
- [2] Getz KA. "New Research from Tufts Center for the Study of Drug Development Characterizes Effectiveness and Variability of Patient Recruitment and Retention Practices." Tufts CSDD Impact Report, January 2013. https://www.fiercebiotech.com/biotech/new-research-from-tufts-center-for-study-of-drug-development-characterizes-effectiveness
- [3] "AI-Driven Clinical Trial Recruitment and Design." Applied Clinical Trials Online, 2025. Citing Tufts CSDD data on extended enrollment timelines. https://www.appliedclinicaltrialsonline.com/view/ai-driven-clinical-trial-recruitment-design
- [4] Getz KA et al. "Improving Protocol Design Feasibility to Drive Drug Development Economics and Performance." Therapeutic Innovation and Regulatory Science, 2014; 48(6):669-677. PMC4053871. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4053871/
- [5] Meystre SM, Heider PM, Kim Y, Aruch DB, Britten CD. "Automatic Trial Eligibility Surveillance Based on Unstructured Clinical Data." International Journal of Medical Informatics, 2019; 129:13-19. DOI: 10.1016/j.ijmedinf.2019.05.018. PubMed PMID: 31445247. https://pubmed.ncbi.nlm.nih.gov/31445247/
- [6] Scherer C, Endres S, Orban M, Kaab S, Massberg S, Winter A, Loebe M. "Implementation of a Clinical Trial Recruitment Support System Based on Fast Healthcare Interoperability Resources (FHIR) in a Cardiology Department." Studies in Health Technology and Informatics, 2022; 294:440-444. DOI: 10.3233/SHTI220497. PubMed PMID: 35612118. https://pubmed.ncbi.nlm.nih.gov/35612118/
- [7] Claessens A et al. "AI-Driven Patient Screening for Clinical Trials in Pancreatic Cancer: The PANCR-AI Pilot Retrospective Comparative Study." JMIR Cancer, 2026; 12:e80268. DOI: 10.2196/80268. PubMed PMID: 41730173. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12928684/
- [8] U.S. Food and Drug Administration. "21 CFR Part 11: Electronic Records; Electronic Signatures." Code of Federal Regulations. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- [9] U.S. Department of Health and Human Services, Office for Civil Rights. "Summary of the HIPAA Security Rule." HHS.gov. https://www.hhs.gov/hipaa/for-professionals/security/laws-regulations/index.html
- [10] International Council for Harmonisation (ICH). "Guideline for Good Clinical Practice E6(R3)." ICH E6(R3) Step 4, finalized January 6, 2025. https://database.ich.org/sites/default/files/ICH_E6%28R3%29_Step4_FinalGuideline_2025_0106.pdf
- [11] Syed M, Hamidi M, Bikkanuri M, Dierschke NA, Katragadda HV, Zozus M, Teixeira AL. "Translating evidence into practice: adapting TrialGPT for real-world clinical trial eligibility screening." Journal of the American Medical Informatics Association, 2026; ocag006. DOI: 10.1093/jamia/ocag006. https://academic.oup.com/jamia/advance-article/doi/10.1093/jamia/ocag006/8460629
- [12] U.S. Food and Drug Administration. "Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff." Final guidance issued January 29, 2026 (supersedes January 6, 2026 version). https://www.fda.gov/media/109618/download
- [13] Regulation (EU) 2016/679 of the European Parliament and of the Council (GDPR). Article 22: Automated individual decision-making, including profiling. Official Journal of the European Union, L 119/1, 4 May 2016. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32016R0679#d1e2818-1-1
Related Articles
Patient Recruitment
AI Patient Recruitment in Clinical Trials
Clinical Operations
Reduce Patient Screening Failures with CTMS Pre-Screening
Clinical Operations
The Hidden Cost of Manual Screening in Clinical Trials
Clinical Operations
When Manual Screening Fails Oncology Trials
AI & Clinical Trials