Contents
Note for publication: This article provides operational and educational context on HIPAA and related compliance frameworks as they apply to AI infrastructure in clinical research. It does not constitute legal advice. HIPAA applicability, business associate status, and de-identification obligations depend on specific factual circumstances and should be reviewed with qualified legal counsel before any compliance determination is made.
Introduction
In early 2024, UnitedHealth Group's Change Healthcare division suffered a ransomware attack that froze pharmacy operations across the United States and cost the company more than $872 million in the first quarter alone, excluding direct breach response expenses [1]. Congressional testimony from the company's CEO suggested the breach may have affected data for roughly a third of all Americans, though the precise scope remains contested [2]. The attack exposed how deeply interconnected the data flows between clinical and administrative healthcare systems have become, and how costly a single infrastructure failure can be.
Healthcare data breaches remain the most expensive of any industry, for the fourteenth consecutive year. According to IBM's 2025 Cost of a Data Breach Report, the global average cost for a healthcare breach was $7.42 million per incident, down from $9.77 million in 2024 but still the highest of any sector [3]. Healthcare breaches also take the longest to contain, averaging 279 days from initial compromise to containment, compared to a global average of 241 days [3].
Clinical research operates inside this environment. And AI is extending its footprint across functions that directly touch participant data: patient matching, protocol generation, eligibility screening, regulatory writing, and monitoring. A 2025 Netskope Threat Labs report found that 88% of healthcare organizations had adopted cloud-based generative AI tools, while 96% were using applications that leverage user data in some form [4]. Many of these deployments have moved faster than the compliance architecture beneath them, including in clinical trial settings where the regulatory stakes extend well beyond data privacy.
This article examines what HIPAA-compliant AI infrastructure actually requires in clinical research, the regulatory layers that apply, the architecture decisions that determine defensibility, and where implementation failures most often occur.
Who HIPAA Actually Covers in a Clinical Trial
Before examining technical requirements, a critical threshold question needs a direct answer: HIPAA does not apply simply because a clinical trial is happening. It applies based on the relationship between the entities handling participant data.
HIPAA's Privacy and Security Rules govern covered entities, which are health plans, healthcare clearinghouses, and healthcare providers that transmit health information in electronic form for standard transactions, and their business associates, which are persons or organizations that perform functions or activities on behalf of a covered entity that involve the creation, receipt, maintenance, or transmission of protected health information (PHI) [5].
Pharmaceutical and device sponsors are typically not covered entities, unless they also operate healthcare services. They receive participant PHI from covered entity sites through one of three mechanisms: a signed research authorization from the participant under 45 CFR §164.508, an IRB- or Privacy Board-approved waiver of authorization under 45 CFR §164.512(i), or a limited data set under a Data Use Agreement (DUA). Note that a limited data set still constitutes PHI under HIPAA, from which only certain direct identifiers have been removed; it is not fully de-identified and continues to require a signed DUA and minimum necessary controls [6],[7]. The AI tools that sponsors deploy to process data received through these pathways must be assessed within that framework. HIPAA obligations attach based on the entity's covered entity or business associate status and the specific permitted disclosure pathway, not based solely on the origin of the data or the clinical trial setting.
CROs occupy a position that requires case-by-case analysis. A CRO is a business associate when it performs functions involving PHI on behalf of a covered entity directly, for example when it manages clinical data for a covered entity investigator site. A CRO working solely for a pharmaceutical sponsor does not automatically become a HIPAA business associate simply by virtue of that relationship, because the sponsor itself is typically not a covered entity. However, if the sponsor received PHI from a covered entity under an authorization or waiver, and shares that PHI with a CRO to perform services, the HIPAA-permitted disclosure conditions governing that data, and related contractual obligations, may constrain how the CRO handles it, and whether a BAA is required warrants careful analysis [5],[8]. A researcher or clinical vendor may also become a business associate if they perform a covered function directly for a covered entity, such as de-identifying PHI or operating a clinical data system that processes identifiable records on the covered entity's behalf [8].
For AI infrastructure decisions, this means the first question is always: does this AI tool process PHI where the organization's role triggers HIPAA coverage, and through what permitted disclosure pathway was that PHI obtained? The answers govern which HIPAA protections apply, to whom, and through what instruments. Where HIPAA does not apply, sponsors may still owe privacy and security duties through clinical trial agreements, informed consent obligations, DUA terms, applicable state privacy statutes, FTC Act enforcement, and FDA and GCP data integrity expectations.
The following table illustrates how entity roles, legal bases, and contractual obligations shift across a common trial data flow. Note that HIPAA BAA requirements attach only where a covered entity or business associate is involved; other obligations such as contractual data protections, research authorizations, DUA terms, and applicable state privacy law may govern the same data even where HIPAA does not require a BAA.
| Party | Role | PHI Legal Basis | Contract Instrument | HIPAA BAA Required? |
|---|---|---|---|---|
| Investigator site (hospital) | Covered entity | N/A (holds PHI as CE) | N/A | N/A |
| Pharmaceutical sponsor | Typically not a CE | Research authorization (45 CFR §164.508) or IRB waiver | Research authorization / CTA | Not automatically; depends on whether sponsor is acting as BA for the CE |
| CRO managing data for sponsor | Analyze per facts | Receives PHI from sponsor | CTA / DPA | Requires formal HIPAA analysis; BA status depends on whether sponsor is itself a CE or BA |
| AI vendor used by CRO | Analyze per facts | Receives identifiable PHI via CRO | BAA or DPA / data processing agreement | Required if CRO is a BA and AI vendor processes ePHI on its behalf; otherwise other contractual/privacy obligations apply |
Each link requires independent analysis. HIPAA is not the only privacy or security regime that may govern sponsor-held trial data; contractual obligations, applicable state law, FTC Act enforcement, Common Rule data protections, and FDA regulations may all apply regardless of HIPAA BA status [5],[6],[8].
Why the Regulatory Environment Has Shifted
The Proposed HIPAA Security Rule Overhaul
On January 6, 2025, HHS OCR published a Notice of Proposed Rulemaking (NPRM) in the Federal Register representing the first substantive update to the HIPAA Security Rule since 2013 [9]. HHS had announced the rulemaking on December 27, 2024. The NPRM, published at 90 FR 800 under docket number RIN 0945-AA22, proposes significant changes that would affect any organization deploying AI to process ePHI.
The core structural change is the proposed elimination of the required/addressable distinction. Under the current rule, several "addressable" specifications can be satisfied with documented equivalent alternatives or risk acceptance. Encryption at rest (§164.312(a)(2)(iv)) is one such addressable specification. The NPRM would change this: if finalized as proposed, all implementation specifications would become required, with very limited exceptions [9]. The NPRM also proposes adding multi-factor authentication as a new explicit required control, a control the current rule does not name by that term. Controls that organizations have historically documented their way around would become mandatory.
The NPRM also introduces explicit obligations for technology asset inventories and network maps showing movement of ePHI through information systems, annual compliance audits, and specific patch management timelines [9]. HHS expects AI software that creates, receives, maintains, or transmits ePHI to be included in these inventories and to be subject to the same controls as any other regulated system.
The finalization timeline, per OCR's regulatory agenda, targets 2026 [10]. Once finalized, covered entities would have 180 days to comply, with business associates receiving an additional 60 days to update agreements, for a total of 240 days. It is essential to note that as of publication, the NPRM has not been finalized. All references to its requirements throughout this article describe proposed changes, not current law. However, organizations building or procuring AI infrastructure now should treat the NPRM's provisions as the likely future baseline.
IBM Breach Data and the AI Detection Gap
IBM's 2025 Cost of a Data Breach Report found that organizations using AI-powered security and automation tools extensively reduced their breach lifecycle by an average of 80 days and saved approximately $1.9 million in breach costs compared to those without such tools [3]. (The 108-day figure cited in earlier literature belongs to IBM's 2023 report; the 2025 figure is 80 days.) The report covered 600 organizations across 16 countries and geographic regions and identified phishing as the leading initial attack vector at 16% of breaches, with supply chain compromise close behind at 15% [3]. For clinical AI platforms that interact with multiple vendor systems, model hosting services, and site data flows, supply chain risk is a distinct exposure, not a subset of general network security.
The same IBM report also found that 63% of breached organizations either lacked an AI governance policy or were still developing one [3]. In clinical research, where AI governance intersects with GCP validation, sponsor oversight requirements, and IRB-approved data use agreements, the absence of a documented AI governance program creates regulatory exposure that extends well beyond data security.
The Regulatory Stack Clinical AI Must Navigate
Clinical trial AI operates at the intersection of multiple overlapping frameworks, each with distinct requirements. Understanding which layer governs which decision prevents both over-compliance (treating every AI tool as a HIPAA-regulated system when it may not be) and under-compliance (treating compliance as a single-checkbox exercise).
HIPAA Privacy and Security Rules
For covered entities and their business associates, the Security Rule requires administrative, physical, and technical safeguards for ePHI. The minimum necessary standard under the Privacy Rule applies to AI tools that access PHI without a research authorization: the tool must use only the PHI strictly necessary for its stated purpose [6]. This standard applies to most disclosures except those made pursuant to a signed individual authorization or for treatment purposes.
De-identification removes data from HIPAA's scope entirely. HIPAA's Privacy Rule at §164.514(a)-(b) specifies two methods: Safe Harbor, which requires removal of 18 defined identifier categories, and Expert Determination, which uses statistical methods to demonstrate that the risk of identifying any individual is very small [11]. Both methods are legally valid; the choice between them has operational consequences for AI model performance, addressed in more detail below.
21 CFR Part 11
FDA-regulated clinical trials conducted under an IND or IDE require that electronic records and electronic signatures meet the requirements of 21 CFR Part 11. Section 11.10(a) specifically mandates validation of computerized systems to ensure accuracy, reliability, consistent intended performance, and the ability to discern invalid or altered records [12]. This requirement applies to AI systems that generate, process, or store clinical data in a regulated trial.
Validation in this context follows established principles requiring documented testing across Installation Qualification (IQ), Operational Qualification (OQ), and Performance Qualification (PQ) phases, consistent with GCP expectations and GAMP guidance for computerized systems [13]. AI systems introduce validation challenges that traditional software did not, including the possibility of behavioral changes between model versions, non-deterministic outputs, and the difficulty of defining pass/fail criteria for generated content. A model update constitutes a system change under GxP principles and should be treated as such in a validation lifecycle, though the appropriate documentation scope depends on the risk category of the system.
NIST AI RMF
The NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0), released on January 26, 2023, is a voluntary cross-sector framework for managing AI risks organized around four core functions: Govern, Map, Measure, and Manage [14]. It is guidance, not regulation. No federal statute requires organizations to adopt the AI RMF, but its use is increasingly reflected in organizational governance programs and vendor qualification processes for clinical AI. NIST also released a Generative AI Profile (NIST AI 600-1) in July 2024 identifying unique risks posed by generative AI and proposing corresponding risk management actions [14]. Both documents are operationally relevant to clinical AI deployments, particularly those involving large language models for regulatory writing or patient screening.
The Map and Measure functions carry the most immediate operational weight for clinical research. Map requires organizations to identify where AI systems interact with sensitive data, including EHR integrations, FHIR APIs, and document generation pipelines. Measure establishes ongoing monitoring of AI behavior, performance, and fairness. OCR investigations have repeatedly found that organizations do not know where all ePHI resides in their systems. Agentic AI deployments, where models take sequential actions across multiple systems without explicit human review of each step, expand that inventory problem substantially.
When Multiple Frameworks Apply Simultaneously
A single clinical AI platform may face all three frameworks at once: HIPAA if it processes ePHI from a covered entity site, 21 CFR Part 11 if it generates or stores regulated records in an FDA-governed trial, and NIST AI RMF as a governance reference for managing AI-specific risks. For trials enrolling data subjects in the Union, GDPR Article 3 establishes territorial scope regardless of where the processing organization is located, and Regulation (EU) 2025/327 on the European Health Data Space (EHDS), which entered into force on March 26, 2025, with staged application dates running through 2029 and 2031, adds obligations for the secondary use of electronic health data, including for research and AI training purposes [22],[26]. The architecture must satisfy the applicable requirements at each decision point across all relevant frameworks.
Architecture Decisions That Determine HIPAA Defensibility
The Business Associate Agreement Is a Threshold, Not a Guarantee
Any AI vendor that creates, receives, maintains, or transmits PHI on behalf of a covered entity or business associate is itself a business associate under HIPAA and requires a signed BAA before PHI flows to that vendor [5]. This includes cloud service providers. HHS has confirmed in published guidance that cloud service providers are business associates when they handle ePHI on behalf of regulated entities, even if they only store it in encrypted form without accessing the content [16].
Major cloud platforms execute BAAs for enterprise customers. AWS offers a self-service BAA process and lists HIPAA-eligible services in its documentation [17]. Comparable arrangements are available from Microsoft Azure and Google Cloud Platform. For AI-specific tools, BAA availability depends on product tier and intended use. OpenAI offers BAAs to eligible API customers and through its ChatGPT for Healthcare and ChatGPT Enterprise products, subject to case-by-case review [21]. BAAs are not available for free consumer accounts, and entering participant PHI into those products constitutes a HIPAA violation regardless of research purpose. Organizations should obtain and review the specific BAA terms from each vendor, not rely on general statements of HIPAA compliance, before PHI flows to any AI system.
Three realities follow from BAA execution. First, the BAA is contractual, not technical. It defines permitted uses, prohibits uses beyond what the contract or law allows, and establishes breach notification obligations. It does not configure infrastructure. Cloud providers operate under a shared responsibility model: the provider secures the underlying physical and virtualization layer, while the customer is responsible for access controls, encryption key management, audit logging, network segmentation, and application-level security [16]. A signed BAA is necessary but never sufficient for HIPAA compliance.
Second, the obligation chain extends downstream. Under 45 CFR §164.308(b)(4), when a business associate engages a subcontractor that handles PHI, that subcontractor is also a business associate and requires its own BAA [5]. Sponsors who contract with a CRO that runs AI tools on cloud infrastructure must confirm that downstream BAA coverage is in place at each link.
Third, the BAA must address AI-specific terms. An agreement that was executed before generative AI was in scope may not address how prompts are handled, whether PHI in prompts is used for model training, what data retention limits apply, and what incident notification obligations attach to model-level exposures. Each of these should be confirmed before PHI flows to a generative AI system [15].
Data Governance: De-identification in Practice
HIPAA's Privacy Rule at §164.514(a)-(b) establishes two de-identification pathways [11]. Safe Harbor requires removal of 18 specific identifier categories from every dataset, including names, dates below year-level, geographic subdivisions smaller than a state, and any other characteristic that could uniquely identify an individual. It is operationally simpler to implement but removes granularity that AI training frequently requires. Safe Harbor-compliant datasets can retain re-identification risk when combined with external data sources, a documented limitation of the method [11].
Expert Determination uses statistical and computational methods to demonstrate that the probability of identifying any individual in the dataset is very small. It requires a qualified expert to examine the data, apply generally accepted statistical and disclosure-control practices, and produce documentation that the de-identification approach is defensible [11]. The resulting dataset retains substantially more analytical utility, which makes Expert Determination the more appropriate method for clinical AI model training in most cases. The cost is expertise and documentation overhead upfront, both of which are necessary regardless.
- • Removes 18 defined identifier categories
- • Simpler to operationalize
- • Reduces dataset granularity
- • May retain re-identification risk when combined with external data
- • Uses statistical and computational assessment
- • Requires qualified expert documentation
- • Preserves more analytical utility
- • Better suited for most clinical AI model training use cases
The minimum necessary standard constrains AI tools in production, even where de-identification handles training data. Tools that process identifiable PHI must be configured to request only the data elements required for their stated function. The natural AI development pattern, where broad access is granted during model development and trimmed after deployment, runs in the wrong direction from a HIPAA design perspective. Access architecture should be designed minimum-necessary-forward, not retrofitted.
Audit Trails: What the Record Must Contain
HIPAA's Security Rule requires covered entities to implement mechanisms that record and examine activity in information systems that contain or use ePHI. For AI systems in clinical research, the practical content of a defensible audit trail goes well beyond a standard access log.
Effective audit trail design should capture six elements for each logged event: who acted (user identity), what action was taken (specific operation or resource accessed), when it occurred (timestamped with time zone), where the request originated (IP address and device details), why the action was taken when clinical context can be captured, and the outcome, including whether the operation succeeded, failed, or modified ePHI [18]. These elements align with HIPAA Security Rule Audit Controls at §164.312(b), which requires covered entities and business associates to implement hardware, software, and procedural mechanisms to record and examine activity in systems containing ePHI [9], with 21 CFR Part 11 audit trail requirements at §11.10(e) [12], and with NIST SP 800-66r2 guidance on implementing the HIPAA Security Rule [25].
Audit logs must be stored with tamper-evident controls. Mutable logs, where unauthorized users can alter or delete audit entries, undermine the evidentiary value of the record and expose the organization in OCR investigations. Retention should follow applicable federal records schedules, and organizations should confirm that their AI vendor's audit records are accessible and exportable under the terms of the BAA. An audit trail that exists within a vendor's infrastructure but cannot be retrieved by the customer provides no practical compliance value.
For 21 CFR Part 11-regulated records, organizations should ensure the AI system's audit trail captures all required elements and remains accessible for the full retention period. AI-generated documents in regulated clinical trials should be treated no differently than other electronic records in this respect.
Encryption, Network Segmentation, and Patch Management
Under current HIPAA rules, encryption at rest and in transit is an addressable specification, meaning an equivalent alternative or documented risk acceptance is technically permissible. The proposed NPRM would change this to required [9]. For organizations building AI infrastructure now, treating encryption as mandatory is the appropriate design posture regardless of when the NPRM is finalized.
Amazon VPC and equivalent network isolation constructs on Azure and GCP provide the foundational layer for separating AI workloads that process ePHI from general organizational networks. Security group restrictions, endpoint detection on compute resources that access PHI, and explicit traffic controls between AI system components are baseline requirements [17]. Supply chain compromise at 15% of 2025 breach initial access vectors is a reminder that model weights, training pipelines, and inference APIs each represent points of potential compromise in a multi-vendor AI environment [3].
The proposed NPRM also introduces explicit patch management requirements, specifying that regulated entities must install patches, updates, and upgrades throughout relevant electronic information systems [9]. For AI platforms with Python library dependencies, model-serving containers, and underlying framework versions, this translates into a documented patching cadence with evidence of application. Framework-level vulnerabilities in common ML libraries have been identified in the past; treating model infrastructure as a security surface that requires patching is not optional under the proposed rule's logic.
A Compliance Architecture Reference Checklist
The following checklist reflects the key controls covered in this article. It is not a substitute for a formal risk analysis but can serve as a scoping tool when evaluating a clinical AI deployment.
| Control Area | Key Question | Primary Reference |
|---|---|---|
| Entity classification | Is this organization a covered entity, business associate, or neither? | 45 CFR §160.103 |
| Data flow mapping | Where does PHI from covered entities enter the AI system? | 45 CFR §164.306(a) (risk analysis); NPRM 90 FR 800 (proposed asset inventory) |
| BAA execution | Is there a signed BAA with every vendor touching ePHI as a BA? | 45 CFR §164.308(b)(1) |
| BAA scope | Does the BAA address AI-specific terms (training, retention, prompts)? | 45 CFR §164.504(e) |
| De-identification | Is data accessed by the AI de-identified by Safe Harbor or Expert Determination? | 45 CFR §164.514(a)-(b) |
| Minimum necessary | Is AI data access limited to what the function requires? | 45 CFR §164.502(b) |
| Access controls | Is role-based access enforced; unique user IDs assigned? | 45 CFR §164.312(a)(1); §164.312(a)(2)(i) |
| Authentication / MFA | Is identity verified before ePHI access? MFA proposed as explicit required control. | 45 CFR §164.312(d); NPRM 90 FR 800 |
| Encryption | Is ePHI encrypted at rest and in transit? (Currently addressable; proposed as required.) | 45 CFR §164.312(a)(2)(iv); §164.312(e)(2)(ii) |
| Audit trail | Do logs capture who, what, when, where, why, and outcome? | 45 CFR §164.312(b); 21 CFR §11.10(e) |
| System validation | Has the AI system been validated per 21 CFR Part 11 / GxP? | 21 CFR §11.10(a) |
| Patch management | Is there a documented patching cadence for AI system components? | NPRM 90 FR 800 |
| AI asset inventory | Is the AI system listed in the ePHI system inventory? | NPRM 90 FR 800 |
| Subcontractor BAAs | Are downstream subcontractors of AI vendors who act as BAs under BAAs? | 45 CFR §164.308(b)(4) |
| Vendor due diligence | Has the AI vendor's PHI handling, training data use, and breach notification been reviewed? | BAA terms; 45 CFR §164.504(e) |
Note: CFR section references above are simplified navigation anchors. During compliance implementation, map each control to the full regulatory text and applicable guidance documents, as implementation specifications may span multiple sections and are subject to the interpretations in the proposed NPRM.
Several patterns appear frequently in compliance assessments of AI deployments in clinical research.
Regulatory and Documentation Considerations
21 CFR Part 11 in Practice for Clinical AI
The validation obligation under 21 CFR Part 11 means that AI systems used to create, process, or store electronic records in FDA-regulated trials require documented evidence that they perform as intended under the conditions of use [12]. The scope of that validation, including how many test cases are needed, how performance criteria are defined, and what constitutes a system change triggering re-validation, follows GxP risk principles rather than a fixed template. GAMP guidance on computerized systems provides a widely accepted framework for applying these risk-proportionate principles [13].
AI-specific validation challenges include non-deterministic output (the same input may produce different outputs across runs), model version changes (a new model version may behave differently even on the same prompts), and the difficulty of constructing pass/fail criteria for generated narrative text. Sponsors and CROs deploying AI for regulatory document generation should document these challenges explicitly in their validation plans, define what constitutes acceptable and unacceptable output, and establish a change control procedure that addresses model version updates.
The Difference Between HIPAA, Common Rule, Part 11, and GDPR
Clinical research personnel frequently encounter multiple privacy and data protection frameworks at once, and their distinct scopes create practical confusion worth clarifying briefly.
HIPAA governs PHI in the hands of covered entities and their business associates, as described above. The Common Rule (45 CFR Part 46) governs human subjects research funded or conducted by federal departments and agencies, and requires IRB review and informed consent, subject to IRB waiver or alteration when the criteria in 45 CFR §46.116(c) and (d) are met [19]. FDA human subjects regulations at 21 CFR Parts 50 and 56 impose parallel requirements for FDA-regulated trials. The three frameworks overlap in their consent and IRB provisions but are administered separately and have different scopes. 21 CFR Part 11 governs electronic records and signatures in FDA-regulated trials and has no HIPAA analog; it applies based on regulatory submission context, not data type. GDPR governs personal data of data subjects in the Union under Article 3, applying to data flows from EU investigator sites regardless of the sponsor's location [22]. Regulation (EU) 2025/327 on the European Health Data Space, in force since March 26, 2025, adds a sector-specific layer governing the secondary use of electronic health data for research and AI development purposes, with relevant provisions applying from March 2029 [26].
For a global clinical trial running across EU and US sites, all four frameworks typically apply simultaneously to different data flows. The AI infrastructure serving that trial must be configured to satisfy each, which in practice means understanding which data elements originate from which regulatory jurisdiction and ensuring the controls governing each flow are appropriate to that jurisdiction's requirements.
AI and Automation Perspective
AI's contribution to clinical trial infrastructure is real, but its compliance properties are consistently overstated during vendor evaluations. Several observations apply across deployment contexts.
AI does not reduce regulatory obligation. Adding an AI document generation tool to a regulatory writing workflow adds a new regulated computerized system requiring validation, audit logging, BAA coverage, and inclusion in the organization's risk analysis. It does not simplify the compliance posture; it adds a component to the posture.
Detection capability is where AI delivers the clearest compliance return. IBM's 2025 breach data shows that organizations using AI and automation extensively in security operations reduced breach lifecycles by 80 days and lowered average breach costs by $1.9 million compared to those without such tools [3]. The operational consequence is faster containment, lower notification costs, and reduced time during which PHI is exposed. Investment in AI for detection is quantifiably different from investment in AI for data processing, from a security economics standpoint.
The minimum necessary standard and AI's performance appetite pull in opposite directions. AI models often improve with broader data access; HIPAA requires limiting access to what the specific function requires. This tension is not resolved by vendors by default. It requires explicit architectural decisions that enforce access limits at the data layer and explicit contractual terms that prevent training data use beyond defined purposes.
Human oversight remains a regulatory expectation regardless of model capability. The FDA's Clinical Decision Support (CDS) guidance distinguishes between non-device software, where clinicians can independently review the basis for AI recommendations, and device software, where decisions are made autonomously or in ways clinicians cannot meaningfully verify [20]. AI tools that generate regulatory documentation for human review are generally on the non-device side of this line. Tools that route trial participants or flag safety signals without reviewable rationale approach the device classification boundary and warrant formal classification analysis before deployment.
How Kitsa Fits Into This Problem
Kitsa's platform, certified to SOC 2, HIPAA, and ISO 27001, and deployed within AWS VPC, is designed to operate inside the compliance stack described above rather than alongside it [23]. KScribe, Kitsa's AI regulatory document generation product, generates protocols, informed consent forms, investigator brochures, DSURs, and clinical study reports within an infrastructure designed for validated, auditable use [24]. The platform's architecture reflects the controls in this article: access controls, encryption, audit trails, and a GCP-aligned validation framework, providing the compliance scaffolding that general-purpose AI tools require organizations to build themselves.
HIPAA-compliant AI infrastructure for clinical trials requires more than a BAA. It needs validated systems, ePHI-aware data flows, encryption, role-based access, audit trails, vendor governance, and human oversight. Kitsa's platform is designed to support regulated clinical research workflows across AI regulatory writing, patient pre-screening, and trial operations within an auditable infrastructure.
Key Takeaways
- •HIPAA applies in clinical trials through the covered entity and business associate framework, not merely because a trial is occurring. Sponsors, CROs, and AI vendors must each assess their specific data relationships before assuming HIPAA obligations attach or do not attach. Where HIPAA does not apply, other obligations including research authorizations, DUA terms, Common Rule data protections, applicable state privacy law, and contractual commitments may still govern identifiable participant data.
- •HHS OCR's proposed HIPAA Security Rule NPRM (January 2025) would eliminate the required/addressable distinction and make controls such as encryption, technology asset inventories, and annual compliance audits mandatory. It also proposes adding multi-factor authentication as a new explicit required control, which the current rule does not name by that term. As of publication, this rule is proposed, not final, with a 2026 finalization target and a 240-day compliance window from publication.
- •Clinical AI platforms in U.S. FDA-regulated trials must comply with applicable regulatory requirements, primarily HIPAA Security Rule and 21 CFR Part 11. The NIST AI RMF is voluntary guidance, not a regulation, but provides a widely referenced governance framework for managing AI-specific risks. Trials enrolling data subjects in the EU face obligations under GDPR and Regulation (EU) 2025/327 (EHDS), which entered into force March 26, 2025, with secondary-use provisions applying from 2029.
- •A signed BAA is a mandatory legal threshold, not evidence of technical compliance. The shared responsibility model places configuration, access controls, encryption, and audit logging with the deploying organization, not the cloud provider.
- •Expert Determination under HIPAA §164.514(a)-(b) is generally more appropriate than Safe Harbor for clinical AI model training because it preserves analytical utility while demonstrating statistical non-identifiability. It requires qualified expertise and documented methodology.
- •AI systems that create, process, or store electronic records in FDA-regulated trials require Computer System Validation under 21 CFR Part 11, including change control procedures that cover model version updates.
- •IBM's 2025 Cost of a Data Breach Report shows that organizations using AI and automation extensively reduced breach lifecycles by 80 days and lowered average breach costs by $1.9 million, making detection investment the clearest operational return from AI in the compliance context [3].
FAQ
Does HIPAA automatically apply to every AI tool used in a clinical trial?
Do I need a separate BAA for each AI vendor in my clinical trial technology stack?
What is the difference between Safe Harbor and Expert Determination for de-identifying clinical trial data?
Does 21 CFR Part 11 apply to AI-generated clinical documents?
What should be in the audit trail for a clinical AI system handling ePHI?
How does HIPAA interact with Common Rule and FDA human subjects regulations in a clinical trial?
References
- [1] Associated Press. "UnitedHealth says Change Healthcare cyberattack cost it $872 million in Q1." AP News, April 2024. https://apnews.com/article/076a2bcbf0db3b5e6c4ffa27589c3692
- [2] GovTech / Ars Technica. "Change Healthcare CEO: Hack May Have Touched a Third of Americans." April 2024. https://www.govtech.com/security/ceo-change-healthcare-hack-may-touch-a-third-of-americans
- [3] IBM Security. "Cost of a Data Breach Report 2025." IBM / Ponemon Institute, July 2025. https://www.ibm.com/reports/data-breach
- [4] Netskope Threat Labs. "Threat Labs Report: Healthcare 2025." Netskope, 2025. https://www.netskope.com/resources/threat-labs-reports/threat-labs-report-healthcare-2025
- [5] U.S. Department of Health and Human Services. "Business Associates." HHS.gov. https://www.hhs.gov/hipaa/for-professionals/privacy/guidance/business-associates/index.html
- [6] U.S. Department of Health and Human Services, Office for Civil Rights. "Research Uses and Disclosures." HHS.gov FAQ. https://www.hhs.gov/hipaa/for-professionals/faq/research-uses-and-disclosures/index.html
- [7] National Institutes of Health. "IRB and the HIPAA Privacy Rule." NIH Office for Human Research Protections. https://privacyruleandresearch.nih.gov/irbandprivacyrule.asp
- [8] Bloomberg Law / Health Law and Business. "Business Associates and Clinical Research: Resolving a HIPAA Compliance Conundrum." Bloomberg Law, 2017. https://news.bloomberglaw.com/health-law-and-business/business-associates-and-clinical-research-resolving-a-hipaa-compliance-conundrum-1
- [9] HHS OCR. "HIPAA Security Rule Notice of Proposed Rulemaking to Strengthen Cybersecurity for Electronic Protected Health Information." Federal Register 90 FR 800, January 6, 2025. https://www.federalregister.gov/documents/2025/01/06/2024-30983/hipaa-security-rule-to-strengthen-the-cybersecurity-of-electronic-protected-health-information
- [10] Office of Information and Regulatory Affairs. "RIN 0945-AA22: Modifications to the HIPAA Security Rule." RegInfo.gov Unified Agenda. https://www.reginfo.gov/public/do/eAgendaViewRule?RIN=0945-AA22
- [11] HHS OCR. "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with HIPAA Privacy Rule." HHS.gov. https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html
- [12] U.S. Food and Drug Administration. "21 CFR Part 11: Electronic Records; Electronic Signatures." eCFR. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- [13] ISPE. "GAMP 5: A Risk-Based Approach to Compliant GxP Computerized Systems." ISPE, 2nd ed., 2022. https://ispe.org/publications/guidance-documents/gamp-5
- [14] NIST. "Artificial Intelligence Risk Management Framework (AI RMF 1.0)." NIST AI 100-1, January 26, 2023. https://doi.org/10.6028/NIST.AI.100-1
- [15] Foley & Lardner LLP. "HIPAA Compliance for AI in Digital Health: What Privacy Officers Need to Know." Foley Health Care Law Today, May 2025. https://www.foley.com/insights/publications/2025/05/hipaa-compliance-ai-digital-health-privacy-officers-need-know/
- [16] HHS OCR. "Does the HIPAA Security Rule Apply to a CSP That Only Stores Encrypted ePHI?" HHS.gov FAQ. https://www.hhs.gov/hipaa/for-professionals/faq/2079/what-if-a-hipaa-covered-entity-or-business-associate-uses-a-csp-to-maintain-ephi-without-first-executing-a-business-associate-agreement-with-that-csp/index.html
- [17] Amazon Web Services. "HIPAA Compliance for Generative AI Solutions on AWS." AWS Industries Blog, October 2025. https://aws.amazon.com/blogs/industries/hipaa-compliance-for-generative-ai-solutions-on-aws/
- [18] Censinet. "The Audit Trail Imperative: Documentation Standards for Healthcare AI." Censinet Perspectives, May 2026. https://censinet.com/perspectives/audit-trail-imperative-documentation-standards-healthcare-ai
- [19] HHS OHRP. "45 CFR Part 46: The Common Rule." OHRP. https://www.hhs.gov/ohrp/regulations-and-policy/regulations/45-cfr-46/index.html
- [20] U.S. Food and Drug Administration. "Clinical Decision Support Software: Frequently Asked Questions." FDA.gov. https://www.fda.gov/medical-devices/software-medical-device-samd/clinical-decision-support-software-frequently-asked-questions-faqs
- [21] OpenAI. "Enterprise Privacy at OpenAI." OpenAI, 2025. (BAA available for eligible API and ChatGPT for Healthcare customers.) https://openai.com/enterprise-privacy/
- [22] European Union. "Regulation (EU) 2016/679 (GDPR), Article 3: Territorial Scope." EUR-Lex, April 27, 2016. https://gdpr-info.eu/art-3-gdpr/
- [23] Kitsa. "Compliance and Certifications." Kitsa.ai. https://kitsa.ai/certifications
- [24] Kitsa. "KScribe: AI Regulatory Document Generation." Kitsa.ai. https://kitsa.ai/regulatory-document-generation
- [25] NIST. "Implementing the HIPAA Security Rule: A Cybersecurity Resource Guide." NIST SP 800-66 Revision 2, July 2022. https://doi.org/10.6028/NIST.SP.800-66r2
- [26] European Union. "Regulation (EU) 2025/327 on the European Health Data Space." OJEU, March 5, 2025. Entry into force March 26, 2025. https://eur-lex.europa.eu/eli/reg/2025/327/oj/eng
