Contents
Introduction
Phase III protocols now average 5.96 million data points per trial, and nearly one-third of the procedures generating those data do not directly support primary or key secondary endpoints, according to a September 2025 collaborative study by TransCelerate BioPharma and the Tufts Center for the Study of Drug Development (CSDD), drawing on 105 Phase II and III protocols across 14 biopharmaceutical companies [1]. The average Phase III trial now carries 3.5 protocol amendments, a figure that has grown more than 50 percent over the prior five years, with protocol deviations per trial averaging 296, nearly triple the level from a decade prior, per Tufts CSDD data presented at the Evolution Summit in May 2025 [2].
Those numbers describe a system under strain, not a collection of individual task failures. Yet the AI tools most organizations have deployed to address that strain are, almost without exception, single-task instruments: one tool drafts an Investigator's Brochure (IB), a second generates a protocol synopsis, a third screens patient eligibility criteria. Each performs its assigned function in isolation, with no awareness of what the others have produced.
The question now appearing in sponsor strategy sessions and regulatory affairs teams is not whether AI helps in clinical research. It clearly does. The question is whether AI designed to execute isolated tasks can address failures that are, at their root, systemic.
System-aware AI maintains trial context across documents. Single-task AI only completes isolated actions.
Why the Clinical Trial Document Ecosystem Breaks Down
A clinical trial produces a family of interdependent documents: the protocol, the Informed Consent Form (ICF), the Investigator's Brochure (IB), the Development Safety Update Report (DSUR), the Clinical Study Report (CSR), and the Statistical Analysis Plan (SAP). These are not independent files. They share common content: the indication, the eligibility criteria, the dosing regimen, the primary and secondary endpoints, the identified risks. A change to any one of these elements is supposed to propagate accurately and consistently across all of them. AI-native platforms such as KScribe are built specifically to generate these documents from a shared data foundation, but even purpose-built tools can only deliver cross-document consistency if their underlying architecture maintains a representation of the trial rather than treating each document as an independent generation task.
In practice, it rarely does. A May 2025 study published in npj Digital Medicine noted that the "implementation gap" in medical AI, where research advances fail to benefit clinical workflows, persists partly because AI systems are evaluated on narrow, task-specific benchmarks rather than on their behavior in connected clinical workflows. Rosenthal et al. observed that by 2024, only 86 randomized trials of machine learning interventions had been conducted worldwide, in part because evaluators struggled to validate how single-model outputs perform when embedded in real-world, multi-step clinical settings [3].
The structural risk is not that an AI tool produces a bad document. The risk is that it produces a document internally consistent but externally inconsistent with every other document in the trial master file (TMF). An eligibility criterion written one way in the protocol, slightly differently in the ICF, and not represented in the patient pre-screening logic creates a deviation waiting to happen. When deviations accumulate, sponsors write amendments. ICH E6(R3), which took effect at the EMA on July 23, 2025, and was adopted at Step 4 of the ICH process on January 6, 2025, explicitly addresses this through Quality-by-Design (QbD) principles requiring sponsors to proactively identify factors critical to trial quality before, not after, the protocol reaches the site [4].
What Single-Task AI Actually Does (and Does Not Do)
Single-task AI, in the clinical research context, refers to any AI application designed to perform one bounded operation: generate a document section, classify a patient record, extract eligibility criteria, predict an adverse event. These tools have demonstrable value within their defined scope. A 2025 scoping review published in npj Digital Medicine, analyzing 142 studies of AI applications in clinical trial risk assessment published between 2013 and 2024, found that AI techniques, including traditional machine learning, deep learning, and large language models, achieved high performance on specific risk prediction tasks, with area under the receiver operating curve (AUROC) values reaching 0.96 for adverse drug event prediction models [5].
High task-level performance is not the same as system reliability. The Rosenthal et al. paper makes the distinction explicitly: the "implementation gap" between AI research and clinical deployment persists because most AI systems are evaluated as standalone models rather than as participants in multi-step, multi-document workflows [3]. Single-task tools have three structural limitations in that context.
First, they lack awareness of upstream and downstream document state. A tool drafting a protocol synopsis has no representation of what the IB has already stated about known risks. A tool generating ICF language has no representation of how the protocol framed the same inclusion criteria. These are not bugs; they are design boundaries. The tools produce outputs, not shared state.
Second, they treat each generation event as independent. Reviewing an AI-generated ICF against a protocol requires a human to hold both documents in working memory and manually check every shared element. When protocols exceed 300 procedures, as the average Phase III protocol now does [2], that check becomes an unreliable manual process rather than a systematic one.
Third, single-task tools have no mechanism for change propagation. When a Phase II interim result forces a dosing amendment, a single-task tool can regenerate the affected section. It cannot identify which other sections, in which other documents, reference the affected dose. A 2016 study by Getz et al. published in Therapeutic Innovation and Regulatory Science, drawing on data from 836 Phase I through IIIb/IV protocols across 15 pharmaceutical companies and CROs, found that 57 percent of protocols had at least one substantial amendment, with a median direct cost of $141,000 per Phase II amendment and $535,000 per Phase III amendment [6]. The downstream costs, re-review, IRB re-submissions, site re-training, and contract change orders, are substantial (see Kitsa's analysis of protocol amendment drivers and prevention). Single-task AI does not directly address one of the important drivers of that cost: cross-document inconsistency and the absence of a change propagation mechanism across the document family.
What System-Aware AI Does Differently
The table below summarizes the structural differences between disconnected single-task AI tools and a system-aware clinical research platform:
| Capability | Disconnected Single-Task Tools | System-Aware Platform |
|---|---|---|
| Cross-document consistency | Each document generated independently; no shared state | All documents generated from a common trial data model |
| Change propagation | Manual identification of affected sections across documents | Data-element changes can be flagged across the full document family |
| Provenance / traceability | Output generated without reference logs | Can maintain section-level logs of source trial data elements |
| Amendment risk | Cannot detect cross-document contradictions before submission | May flag internal contradictions in draft before regulatory review |
| Human review integration | Reviewer receives disconnected outputs | Role-based review against a shared source-of-truth |
| M11 / USDM alignment | Typically not structured against a machine-readable protocol schema | Can be built to align with ICH M11 and CDISC USDM data structures |
Example, a single dose change across four documents. A DSMB recommendation to reduce a Phase II dose from 200 mg to 150 mg touches at least the protocol (dosing schedule, stopping rules), the ICF (risk language, participant expectations), the SAP (exposure-response estimands, subgroup definitions), and eCRF field configurations. In a disconnected tool environment, a medical writer must locate and reconcile each instance manually. In a system-aware platform where dose is a defined trial data element, a change to that element surfaces every downstream reference across the full document family before any document is regenerated.
The operational problem is not generating one revised sentence. It is finding every place where the changed data element matters.
System-aware AI in clinical research describes infrastructure in which AI components operate with shared knowledge of the trial's data model, document state, and regulatory context. Rather than treating each document generation event as a standalone prompt, a system-aware architecture maintains a representation of the trial, its indication, molecule, phase, enrolled population, identified risk profile, and primary endpoints, and uses that representation to constrain and inform every document it touches.
System-aware AI is distinct from simple prompt chaining, where AI outputs are fed sequentially into subsequent AI calls. Prompt chaining automates steps but each call still lacks awareness of the broader trial context, it does not know what other documents exist, what shared data elements they contain, or what regulatory constraints govern their relationship. A system-aware platform maintains a persistent representation of the trial across all generation events.
The operational difference is significant. Consider a scenario where a Phase II interim review leads the Data Safety Monitoring Board to recommend a dose reduction. In a single-task environment, that change requires a medical writer to manually locate every reference to the original dose across the protocol body, the ICF risk section, the pharmacy manual, and the eCRF instructions, a process spanning potentially dozens of sections across at least four documents. A system-aware platform that maintains the dose as a defined data element in its trial data model can flag every downstream location where that element appears, across the full document family, before the change request is even acknowledged.
ICH M11, the Clinical Electronic Structured Harmonised Protocol (CeSHarP) standard, was adopted at Step 4 by the ICH Assembly at its Singapore meeting in November 2025, and FDA published the final M11 CeSHarP guidance on May 22, 2026 [7]. The M11 Technical Specification defines data elements and attributes designed for interoperable electronic exchange of protocol content. TransCelerate's Digital Data Flow initiative, developed in partnership with CDISC through the Unified Study Definition Model (USDM), extends this by creating a common information model capable of representing protocol content in machine-readable form that downstream systems, including AI authoring tools, can reference [8].
Multi-agent AI frameworks, where separate AI components handle specialized tasks but exchange information through a shared data layer, represent one leading architecture for implementing system awareness in practice. A study published in npj Health Systems in March 2026 by Klang, Omar, Raut et al. at the Icahn School of Medicine at Mount Sinai tested four large language models under clinical-scale workloads using two configurations: a single agent handling all tasks, and a multi-agent orchestrator assigning each task to a dedicated worker. Multi-agent accuracy remained at 65.3 percent at 80 concurrent tasks, while single-agent accuracy collapsed from 73.1 percent to 16.6 percent under the same load [9]. That study addressed general clinical data tasks, retrieval, extraction, and dosing, rather than regulatory document generation specifically, but the core finding about accuracy degradation under task accumulation is relevant by analogy to any AI system handling heterogeneous, simultaneous requests.
In clinical trial operations, where decisions include cross-document consistency checks, deviation flagging, and amendment impact assessment, the shift from single-task to system-aware AI has direct operational consequences. The healthcare AI literature reflects this: peer-reviewed studies of multi-agent architectures show meaningful performance advantages over single-agent configurations at scale [9], and the agentic AI scoping review in healthcare identified workflow optimization as the domain with the strongest reported outcomes [14].
The Regulatory and Documentation Argument
FDA's January 2025 draft guidance (Docket FDA-2024-D-4689) introduced a risk-based credibility assessment framework for AI that generates data or analysis intended to directly support regulatory decision-making about drug safety, efficacy, or quality [10]. It is important to note what the guidance does not cover: FDA explicitly excludes AI used for operational efficiencies, including the mechanics of drafting regulatory submissions, when those uses do not affect patient safety, drug quality, or the reliability of nonclinical or clinical study results [10]. A document drafting tool that formats text from inputs provided by the sponsor is outside the guidance's scope. An AI tool that analyzes clinical or pharmacological data whose outputs are then submitted to FDA as evidentiary support for a safety or efficacy determination is within scope.
That distinction matters for the system-aware AI argument. The relevance of FDA-2024-D-4689 to document generation platforms is not direct: it governs AI-generated evidence, not AI-assisted writing. The relevant frameworks for AI-generated clinical documents are the January 14, 2026 joint FDA-EMA Guiding Principles of Good AI Practice in Drug Development, which introduced lifecycle-spanning requirements for human-centric design, data governance, documentation, and lifecycle management applicable across AI use in the drug development process [11]. These principles presuppose that AI tools can be audited, that their outputs are traceable, and that the human oversight layer is clearly defined, all design requirements that distinguish system-aware platforms from collections of independent task tools.
The FDA has also accelerated its own internal AI adoption. On June 2, 2025, the agency launched Elsa, a generative AI tool deployed agency-wide to assist scientific reviewers with clinical protocol review, adverse event summarization, and scientific evaluations. FDA Commissioner Makary stated that, following a pilot with scientific reviewers in which tasks that previously took days were completed in minutes [13], he set a June 30 deadline for agency-wide deployment; Elsa launched ahead of that date [12]. Independent clinical research publications noted accuracy and oversight concerns shortly after launch, a reminder that internal FDA efficiency gains do not remove the need for well-structured, traceable sponsor submissions [13].
Practical Implications for Protocol Design, Startup, and Documentation
The shift from single-task to system-aware AI has its most immediate impact at two points in the trial lifecycle: protocol authoring and the startup-to-activation window.
Protocol authoring. The TransCelerate and Tufts CSDD 2025 study found that nearly one-third of Phase II and III procedures across 14 companies did not directly support primary or key secondary endpoints [1]. One structural cause is that protocol sections are frequently authored by different teams at different times, with no mechanism for checking whether the eligibility criteria in Section 4 are consistent with the risk language in Section 8 or the stopping rules in Section 9. A system-aware AI platform that holds the complete trial data model during authoring can flag those inconsistencies in draft before the protocol reaches regulatory review, which may reduce the risk that an internal contradiction drives a post-submission amendment.
Trial startup. The same Tufts CSDD data that show 3.5 amendments and 296 deviations per average Phase III trial also show a 60-percent increase in procedures since 2015 [2]. Much of that complexity lands on sites during the period between protocol finalization and first patient in, when site staff must interpret the protocol, build study-specific templates, and complete role-specific training before they can enroll a single patient. Sites must translate increasingly complex protocols into execution-ready materials, a process that compounds when the protocol, ICF, and site-facing documents contain slightly different versions of shared content. When a system-aware platform generates those materials from the same underlying trial data model, that interpretation burden is lower and the risk of site-level deviation from the intended design is reduced.
Regulatory submissions. The FDA-EMA joint principles require that AI tools be designed for human oversight and that their outputs be traceable throughout the product lifecycle [11]. A system-aware platform that generates each document section with a log of which trial data elements informed the output is structurally better positioned to satisfy post-hoc audit requests than a collection of independently generated documents assembled after the fact.
Limitations and Cautions
None of this means system-aware AI platforms are ready to operate without expert oversight. ICH E6(R3), effective at the EMA from July 23, 2025, requires that quality management systems include mechanisms for oversight and accountability that cannot be delegated to automated systems [4]. The FDA-EMA joint principles published January 14, 2026 include human-centric design as their first principle, meaning that the reduction of human review burden is a goal but not a license to eliminate it [11].
A scoping review published in npj Digital Medicine (2026), analyzing agentic AI applications across emergency medicine, oncology, radiology, and rehabilitation, found that systems featuring multi-agent collaboration demonstrated the strongest outcomes in workflow optimization tasks. The authors identified only seven eligible studies in total and noted that most were exploratory, limited in scope, and lacked prospective clinical validation at scale [14].
The practical caution for document generation is this: a system-aware platform that confidently propagates an incorrect piece of trial data across multiple documents can produce downstream errors at scale. Human expert review, by someone who understands the trial data model well enough to catch a plausible-sounding but factually wrong output, remains a necessary control regardless of the sophistication of the underlying platform.
How Kitsa Fits Into This Problem
Kitsa's approach at kitsa.ai is built on the premise that clinical research AI should be trial-aware, not just task-capable. KScribe, the platform's AI regulatory document generation product (kitsa.ai/regulatory-document-generation), generates protocols, ICFs, IBs, DSURs, and CSRs from a structured trial knowledge model rather than as independent generation tasks [16]. Kitsa's published approach emphasizes cross-document consistency, source-linked evidence, human review, audit trails, and documented approval workflows before regulated content leaves the system [17],[18],[19], consistent with the auditability expectations reflected in both ICH E6(R3) and the FDA-EMA joint AI principles for tools operating in regulated clinical research contexts.
Single-task AI can complete isolated clinical research tasks, but system-aware AI is designed to maintain context across the full trial ecosystem. Kitsa is built around this connected model: KScribe supports AI regulatory document generation from a structured trial knowledge model, KScout supports site selection intelligence, and KScreener supports patient pre-screening workflows, helping sponsors move from disconnected task automation toward trial-aware clinical research infrastructure.
Key Takeaways
- Phase III protocols now average 5.96 million data points, and nearly one-third of associated procedures do not directly support primary endpoints, per the 2025 TransCelerate/Tufts CSDD study across 105 protocols and 14 companies [1].
- Disconnected single-task AI tools perform well on bounded tasks but typically cannot maintain cross-document consistency, propagate changes across a document family, or establish the kind of traceable output lineage that the January 2026 FDA-EMA joint AI principles emphasize [11].
- System-aware AI operates from a shared trial data model, enabling consistency checks, change impact flagging, and traceable output generation across the full document ecosystem.
- ICH M11 CeSHarP was adopted at ICH Step 4 in November 2025, and FDA published the final guidance on May 22, 2026; together with the CDISC USDM, these standards provide the technical foundation for machine-readable, system-aware protocol authoring [7],[8].
- ICH E6(R3), effective at the EMA from July 23, 2025, requires Quality-by-Design approaches that can be better aligned with system-aware platforms than with collections of disconnected point tools [4].
- Human oversight of AI-generated clinical documents is a regulatory and scientific requirement under both ICH E6(R3) and the FDA-EMA joint principles. System-aware AI reduces review burden; it does not replace expert judgment.
- FDA-2024-D-4689 governs AI that generates evidentiary data for regulatory decisions, not AI used for operational drafting. The more directly applicable frameworks for document generation platforms are the broader FDA-EMA joint principles, which address AI governance across the full drug development lifecycle [10],[11].
Frequently Asked Questions
What is the difference between single-task AI and system-aware AI in clinical trials?
Why does cross-document inconsistency in clinical trials matter?
What does ICH E6(R3) require that is relevant to AI-generated documents?
Does FDA's January 2025 AI guidance apply to clinical document drafting tools?
What is ICH M11 and why does it matter for AI authoring platforms?
Can system-aware AI replace human medical writers in clinical research?
References
- [1]Getz KA, et al. (TransCelerate BioPharma / Tufts CSDD). "Insights Informing Strategies for Optimizing the Collection of Clinical Trial Data." Pre-print submitted to Therapeutic Innovation and Regulatory Science (DIA TIRS), September 2025. https://www.prnewswire.com/news-releases/transcelerate-and-tufts-csdd-uncover-opportunities-to-rethink-data-collection-and-optimize-protocol-design-302556373.html
- [2]Getz KA. Tufts CSDD, Evolution Summit Keynote, May 2025. Reported in: CRIO. "The Rising Complexity of Study Design: What It Means for Clinical Research Sites." 2025. Secondary reporting of keynote data; no publicly accessible primary Tufts CSDD publication for these specific figures. https://clinicalresearch.io/blog/the-rising-complexity-of-study-design-what-it-means-for-clinical-research-sites/
- [3]Rosenthal JT, Beecy A, Sabuncu MR. "Rethinking clinical trials for medical AI with dynamic deployments of adaptive systems." npj Digital Medicine 8:252 (2025). DOI: 10.1038/s41746-025-01674-3. https://www.nature.com/articles/s41746-025-01674-3
- [4]International Council for Harmonisation. "ICH E6(R3) Guideline for Good Clinical Practice." Endorsed at Step 4, January 6, 2025; EMA effective date July 23, 2025. https://www.ema.europa.eu/en/ich-e6-good-clinical-practice-scientific-guideline
- [5]Teodoro D, Naderi N, Yazdani A, Zhang B, Bornet A. "A scoping review of artificial intelligence applications in clinical trial risk assessment." npj Digital Medicine 8:486 (2025). DOI: 10.1038/s41746-025-01886-7. https://www.nature.com/articles/s41746-025-01886-7
- [6]Getz KA, Stergiopoulos S, Short M, et al. "The Impact of Protocol Amendments on Clinical Trial Performance and Cost." Therapeutic Innovation and Regulatory Science (2016). DOI: 10.1177/2168479016632271. https://journals.sagepub.com/doi/abs/10.1177/2168479016632271
- [7]ICH Assembly, Singapore, November 2025. "M11 Guideline on Clinical Electronic Structured Harmonized Protocol (CeSHarP) adopted at Step 4." Final FDA guidance: Federal Register, May 22, 2026. https://www.federalregister.gov/documents/2026/05/22/2026-10295/m11-clinical-electronic-structured-harmonised-protocol-cesharp-international-council-for
- [8]TransCelerate BioPharma. "Digital Data Flow Initiative." CDISC. "Digital Data Flow / Unified Study Definition Model (USDM)." https://www.cdisc.org/ddf
- [9]Klang E, Omar M, Raut G, et al. "Orchestrated multi agents sustain accuracy under clinical-scale workloads compared to a single agent." npj Health Systems 3:23 (2026). DOI: 10.1038/s44401-026-00077-0. https://www.nature.com/articles/s44401-026-00077-0
- [10]U.S. Food and Drug Administration. "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products." Draft Guidance, Docket FDA-2024-D-4689, January 7, 2025. https://www.federalregister.gov/documents/2025/01/07/2024-31542/considerations-for-the-use-of-artificial-intelligence-to-support-regulatory-decision-making-for-drug
- [11]U.S. FDA and European Medicines Agency. "Guiding Principles of Good AI Practice in Drug Development." January 14, 2026. https://www.fda.gov/about-fda/artificial-intelligence-drug-development/guiding-principles-good-ai-practice-drug-development
- [12]U.S. Food and Drug Administration. "FDA Launches Agency-Wide AI Tool to Optimize Performance for the American People." Press Release, June 2, 2025. https://www.fda.gov/news-events/press-announcements/fda-launches-agency-wide-ai-tool-optimize-performance-american-people
- [13]Applied Clinical Trials. "FDA's Elsa AI Tool Raises Accuracy and Oversight Concerns." June 2025. https://www.appliedclinicaltrialsonline.com/view/fda-elsa-ai-tool-raises-accuracy-and-oversight-concerns
- [14]Collaco BG, Haider SA, Prabha S, Gomez-Cabello CA, Genovese A, Wood NG, Bagaria SP, Gopala N, Tao C, Forte AJ. "The role of agentic artificial intelligence in healthcare: a scoping review." npj Digital Medicine 9:345 (2026). DOI: 10.1038/s41746-026-02517-5. https://www.nature.com/articles/s41746-026-02517-5
- [15]Getz K, Smith Z, Jain A, Krauss R. "Benchmarking Protocol Deviations and Their Variation by Major Disease Categories." Therapeutic Innovation & Regulatory Science 56:632-636 (2022). DOI: 10.1007/s43441-022-00401-4. PMID: 35378712. https://link.springer.com/article/10.1007/s43441-022-00401-4
- [16]Kitsa. "KScribe: AI Regulatory Document Generation." https://kitsa.ai/regulatory-document-generation
- [17]Kitsa. "How KScribe Uses Structured Clinical Intelligence." https://kitsa.ai/blog/kscribe-structured-clinical-intelligence-regulatory-documents
- [18]Kitsa. "Audit Trails and Traceability in AI Regulatory Systems." https://kitsa.ai/blog/audit-trails-traceability-ai-regulatory-systems-clinical-trials
- [19]Kitsa. "Human-in-the-Loop AI in Clinical Trial Documentation." https://kitsa.ai/blog/human-in-the-loop-ai-clinical-trial-documentation
