Contents
The number of vendors promising to automate clinical and regulatory document generation has grown sharply since late 2023, but the regulatory infrastructure required to use these tools responsibly has only recently begun to catch up. Between the FDA's draft guidance on AI in regulatory decision-making issued in January 2025 [1], the joint FDA-EMA guiding principles published in January 2026 [2], and the ISPE GAMP Guide: Artificial Intelligence released in July 2025 [3], sponsors now have a clearer set of expectations to work from. The question is no longer whether AI can accelerate regulatory document authoring. Several real-world implementations have already established that it can, particularly for CSRs and IND nonclinical summaries. The question is what separates a platform that can withstand sponsor QA, regulatory inspection, and reviewer scrutiny from one that creates compliance liability.
This article offers a structured evaluation framework for sponsors making procurement decisions on AI regulatory writing platforms. Each criterion maps to a documented compliance expectation, not a marketing checklist.
Why the Evaluation Pressure Has Intensified
The commercial case for AI-assisted regulatory writing is increasingly well documented. A McKinsey analysis of leading pharma regulatory operations found that generative AI can accelerate CSR completion by approximately 40 percent, compressing a timeline of 8 to 14 weeks down to 5 to 8 weeks and adding roughly $15 million to $30 million in net present value per asset [4]. A McKinsey-Merck collaboration reported an even sharper reduction: first-draft CSR writing time fell from 180 hours to 80 hours, with errors across data, messaging, citations, and terminology cut by 50 percent [5].
Those figures are persuasive. Both represent structured, validated implementations by organizations with established medical writing operations and formal human oversight workflows already in place; neither should be treated as a default capability expectation for a first deployment. Sponsors that skip the validation architecture because the efficiency numbers look attractive will learn this distinction at the cost of a regulatory query or, in worse cases, a submission rejection.
Tufts Center for the Study of Drug Development (CSDD) data makes the stakes plain. A Tufts CSDD follow-up study analyzing 950 protocols from 16 pharmaceutical companies and CROs found that 76 percent of Phase I-IV protocols now require at least one amendment, up from 57 percent in 2015 [6]. The median direct cost to implement a substantial amendment is $535,000 for a Phase 3 protocol and $141,000 for Phase 2, excluding site disruption and timeline delays, per Getz et al. (2016) in Therapeutic Innovation and Regulatory Science [7]. Protocol content generated by an unvalidated AI tool, if it introduces inconsistencies that require downstream correction, compounds exactly this problem rather than reducing it.
Sponsors face a legitimate choice between tools that reduce that burden and tools that transfer it elsewhere, including to regulatory reviewers.
The Regulatory Backdrop Sponsors Must Internalize
Any vendor evaluation starts with a shared understanding of the current regulatory environment. Three documents define the floor.
The FDA's draft guidance FDA-2024-D-4689, issued January 7, 2025, introduced a structured credibility assessment framework for AI used to support regulatory decision-making [1]. The framework recommends that sponsors define a context of use (COU) for each AI application, assess risk relative to that COU, document training data and model architecture, and demonstrate ongoing lifecycle management, including post-deployment monitoring for performance degradation [1]. As draft guidance, the document is not legally binding but reflects current FDA expectations and informs what reviewers will look for in submissions involving AI-generated content. The draft guidance was informed by over 500 AI-related regulatory submissions CDER reviewed between 2016 and 2023 [8], which means its expectations reflect observed industry failures, not theoretical ones.
The FDA-EMA joint guiding principles published January 14, 2026, extended this into a shared transatlantic framework covering the full medicinal product lifecycle from clinical trials through manufacturing and pharmacovigilance [2]. The 10 principles emphasize human-centric design, clear context of use, risk-based governance, data integrity, multidisciplinary oversight, and continuous lifecycle monitoring for performance drift [2]. While not legally binding, they signal what reviewers on both sides of the Atlantic will expect to see documented in future submissions.
ICH E6(R3), finalized January 6, 2025 and effective in the EU from July 23, 2025 [9], significantly expanded sponsor obligations around computerized system oversight. Under GCP Principle 10.1 and 10.2, sponsors retain overall responsibility for their activities even when those activities are transferred or delegated to service providers [9]. Section 4 of the same Guideline (Data Governance) sets out that both sponsors and investigators must maintain documented policies for data integrity, traceability, and security across all systems that handle trial data [9]. Delegating document generation to an AI vendor does not discharge these obligations.
The ISPE GAMP Guide: Artificial Intelligence, published July 2025, provides the most detailed operational framework [3]. The 290-page guide covers the full AI lifecycle from concept through retirement and establishes that third-party AI systems used in GxP contexts require rigorous supplier qualification. Sponsors cannot rely on a vendor's self-certification; documented audit rights, quality agreements, and technical validation evidence are expected [3].
These are the four primary sources a well-prepared evaluation committee should have read before a single vendor demonstration. A note on authority: 21 CFR Part 11 is binding federal regulation; FDA-2024-D-4689 and the FDA-EMA joint principles are non-binding guidance documents reflecting current agency expectations; and the ISPE GAMP AI Guide represents industry good practice consensus. All four are relevant to procurement decisions, but they carry different legal weight.
Criterion 1: Validation Architecture and Regulatory Documentation Readiness
The first question for any sponsor evaluation committee is not what the platform can generate but how it was built and validated for regulated use.
Under 21 CFR Part 11, electronic records created, modified, maintained, archived, retrieved, or transmitted to meet FDA predicate-rule record requirements, or submitted electronically to FDA, are subject to requirements for validated software, secure audit trails, access controls, and controlled time-stamped entry [10]. AI-generated regulatory documents that are created and managed in electronic form to satisfy these requirements fall directly within that scope. The GAMP AI Guide reinforces this, adding that AI-specific validation should address model performance qualification, test dataset documentation, and change management protocols when the underlying model is updated [3].
Sponsors should require vendors to produce a validation summary report covering user requirement specifications, functional specifications, and executed test protocols. For platforms that generate regulatory documents, the appropriate GAMP category, typically Category 4 (configurable) or Category 5 (custom or significantly configured), depends on the degree of vendor customization, intended GxP impact, and how the system is deployed in the sponsor's environment. Sponsors should confirm the categorization with the vendor during procurement and verify that the validation evidence matches that category [3]. If a vendor cannot provide a complete IQ/OQ/PQ evidence set, or if their validation documentation refers only to the hosting infrastructure rather than the AI model itself, that is a material gap.
Ask specifically whether the vendor has a change management protocol governing model updates. An AI platform updated with new training data or a new model version without a documented change control process can render prior validation evidence obsolete without notice. The ISPE GAMP AI Guide identifies continuous monitoring for data drift as a recommended element of lifecycle management under GAMP good practice [3], and the FDA-EMA joint principles echo this expectation [2].
Criterion 2: Hallucination Controls and Factual Accuracy Mechanisms
The hallucination problem in generative AI is not a theoretical risk in clinical document generation. It is a documented, reproducible failure mode. A 2024 peer-reviewed study in the Journal of Medical Internet Research by Chelli et al. measured hallucination rates directly: 39.6 percent for GPT-3.5, 28.6 percent for GPT-4, and 91.4 percent for Bard when each model was tasked with biomedical literature searches using standardized inclusion criteria [11]. A 2023 study by Walters and Wilder, published in Scientific Reports, prompted GPT-3.5 and GPT-4 to generate 636 references across 42 multidisciplinary topics and found fabrication rates of 55 percent and 18 percent respectively, establishing a quantitative baseline that has since been replicated across clinical and professional domains [12]. A 2025 arXiv preprint by Eser et al. (not peer-reviewed) evaluating AutoIND, a platform for generating IND nonclinical written summaries, found that AI-assisted drafting reduced initial composition time by approximately 97 percent but produced quality scores of 69 to 78 percent, with systematic deficiencies in emphasis, conciseness, and clarity; the authors concluded that expert regulatory writers remain essential for maturing outputs to submission-ready quality [13].
The problem is not that AI models are inherently unreliable but that, left unconstrained, they generate text with statistical plausibility rather than factual fidelity. In regulatory documents, those two properties are not equivalent.
When evaluating a platform, sponsors should ask three specific questions. First: does the platform generate output grounded strictly in sponsor-provided source documents, or can it draw on general training data to fill gaps? Platforms that use retrieval-augmented generation (RAG) architectures anchored to sponsor source files offer meaningfully lower hallucination risk in grounded tasks than those that rely on base model knowledge. Second: what validation check does the system apply to numerical claims, such as dosing, endpoints, and statistical results? A platform that copies a numerical value from a source table is categorically different from one that synthesizes a number from language context. Third: does the platform flag low-confidence outputs for human review rather than presenting all generated text with equivalent confidence? The latter behavior, common in general-purpose LLMs, is directly contrary to the FDA-EMA principle requiring transparency about AI system limitations [2].
Criterion 3: Cross-Document Consistency Management
Regulatory submissions are not collections of independent documents. The protocol, Investigator's Brochure (IB), Informed Consent Form (ICF), Development Safety Update Report (DSUR), and Clinical Study Report (CSR) must maintain factual consistency across shared elements: eligibility criteria, endpoints, dosing regimens, safety definitions, and statistical analysis plan references. When any of those elements change, all affected documents must be updated correspondingly.
Manual management of cross-document consistency is one of the most reliable sources of regulatory queries. A protocol amendment that is reflected in the CSR but not the ICF, or an efficacy endpoint described differently in the protocol versus the clinical overview, can generate reviewer queries that add weeks to submission timelines. The McKinsey regulatory benchmarking report from August 2025 identified cross-document data consistency as one of the core challenges for sponsors attempting to compress submission timelines from months to weeks [4].
Platforms that generate regulatory documents in isolation, where each document is an independent generation task with no persistent cross-document memory or controlled data model, cannot solve this problem architecturally. Sponsors should ask whether the platform maintains a shared data layer across document types, whether protocol-level parameters flow automatically into downstream documents, and whether a change in one document triggers a propagation check across linked documents. These are design questions with concrete, verifiable answers. A vendor that cannot explain the data architecture underlying cross-document consistency does not have a production-ready answer to the problem.
Criterion 4: Audit Trail and Data Governance
ICH E6(R3) Section 4 (Data Governance) sets out that both sponsors and investigators maintain documented policies for data integrity, traceability, and security [9]. For AI-generated regulatory documents, this means every generated document should have a traceable, timestamped record of what source data was used as input, which model version produced the output, what human review actions were taken, what edits were made, and by whom.
21 CFR Part 11 requires that audit trails for electronic records be computer-generated and include the date and time of any change, the identity of the person making the change, and a record of the prior entry [10]. Where an AI-generated document constitutes a regulated electronic record under Part 11 predicate rules, storing only the final approved version, without the generative history and review log, does not satisfy the audit trail requirements of 21 CFR Part 11 [10].
Sponsors should request a demonstration of the audit trail interface, not just a description of it. Key questions include: are source citations for each generated passage traceable back to the specific document and section the AI drew from? Is the human review step an enforced workflow gate or an optional feature that a user can bypass? Can the audit log be exported in a format suitable for regulatory inspection?
On data governance, sponsors must also ask how the platform handles sponsor data. Specifically: is sponsor trial data used to train or fine-tune the underlying model? If so, what are the data residency, segregation, and deletion guarantees? The FDA draft guidance recommends that AI systems used in regulated contexts have documented training dataset provenance [1]. A vendor that cannot produce that documentation for their core model is asking sponsors to accept an undocumented input into a regulated workflow.
Criterion 5: Human-in-the-Loop Design as an Architecture Commitment
The FDA-EMA joint principles state explicitly that AI in drug development should support ethical, human decision-making and not replace it [2]. ICH E6(R3) extends this in a more operational direction: under GCP Principle 10.2, responsibility for the conduct of the trial, including the quality and integrity of trial data, resides with the sponsor regardless of which tools or vendors generated the content [9]. The combination of these two expectations means that a sponsor cannot credibly argue in a regulatory context that they relied on an AI platform's output without applying substantive human expert review.
The relevant evaluation question is therefore not whether a vendor claims to support human oversight but whether that oversight is an architectural feature or a user preference. These are different things. A platform that presents AI-generated text in a seamless word-processor interface, where human edits and AI-generated content are visually indistinguishable, makes human oversight harder, not easier. A platform that presents AI-generated passages with explicit sourcing, confidence signals, and a structured review workflow makes human expertise the decisive step rather than a formality.
Sponsors should ask to see the review workflow, not the generation workflow. The generation speed is visible in any demo. The oversight architecture is not.
Criterion 6: Regulatory Content Scope and Depth
Not all AI writing platforms cover the same document types, and the clinical and regulatory requirements differ substantially across them. A platform optimized for CSR narrative generation may not have the structured data architecture needed for ICF development, where FDA guidance under 21 CFR 50 and the ICH E6(R3) requirements for participant-facing language, reading level, and IRB-traceable amendments apply [9]. An IB generation tool must handle the structure set out in ICH E6(R3) Appendix A, which defines the required content sections including summary, introduction, physical and chemical properties, nonclinical studies, and effects in humans [9]. DSUR generation requires direct integration with the sponsor's pharmacovigilance data and adherence to ICH E2F format specifications.
Sponsors should map each document type in their development pipeline against what the platform has validated it can produce. "We support all regulatory documents" is not an evaluation-ready answer. Request specific examples, with reference document types, format compliance evidence, and user acceptance testing reports for each document category the sponsor intends to use.
Protocol generation, to take the most consequential example, is a structured task that draws on multiple source inputs, including prior clinical data, competitive positioning analysis, and scientific rationale, and produces a document that must satisfy both scientific review by IRBs and regulatory review by FDA, EMA, or national competent authorities. A platform producing protocol first drafts without documented handling of the ICH M11 CeSHarP structure requirements [14], finalized by FDA in May 2026, is not current with agency expectations for protocol format.
Criterion 7: Sponsor Accountability Transfer and Quality Agreement Structure
Under both ICH E6(R3) and the GAMP AI Guide, the regulated company, the sponsor, retains ultimate responsibility for the quality of outsourced or AI-assisted activities [3],[9]. This principle has a direct contractual implication: sponsors need quality agreements with AI platform vendors that specify each party's obligations for data integrity, change notification, incident response, and ongoing validation maintenance.
Most enterprise software agreements are not quality agreements in the ICH/GCP sense. They describe service levels, liability limits, and data processing terms. Quality agreements for regulated software need to cover: the vendor's obligation to notify sponsors of model version changes before deployment; the vendor's document retention obligations for validation evidence; the sponsor's rights to audit the vendor's quality management system; and the protocol for managing compliance deviations.
Sponsors evaluating AI regulatory writing platforms should request a draft quality agreement as part of the procurement process, not as a post-contract deliverable. Vendors that decline to provide quality agreements on the grounds that their platform is a software tool rather than a regulated service are not fully accounting for how the platform will be used in a regulated context.
Red Flags in Vendor Demonstrations
A structured vendor evaluation requires specific artifacts, not just platform walkthroughs. Sponsors should treat any of the following as signals that a platform is not ready for regulated deployment.
No validation package available. A vendor that cannot produce user requirement specifications, functional design specs, and executed test protocols on request has not undergone regulated software validation. A summary slide deck is not a validation package.
Model updates deployed without change notification. Ask directly: how does the vendor notify customers before a model version change goes live? If the answer involves no pre-deployment notice, no re-validation trigger, and no customer opt-out, the vendor's change management does not meet GAMP AI Guide expectations [3].
No source traceability in generated output. A generated regulatory document should be traceable at the passage level to the specific source document and section the AI drew from. If the demo cannot show this, the platform cannot support inspection-ready audit trails under 21 CFR Part 11 [10].
No evidence of tenant isolation or data segregation controls. Multi-tenant architectures are not automatically disqualifying, but a vendor must provide clear evidence of logical data segregation, access controls, encryption at rest and in transit, and contractual guarantees that sponsor trial data is not accessible to other customers or used to train shared models. The absence of that documentation presents a data governance risk under ICH E6(R3) Section 4.3 [9] and HIPAA where applicable.
No quality agreement offered. A vendor that declines a quality agreement or treats it as a legal novelty has not operated in a GxP-regulated sponsor context before. The ISPE GAMP AI Guide is explicit that the regulated company retains responsibility for qualifying AI suppliers, which requires a documented quality relationship [3].
Demo shows only generation speed, not review workflow. Demos that emphasize how fast the first draft appears, without showing the review gates, approval workflow, and audit trail that follow, are optimizing the wrong end of the workflow for regulatory purposes.
A practical evidence request list. When a vendor demo goes well and procurement moves forward, sponsors should formally request these six artifacts before contract signature:
| Artifact | What to look for | Regulatory basis |
|---|---|---|
| Validation summary report | User requirements, functional specs, IQ/OQ/PQ protocols, executed test evidence | Good practiceBinding where applicable ISPE GAMP AI Guide; 21 CFR Part 11 (where predicate rules apply) |
| Source-traceability demonstration | Passage-level linkage from generated text to specific input document and section | Binding where applicable ICH E6(R3) Section 4 (data integrity); 21 CFR Part 11 audit trail |
| Model-update SOP | Pre-deployment notification timeline, re-validation trigger criteria, customer opt-out mechanism | Good practiceNon-binding guidance ISPE GAMP AI Guide; FDA-2024-D-4689 lifecycle management |
| Quality agreement draft | Validation obligations, change notification, audit rights, deviation protocol | Good practice ICH E6(R3) Principle 10.2 / GAMP AI Guide |
| Audit log sample export | Computer-generated timestamps, user identity, prior-entry record, exportable format | Binding where applicable 21 CFR Part 11 Section 11.10 |
| Data-use policy | Confirmation that sponsor trial data is not used for model training; data residency and deletion terms | Binding where applicable ICH E6(R3) Section 4.3; HIPAA (where applicable) |
A strong demo is not enough. Sponsors should request validation evidence, source traceability, model-update controls, quality agreement terms, audit log samples, and data-use documentation before contract signature.
Regulatory and Documentation Considerations
The current regulatory framework for AI in regulated clinical contexts, while still evolving, has moved well beyond the absence of guidance that characterized the early 2020s. The FDA's draft guidance FDA-2024-D-4689 introduced a five-step credibility assessment process: define the context of use, assess model risk, characterize model performance, evaluate data adequacy, and demonstrate ongoing lifecycle monitoring [1]. Sponsors deploying AI regulatory writing platforms should apply this framework at the platform level, not only at the level of individual submissions that draw on AI-generated content.
The ISPE GAMP AI Guide's risk categorization is also relevant here. AI systems with direct impact on patient safety data, submission content, or essential trial records are subject to the highest validation tier [3]. Regulatory document generation systems, which directly produce content that regulators review to make approval decisions, will likely fall into the higher validation tiers of that framework for most sponsor use cases, though the exact tier depends on system configuration, GxP impact, and intended use.
One practical implication: sponsors should conduct a platform impact assessment before full deployment, classifying the platform's outputs by document type and mapping each against the relevant regulatory standard. For documents that directly enter regulatory submissions, validation evidence requirements are substantially more demanding than for internal operational documents. A single validation package that treats a protocol template and an email draft equivalently does not reflect this difference and may not withstand inspection scrutiny.
AI and Automation Perspective
The efficiency case for AI-assisted regulatory writing is established and growing, but the compliance case is more complex. The primary risk is not that AI tools are slow to adopt, nor that they will be rejected by regulators in principle. Both FDA and EMA have signaled clear openness to AI in regulated drug development contexts [1],[2]. The risk is that sponsors deploy platforms with insufficient validation, insufficient sourcing controls, and insufficient human oversight, and then encounter performance failures in contexts where the cost of failure is high.
The human-AI collaboration model described in the emerging regulatory writing literature, where AI handles systematic compilation and first-draft synthesis while human experts provide substantive review and accountability, addresses this risk architecturally [13]. Platforms designed for this model, with structured generation workflows, traceable sourcing, enforced review gates, and documented validation histories, are fundamentally different from platforms that optimize for the appearance of speed at the generation stage without addressing the downstream accountability chain.
Continuous monitoring for data drift deserves specific mention. The FDA-EMA joint principles require that AI systems be monitored for performance changes over time as the underlying data environment evolves [2]. The GAMP AI Guide operationalizes this as an ongoing obligation, not a one-time validation event [3]. For sponsors using vendor-hosted AI platforms, this means verifying that the vendor has a documented monitoring program and a defined threshold for when sponsor notification or revalidation is required.
How Kitsa Fits Into This Problem
Kitsa built KScribe with the compliance architecture this article describes as an operational requirement rather than a post-launch feature. (See KScribe: AI Regulatory Document Generation for full product details.) According to Kitsa, the platform grounds document generation in sponsor-provided source materials, maintains a shared data layer across protocol, ICF, IB, DSUR, and CSR output types to support cross-document consistency, and produces a traceable audit record for every generation event. Kitsa states that KScribe operates within SOC 2 Type II and ISO 27001 certified infrastructure, with HIPAA-aligned controls and deployment within an AWS Virtual Private Cloud (VPC). Sponsors conducting formal procurement evaluations can request Kitsa's validation documentation package, quality agreement framework, and SOC 2 Type II and ISO 27001 certificates as part of due diligence, consistent with the GAMP AI Guide's expectation that regulated companies qualify their AI suppliers [3]. As with any AI platform, sponsors should verify these claims directly against vendor-provided evidence during procurement. Kitsa states that these materials, including validation documentation, security certifications, and quality agreement drafts, are available to sponsors as part of a formal due-diligence process. The review team should include sponsor QA, regulatory affairs, IT security, and legal functions, each evaluating the aspects of the platform that fall within their remit.
Sponsors evaluating AI regulatory writing platforms need more than fast first drafts. They need source-grounded generation, cross-document consistency, audit-ready records, validation evidence, human review workflows, and vendor documentation that can withstand QA and regulatory scrutiny. KScribe is designed for regulated clinical document generation across protocols, ICFs, IBs, DSURs, and CSRs, with sponsor due diligence materials available for formal procurement review.
Explore KScribeKey Takeaways
- The FDA's credibility assessment framework (FDA-2024-D-4689, January 2025) recommends that sponsors define context of use, assess model risk, document training data, and monitor AI performance throughout the system lifecycle; these expectations apply to third-party AI writing platforms deployed for regulated document generation.
- Tufts CSDD data shows that 75 to 76 percent of clinical trial protocols require at least one major amendment, with costs of $141,000 to $535,000 per amendment in Phase 3, which means document inconsistencies introduced by unvalidated AI tools carry concrete financial and operational consequences.
- LLM hallucination is a documented and quantified risk in regulated clinical contexts; sponsors should require vendors to demonstrate retrieval-grounded generation architectures, not simply cite general model accuracy benchmarks.
- Cross-document consistency management, covering shared parameters across protocols, IBs, ICFs, DSURs, and CSRs, is an architectural requirement for any AI platform used across the full regulatory document set, not a feature sponsors should implement manually.
- Under ICH E6(R3) GCP Principle 10.1 and 10.2, sponsors retain overall responsibility for AI-generated document quality even when the platform is vendor-hosted; quality agreements specifying validation obligations, change notification, and audit rights are contractually necessary.
- The ISPE GAMP Guide: Artificial Intelligence (July 2025) established that AI systems in GxP contexts require supplier qualification with documented audit rights, quality agreements, and technical validation evidence from the vendor.
- Human-in-the-loop design should be evaluated as an architectural feature: platforms with enforced review workflows and traceable generation histories support genuine oversight; those that simply allow human editing do not.
FAQ
What regulatory guidance currently applies to AI-generated regulatory documents?
What is the hallucination risk in clinical regulatory writing, and how should sponsors assess it?
Does using an AI writing platform transfer regulatory liability to the vendor?
How does 21 CFR Part 11 apply to AI-generated regulatory documents?
How should sponsors evaluate cross-document consistency in an AI regulatory writing platform?
What is the right validation documentation to request from an AI regulatory writing vendor?
References
- [1]U.S. Food and Drug Administration. "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products." Draft Guidance, FDA-2024-D-4689. January 7, 2025. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological
- [2]U.S. Food and Drug Administration and European Medicines Agency. "Guiding Principles of Good AI Practice in Drug Development." January 14, 2026. https://www.ema.europa.eu/en/news/ema-fda-set-common-principles-ai-medicine-development-0
- [3]International Society for Pharmaceutical Engineering (ISPE). "ISPE GAMP Guide: Artificial Intelligence." July 2025. https://ispe.org/publications/guidance-documents/gamp-guide-artificial-intelligence
- [4]Mihic, A., Agrawal, G., Berghauser Pont, L., et al. "Rewiring Pharma's Regulatory Submissions with AI and Zero-Based Design." McKinsey & Company. August 1, 2025. https://www.mckinsey.com/industries/life-sciences/our-insights/rewiring-pharmas-regulatory-submissions-with-ai-and-zero-based-design
- [5]McKinsey & Company. "With Gen AI, Merck and McKinsey Transform Clinical Authoring." McKinsey Blog. June 27, 2025. Also: "Merck Expands Innovative Internal Generative AI Solutions Helping to Deliver Medicines to Patients Faster." Merck Press Release. June 25, 2025. Note: Both [4] and [5] represent commercial benchmarking and implementation reporting, not peer-reviewed evidence.
- [6]Getz K, et al. "New Benchmarks on Protocol Amendment Practices, Trends and their Impact on Clinical Trial Performance." Therapeutic Innovation and Regulatory Science. 2024. PubMed PMID: 38438658. DOI: 10.1007/s43441-024-00622-9. https://link.springer.com/article/10.1007/s43441-024-00622-9
- [7]Getz KA, Stergiopoulos S, Short M, et al. "The Impact of Protocol Amendments on Clinical Trial Performance and Cost." Therapeutic Innovation and Regulatory Science. 2016;50(4):436-441. DOI: 10.1177/2168479016632271. PMID: 30227022. https://link.springer.com/article/10.1177/2168479016632271
- [8]U.S. Food and Drug Administration. "Artificial Intelligence for Drug Development." CDER. https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development
- [9]International Council for Harmonisation. "ICH E6(R3): Guideline for Good Clinical Practice." Final version adopted January 6, 2025. EMA effective July 23, 2025. FDA posted September 9, 2025. UK effective April 28, 2026. https://database.ich.org/sites/default/files/ICH_E6%28R3%29_Step4_FinalGuideline_2025_0106.pdf
- [10]U.S. Food and Drug Administration. "21 CFR Part 11: Electronic Records; Electronic Signatures." Code of Federal Regulations. https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- [11]Chelli M, Descamps J, Lavoue V, et al. "Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis." Journal of Medical Internet Research. 2024;26:e53164. DOI: 10.2196/53164. https://www.jmir.org/2024/1/e53164
- [12]Walters WH, Wilder EI. "Fabrication and errors in the bibliographic citations generated by ChatGPT." Scientific Reports. 2023;13(1):14045. DOI: 10.1038/s41598-023-41032-5. https://www.nature.com/articles/s41598-023-41032-5
- [13]Eser U, et al. "Human-AI Collaboration Increases Efficiency in Regulatory Writing." arXiv preprint. September 2025. Note: Preprint, not peer-reviewed. Evaluates AutoIND for IND nonclinical written summaries (eCTD modules 2.6.2, 2.6.4, 2.6.6). https://arxiv.org/abs/2509.09738
- [14]U.S. Food and Drug Administration. "M11 Clinical Electronic Structured Harmonised Protocol (CeSHarP)." Final guidance. Federal Register published May 22, 2026. https://www.federalregister.gov/documents/2026/05/22/2026-10295/m11-clinical-electronic-structured-harmonised-protocol-cesharp-international-council-for
