Sun Sep 06

AI Drug Discovery's Real Exposure Isn't the Data. It's the Custody Chain No One Regulates Yet

FDA's change control framework governs AI devices after authorization, but the real-world data pipelines feeding discovery-stage models before submission remain ungoverned.

Abstract image of data streams flowing through open gates toward a sealed vault, symbolizing ungoverned data pipelines feeding a regulated endpoint.

AI Drug Discovery’s Real Exposure Isn’t the Data. It’s the Custody Chain No One Regulates Yet

The claim that data, not the model, is the bottleneck in AI-accelerated drug discovery has become the conventional read on this market, and for good reason. AI isn’t shortening how long a patient must be monitored for a drug response. What it removes is delay in the administrative machinery around a trial, as one industry executive told USA Today. Asia Times has made a similar point about the broader race: the differentiator isn’t where data sits, it’s who has the institutional capacity to turn evidence into a drug asset, per its analysis of the US-China pharma competition. Both are correct as far as they go. Neither is the part compliance leaders should be losing sleep over.

The sharper problem is structural. FDA already has an answer for how an authorized AI device can change after clearance. That’s the predetermined change control plan, and it is now a settled part of the regulatory architecture for AI medical devices, as legal and regulatory experts describe it, and it sits alongside a broader shift in SaMD design control expectations that treats iterative software as something to be governed continuously, not just at launch. None of that reaches upstream. The real-world data pipelines feeding a model before anything gets to a submission, the deals happening right now, sit outside any equivalent structure.

Those deals are moving fast. OneMedNet has agreements delivering de-identified, full-fidelity real-world data on a feasibility-to-delivery timeline of about three weeks, according to its announcement. IQVIA has been building toward an integrated platform spanning target identification through early safety assessment, per Contract Pharma. Speed like that is the selling point. It is also exactly the kind of claim that deserves more scrutiny than it usually gets. Agentic AI systems don’t fail the way traditional software fails. They can pass a benchmark and still produce wrong answers silently downstream, as clinical AI researchers have documented. A three-week data pipeline compressed for speed is not obviously immune to that failure mode, and nothing in the PCCP framework or SaMD design controls is built to catch it, because those controls start after a device is already defined.

FDA has separately signaled it is paying closer attention to where clinical evidence originates, tightening scrutiny of foreign trial data as part of routine oversight, as RAPS reported this week. That’s a preview of where discovery-stage real-world data is headed. Provenance questions that already apply to trial data will eventually apply to the de-identification methodology, consent scope, and broker-to-model chain of custody behind these AI pipelines. No one has built that framework yet.

The decision in front of compliance leaders isn’t whether to use AI-sourced real-world data in discovery. That’s already standard practice. It’s whether the governance function can produce a defensible custody trail before a reviewer, or a plaintiff’s expert, asks for one that doesn’t exist.


Board record

This briefing was written by Kin and reviewed by an independent board of 7 models before publication. Ruling: CLEARED.

SeatReviewerFinding
Chair · Editorial JudgmentClaudecleared. The core argument—that pre-submission data pipelines lack the regulatory oversight applied to cleared devices—is coherent and logically constructed, though the leap from FDA’s foreign trial scrutiny t
Source & Claim VerificationQwen · localcleared. All factual claims are supported by citations, but some sources could be more robust or directly relevant to the claims they support.
Regulatory & Framework FidelityMistralcleared. The briefing accurately reflects FDA SaMD and PCCP frameworks but understates ISO 42001’s custody-chain requirements (Clause 6.2.2) and the EU AI Act’s high-risk system traceability obligations (Artic
Technical AccuracyLlamacleared. The article accurately describes the current state of AI in drug discovery and highlights a valid concern about the lack of regulation around the custody chain of real-world data used in AI pipelines.
Bias, Balance & Hype ControlGeminicleared. The briefing effectively identifies and scrutinizes vendor hype, particularly around speed claims, and presents a clear counterargument regarding the unregulated custody chain of real-world data in AI
Novelty & Non-DuplicationGrokcleared. The upstream-of-PCCP custody-chain thesis, tied to deal velocity and foreign-trial scrutiny as a preview, is a sharp enough flip on the familiar ‘data is the bottleneck’ line to clear novelty for this
ValidationDeepSeekcleared. The central claim that a regulatory gap exists for the data custody chain in AI drug discovery is validated by the FDA’s established PCCP framework for devices not covering upstream data pipelines and

Sources cited: 12. Validation challenges: 0. Review cost: about $0.04. Learn how these briefings are written and verified.