[
Blog Post
|
September 30, 2026
]

Can You Trust AI With Your CMC Data?

TL;DR

  • AI can pass every demo and still return a supplier impact list that is correct but incomplete, with no warning.
  • The cause is structural: CMC relationships between supplier, material, equipment, and process are implied in documents, and smarter models cannot certify completeness.
  • QbDVision measures AI on two tracks, Import (extraction) and Insights (recall), against expert-curated ground truth.
  • Measure before you deploy: define acceptance criteria, validate on your data, then monitor continuously.

"We have an AI system that extracts data from our manufacturing documents and answers questions about our drug formulations. Should we trust it?"

We hear some version of that question every day. It's a fair one, and the honest answer starts with how these systems are judged.

It often starts with a demo. Someone hands the system a clean manufacturing specification and asks for the target particle size. The answer comes back in seconds, every time. Your team gets excited, your executives see the potential; the budget gets approved, and the system goes live.

Then a real question surfaces.

A quality leader has just heard that Supplier X received a warning letter, and she needs to know which of your materials come from them. She needs more than the three closest matches. She needs every affected material, plus every document that mentions the supplier relationship, specifies where a material comes from, or references equipment sourced from Supplier X.

The system answers quickly. The list looks complete. It reads as complete.

But seven materials are affected and the list has three.

Nothing on the screen says anything is missing. There is no count to check against and no warning. Four risks leave the room with a quality leader who believes she has the full picture. In CMC, that is a compliance gap, and a wrong answer that looks right is worse than no answer at all.

[
Supplier impact query
]
"Which of our materials come from Supplier X?"
What exists
7 affected material records
What the system returned
3 records, all correct. 4 missing, with no warning.
100%
Of returned answers correct
43%
Of affected materials found
0
Warnings that anything is missing
Fail
For a regulated workflow
A wrong answer that looks right is worse than no answer.

Illustrative example.

Why the System Missed Them

The obvious suspect is the model, but the model is only part of the story.

Most AI performance claims describe the model: how it scores on scientific reasoning benchmarks, how well it answers questions. They say little about whether a system can do the specific work CMC requires. A model can score well and still miss a process parameter, lose the link between a unit operation and its steps, or return three correct materials when seven exist. In CMC, correct and complete are different measures. So are accurate and trustworthy.

The deeper cause sits in the data. Anthropic ran into the same wall studying why AI agents for biology are advancing more slowly than expected, and described it this way:

"Using AI agents to navigate biological data infrastructure is like driving through an old city that was designed before cars: the infrastructure may be beautiful and even thoughtful, but it's full of narrow, winding streets that are difficult for modern vehicles to navigate."

CMC has the same streets. Manufacturing knowledge comes from decades of systems: digital records, scanned paper, handwritten notes. It lives in separate databases. Sourcing information is scattered across specifications, procurement records, and supplier documents, and changes are recorded in dozens of different ways. The relationship our quality leader needed, from supplier to material to equipment to process, was never written down as a relationship. It was implied, spread across narrative text in documents built for people to read.

[
Relationships CMC depends on
]
In most CMC documents todayImplicit, buried in narrative text across specifications, procurement records, and supplier documents.
What AI needsExplicit, machine-readable relationships it can traverse and verify.

The tempting response is to wait for smarter models, and frontier models are improving. They iterate: they ask, read, refine, and recover more than single-shot retrieval ever could. They will find more of the seven. Two things stay out of their reach. A model cannot certify that it found everything, and it cannot tell you whether 85% is acceptable or dangerously inadequate for your workflow.

That is why each frontier advance, from RAG to GraphRAG to agents, has moved toward structure. Structure is where the question "Can I trust this?" gets answered, and it's where QbDVision started.

Validate, Yes. But Against What?

If smarter models can't settle the question, guidance is the next place to look. The FDA's January 2025 draft guidance on AI for regulatory decision-making asks sponsors to describe the performance metrics used to evaluate a model, to "specify how data independence was achieved between development (training and tuning data) and test data," and to document the model's limitations. Data specialists point the same way. IntuitionLabs' benchmark analysis of pharma document AI found typical accuracy of 80 to 95% on mixed pharma documents and concluded that "custom domain-aware pipelines are essential."

The direction is clear: generic benchmarks cannot tell you whether AI works in CMC. What the industry still lacks is a widely adopted, published standard for what CMC-specific measurement looks like.

So the practical questions stay open. A system reports 90% accuracy on batch parameters: accuracy against clean data, or messy real-world data? A system finds supplier relationships in 8 of 10 documents: did it miss any, and is 80% good enough for this use? And for our quality leader: how do you know a list is complete when nobody holds the complete list?

ISPE provides validation principles. Vendors report accuracy. Neither defines what good looks like in CMC. Measurement does.

How She Would Have Known

Start by trading "Is it accurate?" for a better question: "Can I trust it for this specific job?" In CMC, AI does two different jobs, and each fails in its own way, so we measure them on two tracks.

[
Two-track measurement framework
]

Track 1

Import

Structured data extraction
AccuracyDoes it extract the correct values?
FidelityDoes it preserve relationships and context?
TraceabilityCan every value be traced to its source document?
CompletenessDoes it find every material from a supplier?

Track 2

Insights

Recall and knowledge Q&A
RecallCan it find relevant information across the corpus?
AccuracyIs the information correct?
CompletenessDid it find every relevant document, or only the top few?
ReasoningCan you see why it returned these documents?

Import is extraction: the AI pulls data from manufacturing documents into your database. Accuracy is the entry point. Fidelity asks whether a value kept its context, so a 1000 L batch size that belongs to Process A variant 2 stays attached to that variant and only that variant. Traceability asks whether every value leads back to its source document for audit. A system can extract batch parameters at 92% accuracy and still fail on both, which is why fidelity and traceability carry the most weight here. A system at 88% accuracy with complete traceability can be the stronger choice over one at 95% with questionable sources, because regulators expect every value to trace to its source.

Insights is recall: the AI answers questions about manufacturing history, regulatory relationships, and technical rationale. Here the gap is between "documents that mention Supplier X" and "all documents that mention Supplier X," so recall and completeness carry the most weight. This is the track that catches our quality leader's problem. Graded on correctness alone, three correct materials scores 100%. Graded for completeness against expert-curated ground truth, it scores 43%, and it fails. For a supplier impact question, finding 80% of affected materials can be more dangerous than finding none, because it creates false confidence.

Two more dimensions apply across both tracks: regulatory compliance, meaning whether outputs meet FDA data integrity expectations, and stress resilience, meaning how the system holds up against OCR errors, handwriting, outdated formats, and missing data. Real documents are rarely as clean as the demo.

We hold our own AI to this standard. QbDVision's Qurio AI capabilities are evaluated on both tracks against domain-expert-curated ground truth.

[
The framework in practice: Qurio evaluation results
]
225+
evaluation tasks: 125 Import (extraction) and 102 Insights (recall), measured against domain-expert-curated CMC ground truth.
97.5%
Import completeness
Expected information identified from source documents.
90.7%
Import correctness
Extracted values correct against ground truth.
94.7%
Insights correctness
Retrieved answers correct against ground truth.
92.7%
Average correctness
Mean of Import and Insights correctness.

Results as of August 18, 2026.

Measure Before You Deploy

Look again at the order of events in that opening demo: deploy, hope it works, find problems, patch, and eventually, maybe, arrive at a validated system. Measurement reverses that order.

[
Flip the timeline
]
The common path
  1. Deploy
  2. Hope it works
  3. Find problems
  4. Patch
  5. Validated system, eventually (maybe)
The measurement-first path
  1. Define metrics
  2. Validate rigorously
  3. Deploy with evidence
  4. Monitor continuously
  5. Evidence ready for regulatory review

Measuring first puts the work up front and avoids the rework. It also removes the most expensive failure mode: a system that looks like it's working when it isn't. If you're evaluating AI for CMC, three steps make that concrete.

  1. Ask vendors for CMC-specific evidence. MMLU scores and general extraction benchmarks will not answer the question. Ask for performance on your document types, your data structures, and your regulatory requirements; performance under stress, such as OCR errors and missing data; and traceability evidence. A vendor that cannot provide this has not validated in CMC.
  2. Define acceptance criteria before you deploy. Assess the risk of each use case, since high-risk uses demand stricter criteria, and document what "good enough" means for each workflow. That record becomes your validation evidence.
  3. Keep measuring. Performance drifts, and new document formats appear. Track whether accuracy holds steady, when performance degrades enough to require retraining, when people override the AI and why, and whether the system is reducing cycle time as expected.

Back to Supplier X

Picture the same call with measurement in place. The system still has to find seven materials across decades of documents. This time, the quality leader knows how it performs on completeness for exactly this kind of question, what threshold her team set before it went live, and what evidence sits behind that number. She knows how far to trust the list, and when to ask for more.

The timing matters. FDA expectations for AI credibility are taking shape, failed pilots are expensive, and boards now ask for evidence before they fund. The organizations that publish CMC AI validation frameworks first will shape the standard that others cite in RFPs and industry guidance.

"Can AI be trusted in CMC?" has a real answer. It comes from measurement, and our white paper lays out how we do it.

[
White paper
]
Measuring AI Performance in CMC

Defined CMC tasks, task-specific benchmarks, ground-truth evaluation, performance by task, model, and cost, and acceptance criteria based on intended use and risk.

[

Chief Marketing Officer

]

Luke Guerrero

Luke is an experienced technology leader with a background in growing teams, building software products, and running business operations. As Chief Marketing Officer at QbDVision, he leads marketing and brings operational and cross-industry experience to the company's brand and go-to-market.

HubSpot form (set the "Hubspot form ID" property) renders here on the published site.
[
Blog Post
 I
August 10, 2026
]

Scalable AI for CMC: The Role of Credibility

Artificial Intelligence
[
Blog Post
 I
June 8, 2026
]

Building Scalable AI for CMC

Artificial Intelligence
[
Blog Post
 I
January 26, 2026
]

JPM26 takeaways: The race to AI meets the reality of CMC data maturity

Artificial Intelligence
[
Blog Post
 I
September 30, 2026
]

Can You Trust AI With Your CMC Data?

Artificial Intelligence
[
Customer Story
 I
November 8, 2023
]

From Paper to 3X Faster Technology Transfers

Tech Transfer
[
Customer Story
 I
January 8, 2024
]

Streamlining CMC Workflow for Xeris Biopharma

Tech Transfer

Ready to build scalable AI for CMC?