Scalable AI for CMC: The Role of Credibility
Most of us use navigation systems without thinking twice. We type a destination, follow the recommended route, and expect to arrive safely. These systems combine GPS with maps, traffic data, road closures, and other information to guide us. But if a navigation system directed you onto a closed road, through a lake, or into oncoming traffic, our confidence in it would disappear instantly. Confidence does not come simply because a system provides an answer. It comes from understanding how that answer was produced, being able to verify the recommendation, recognizing the system’s limitations, and knowing when human judgment should take over. These are not simply matters of trust; they are outcomes of effective governance.
AI in pharmaceutical development is no different. Organizations need confidence that AI operates within defined boundaries, that its outputs can be evaluated, and that qualified people remain accountable for critical decisions. That confidence depends on how the AI is governed.
In our previous article, Building Scalable AI for CMC, we discussed what it takes to build AI that can scale in regulated environments. We introduced four foundational pillars: Structured, Governed, Sustainable, and Trusted. Together, these pillars create the foundation required to deliver repeatable, reliable, traceable, and compliant AI outcomes in Chemistry, Manufacturing, and Controls (CMC). This article focuses on the Governed pillar: the controls, evaluation methods, and evidence used to establish AI credibility.
With that foundation in place, we establish that managing credibility is central to governance, and thus a new question emerges: “How do you establish that an AI capability is appropriate for its intended use?” For decades, the standard answer for GxP software systems was validation: documented evidence that a system consistently performs as intended and meets predefined requirements. Validation remains essential, but AI changes the conversation.
As non-deterministic systems such as large language models (LLMs) become more common in regulated environments, organizations need a way to demonstrate not only that a system functions as intended, but that its outputs are appropriate, reliable, and defensible for their intended use.
In a regulated GxP environment, even when there is a human in the loop, an unreliable AI output is not just a minor technical glitch. An answer may contain fabricated information, provide an incorrect result, or omit expected content. In CMC workflows, these failures can contribute to errors in a batch record, derail a tech transfer, delay a drug launch, potentially costing an enterprise $1–5 million in unrealized revenue for each day of delay, or result in a significant regulatory inspection finding. Building scalable AI for CMC therefore requires more than a system that simply “gives an answer.” It requires governed AI: a framework of controls, evaluation methods, and evidence used to determine whether outputs meet defined credibility thresholds and are reliable, traceable, compliant, and appropriate for their intended use. By operationalizing credibility in this way, the Governed pillar provides the foundation for AI that can scale across regulated CMC environments.
Validation Was Built for Deterministic Systems
Traditional software validation is built on a simple principle: Expected behavior = Actual behavior.
A user performs an action, the system produces a result, and that result can be compared directly against a predefined expectation. If the behavior matches, the test passes. If it doesn’t, the test fails.
Many aspects of AI systems can still be evaluated this way. Access controls, security permissions, audit trails, and data boundaries remain deterministic and should continue to be validated.
But AI-generated outputs introduce a different challenge. An AI model can have fixed weights and remain unchanged during use while still producing probabilistic or variable outputs. The same prompt may produce slightly different outputs, and there may be multiple acceptable ways to answer the same question. In these situations, the goal is no longer to determine whether an output is identical to an expected result. The goal is to determine whether the output is appropriate for its intended use.
This is where credibility assessment comes in. Rather than asking, “Did the system produce the exact expected answer?”, credibility asks a different question:
“Did the system produce an answer that meets a predefined standard of quality, reliability, and fitness for its purpose?”
Validation remains essential. However, how traditional validation practices should be applied to probabilistic AI outputs is still evolving. Emerging FDA thinking points to credibility assessment as a complementary, risk-based approach for evaluating whether an AI capability performs appropriately for its defined intended use.
From Validation to Credibility Assessment
The distinction between validation and credibility may seem subtle, but it is foundational. Validation remains essential for deterministic system behaviors such as security controls, user permissions, audit trails, and data boundaries. But when it comes to non-deterministic AI outputs, organizations need a way to assess whether a system’s responses are appropriate, reliable, and fit for their intended use.
Rather than replacing validation, credibility assessment extends it by providing a structured, risk-based framework for evaluating probabilistic AI outputs.
At QbDVision, we approach this challenge through a framework centered around credibility assessment.
Introducing CAMP: QbDVision's Framework for AI Credibility
At QbDVision, we believe organizations need a structured and repeatable way to assess whether AI systems are appropriate for their intended use. To operationalize this, we’ve developed the Credibility Assessment Master Plan (CAMP): a practical governance framework heavily informed by the FDA’s guidance on AI model credibility assessment and risk-based evaluation.
CAMP is not intended to replace established validation, engineering, or risk-management frameworks. It translates the risk-based credibility concepts reflected in FDA guidance into a practical method for evaluating AI capabilities used in CMC, while complementing GAMP 5’s risk-based, fit-for-intended-use approach to computerized systems. Its specific focus is the additional evidence needed to establish whether probabilistic AI outputs are credible for a defined CMC use case.
The CAMP serves as the governance framework for AI credibility, much like a Validation Master Plan serves as the governance framework for software validation. It establishes the standards, methodologies (including risk assessment), acceptance criteria, and oversight processes used to evaluate AI systems throughout their lifecycle.
Individual AI capabilities are then assessed through feature-specific Credibility Assessment Plans (CAPs) that apply the CAMP framework to a particular use case.
In simple terms:
- CAMP defines how credibility is established and maintained across AI systems.
- CAP applies that framework to a specific AI capability through a structured evaluation process.
Why AI Credibility Requires Engineering, Quality, and CMC Expertise
Technical benchmarks and model-evaluation tools answer only one part of the question: how well a model performs on a defined test. Regulated Quality and CMC teams must also determine whether that performance is acceptable for a specific intended use, risk profile, and operating environment. CAMP was developed at this intersection. It combines AI engineering and evaluation methods, including representative test scenarios, task-specific metrics, uncertainty assessment, and performance monitoring, with CMC, GxP, Quality, and validation expertise. This allows model-level evidence to be translated into intended-use boundaries, risk-based acceptance criteria, human-accountability controls, release decisions, and lifecycle governance.
Together, the CAMP and CAP provide a structured, repeatable approach for assessing, monitoring, and governing AI credibility throughout the system lifecycle.
Traditional Software | AI Systems |
Validation Master Plan | CAMP |
Validation Plan | CAP |
Deterministic testing | Probabilistic evaluation |
Why Now: AI Governance Is Moving from Principle to Practice
Unless you’ve been living under a rock, you’ve noticed that interest in AI in pharma is accelerating, and regulatory expectations for AI in pharmaceutical development are taking shape quickly. In January 2025, the FDA issued its draft guidance on the use of AI to support regulatory decision-making for drug and biological products, introducing a risk-based credibility assessment framework. Since then, the EU Artificial Intelligence Act has continued its phased application, the European Commission has proposed draft EU GMP Annex 22 on Artificial Intelligence, and the FDA and EMA have issued joint Guiding Principles of Good AI Practice in Drug Development. The FDA’s draft AI guidance introduces credibility assessment for AI used in regulatory decision-making, while the FDA–EMA Good AI Practice principles reinforce similar expectations across drug development. The EU AI Act adds a broader risk-based governance framework. Draft EU GMP Annex 22 does not directly cover generative AI or LLMs in critical GMP applications, but several of its underlying principles, including defined requirements, risk-based controls, human review, performance criteria, and ongoing monitoring, remain relevant to the broader evolution of AI governance in regulated environments.
While not all of the above guidance is directly applicable to systems like QbDVision, the frameworks are used to inform industry best practice for many different types of AI use in drug development. Teams need practical ways to govern AI across the systems landscape.
At the center of this shift is credibility: documented, risk-based evidence that an AI capability performs appropriately for its defined intended use. Credibility complements traditional validation by extending evaluation beyond deterministic system controls to the quality, reliability, and consistency of AI-generated outputs. At QbDVision, we operationalize this approach through CAMP and feature-specific CAPs.
QbDVision customers can use this evidence as an input to their own risk assessment and validation strategy, while responsibility for the final validation approach remains with each organization.
Step 1: Start with Intended Use
Every credibility assessment begins with a question: “What is this AI feature intended to do, and not do?” Credibility is not determined by the model itself, but by the context in which it is planned to be used. Before an AI capability can be assessed, organizations must clearly define its intended use, expected user interactions, potential misuse scenarios, and explicit boundaries. For example, an AI feature may be designed to summarize information or recommend content, but not make GxP decisions. This principle aligns closely with FDA guidance: credibility is established relative to intended use, not model architecture.
Step 2: Translate Intended Use into Evaluation
Once the intended use is established, it must be translated into a measurable evaluation strategy. At QbDVision, this happens through a feature-specific Credibility Assessment Plan (CAP), which includes an evaluation suite designed to test how an AI capability performs in realistic conditions. Rather than relying on traditional test scripts, we evaluate AI against real-world CMC scenarios, running a multitude of tasks to assess performance, including consistency and reliability. This allows us to measure not only whether the AI works, but whether it is credible enough to support its intended purpose.
Step 3: Measure What Matters
AI systems are commonly evaluated through task-specific benchmarks. For example, foundation models may be assessed using benchmarks such as SWE-bench, which measures performance against representative, real-world software engineering tasks rather than relying on a general assessment of model quality.
The same principle applies to AI in CMC: credibility must be measured against the tasks the capability is intended to perform. An AI-powered search experience, for example, requires different evaluation criteria than a content-generation or recommendation capability.
Within QbDVision’s framework, the CAMP establishes the overarching standards for evaluation, while each feature-specific CAP defines the relevant benchmark suite: representative CMC scenarios, expected assertions, grading criteria, performance metrics, and acceptance thresholds.
Depending on the capability and its risk profile, evaluation criteria may include accuracy, completeness, hallucination rate, formatting compliance, regulatory alignment, and language quality. Each output is assessed against predefined criteria, producing measurable evidence of whether the capability performs appropriately for its intended use.
Over time, these results establish credibility baselines, support comparisons across releases, identify performance drift, and provide evidence that the capability continues to meet its defined acceptance criteria.
Step 4: Establish Statistical Confidence
Unlike traditional software validation, credibility is not established through a single pass/fail test. It is established through evidence gathered across many evaluations.Â
The level of evidence required depends on the risk profile and business impact of the AI capability. Higher-risk use cases may require more extensive evaluation, stricter acceptance thresholds, and greater oversight than lower-risk use cases. The appropriate level of scrutiny should be determined through a risk-based justification that considers intended use, potential impact, and the consequences of an incorrect or misleading output.
By assessing performance across a representative set of real-world scenarios, organizations can establish confidence that an AI system performs reliably, consistently, and appropriately for its intended use, not just under controlled test conditions.
In other words, not every AI capability requires the same level of evidence. Credibility, like validation, should be commensurate with risk.
The Critical Principle: Human-in-the-Loop
Human oversight should be defined upstream, beginning with the intended use of each AI capability and carried through its design, evaluation, release, and operation. The role of the AI, the decisions reserved for qualified personnel, and the conditions requiring review or escalation should be established before the capability is deployed.
In GxP workflows, AI may support information retrieval, pattern identification, content generation, and recommendations, but it should not assume accountability for regulated decisions. That accountability remains with appropriately qualified personnel operating within defined controls.
This approach should be embedded in the organization’s AI governance and software development lifecycle. By making review responsibilities, approval points, and escalation paths explicit, human oversight becomes an operational control rather than a general expectation. This approach operationalizes the FDA–EMA principles of human-centric, risk-based oversight and aligns with the EU AI Act’s expectation that human oversight be proportionate to risk and enable users to understand system limitations, monitor performance, and disregard or override an AI-generated output when appropriate.
What Comes Next
AI adoption in pharmaceutical development and manufacturing is accelerating, but confidence cannot be assumed. It must be established and maintained. It must be overall governed.
AI in pharma requires validation and credibility assessments working together to provide evidence that AI systems are appropriate for their intended use.
At QbDVision, we apply this approach to individual AI capabilities. For Import by Qurio, a powerful capability that turns PDF knowledge into structured data, we assess outputs against curated ground-truth data across dimensions including completeness and correctness (ex.missing or partial content, hallucination incidence), and consistency across repeated trials.Â
For each release, the designated evaluation run is assessed against feature-specific acceptance criteria based on the capability’s intended use and risk profile, informing the go/no-go decision for that release. Results from repeated runs are also tracked over time to identify performance changes, emerging trends, and potential drift.
In an upcoming research paper, we will examine how structured, intentional AI architecture influences performance across measures such as reliability, accuracy, reproducibility, and performance across models. The paper will compare engineered, schema-driven approaches with more general-purpose LLM use and show how AI performance can be measured for industrial CMC workflows.
At QbDVision, we’re helping define what that future looks like through four pillars of scalable AI: Structured, Governed, Sustainable, and Trusted. Together, they provide the foundation, controls, architecture, and evidence required to deploy AI responsibly across regulated CMC environments.
GET IN TOUCH
Ready to Build Scalable AI for CMC?
Learn how QbDVision combines structured CMC knowledge, governance, human oversight, and credibility assessment to help organizations adopt AI with confidence.

