Skip to main content

Compliance

From Annex 22 to Agentic AI: A Practical Language for Regulated AI

A practical language for distinguishing predictive, generative, and agentic AI, and matching testing and oversight to the role AI plays in regulated work.

Published 2026-09-10QikSolve

The draft EU GMP Annex 22 consultation guideline has made one question more visible for regulated organisations: what kind of artificial intelligence is being used, and what does it do in the process?

That question matters because the same word, AI, can describe a model that predicts an outcome, a tool that drafts content, or a system that coordinates actions. Those uses do not create the same testing, review, evidence, or change-control questions.

This article proposes a practical language for discussing the difference. It is a planning model for regulated AI, not a regulatory determination, validation protocol, or legal advice. The organisation's intended use, process impact, jurisdiction, and quality system remain the controlling context.

What the draft Annex 22 says about scope

The draft Annex 22 consultation guideline addresses certain AI use cases in GMP environments. It distinguishes the AI and machine-learning applications within its scope from generative AI and large language models, and states that generative AI and LLMs should not be used in critical GMP applications. It also describes human responsibility for reviewing outputs in non-critical uses where those models are used.

Read the draft EU GMP Annex 22 consultation guideline for the source text and its defined context. The draft is not a blanket answer to every future AI use case. A team still needs to describe the intended use and determine what the output does next.

A useful language: predict, generate, act

Technology labels can obscure the practical control question. A simpler starting point is to classify the role of the output in the process.

Predictive AI predicts data

Predictive AI produces classifications, measurements, probabilities, or predictions. Examples may include defect detection, process anomaly detection, predictive maintenance, yield prediction, or process trend analysis.

The central question is usually whether the prediction performs as intended against suitable data and acceptance criteria. Accuracy, precision, sensitivity, specificity, and false-positive or false-negative rates may be relevant, depending on the use case.

This does not make predictive AI automatically suitable for a GMP process. The data, intended use, decision impact, review, and ongoing control still need to be assessed.

Generative AI generates content

Generative AI produces text, explanations, summaries, or other content. Examples may include drafting a deviation summary, organising an investigation, suggesting an SOP structure, or explaining a trend for a person to review.

There may not be one exact correct answer. Evaluation therefore asks whether the output is accurate, complete, grounded in the available evidence, consistent enough for the intended use, and acceptable to the accountable reviewer. The original context and the review disposition may also need to remain understandable and retrievable.

Human review is not a substitute for defining the review. The team should specify who reviews the output, what evidence they check, what acceptance criteria apply, and what happens when the output is incomplete, uncertain, or wrong.

Agentic AI coordinates actions

Agentic AI uses content generation as part of a larger process involving planning, orchestration, tool use, or action. An agent might retrieve records, search procedures, prepare findings, identify missing information, recommend escalation, or create a task for a person to review.

The important change is that the output is no longer only an answer. It can influence what happens next. The governance question becomes whether the agent followed the intended workflow, used approved information, avoided prohibited actions, escalated when required, and obtained human approval at the correct point.

Agentic AI therefore combines two kinds of evaluation: generative-AI evaluation for the quality and grounding of content, and workflow or process testing for sequencing, permissions, exceptions, escalation, and evidence.

Why the testing approach changes

The question should move from "Was the model answer correct?" to "Did the AI-assisted process behave as intended for this use case?”

AI rolePrimary outputPractical testing emphasis
Predictive AIPrediction or classificationPerformance against suitable data and acceptance criteria
Generative AIContent or explanationQuality, grounding, completeness, consistency, and human review
Agentic AIAction or workflow progressionProcess sequence, permissions, escalation, human approval, and traceability

This table is a planning aid, not a universal validation classification. A single solution may contain more than one role, and the same model may require a different control response when its output is used in a different process.

Start with process impact

The tool does not set the control boundary. Process impact does.

Ask five questions before selecting a governance response:

  1. What is the AI allowed to do, and what must it never be relied on to do?
  2. What information does it receive, and which sources are approved?
  3. Where does its output go next: a draft, a workflow, a decision, an approval, a release, or a controlled record?
  4. Who is accountable for reviewing, accepting, rejecting, reworking, or escalating the output?
  5. What evidence will show what happened, which version or configuration was involved, and what changed later?

These questions connect the AI use case to existing quality-system disciplines rather than creating a parallel governance vocabulary.

From model language to operating language

For a quality team, the three roles can be understood through familiar work:

  • A monitoring or statistical tool that identifies an emerging trend resembles predictive AI.
  • A quality professional using assistance to prepare a draft resembles generative AI.
  • A quality professional coordinating an investigation, assigning actions, and managing escalation resembles the process role of agentic AI.

The analogy is not a substitute for assessing the actual system. It helps people ask where accountability, review, evidence, and permission belong as AI moves from describing information to influencing action.

Continue the decision

The Practical, Governed AI for GMP Quality Operations hub connects this language to use-case selection, output boundaries, human oversight, evidence and traceability, evaluation, controlled change, and scaling through existing quality systems.

Relevant next questions include:

For platform and implementation depth, the hub refers visitors to the QxAIOS solution. For a discussion of a defined use case and operating context, contact QikSolve.