AIG-028 Hallucination and Factual Accuracy Controls
Description
Each generative use case whose outputs a user may rely on as factual carries a documented acceptable hallucination rate and a measured rate from an evaluation against a named dataset, recorded before deployment and repeated at defined intervals. At least one grounding or verification mechanism is active in the serving path for each such use case, for example retrieval with source citation, a post-processing verification step or an uncertainty score carried with the output. Use cases documented as safety-critical carry a human verification step before the output is acted on. A measured rate above the documented threshold carries a recorded remediation decision and a re-measurement result.
Rationale
Hallucination is a property of the model, not a defect to be patched, so the control is a measured rate against a threshold the organisation has set, not the presence of a named technique. That is also why the old wording failed: any two items from a list could be present while the rate went unmeasured. The mechanism list is illustrative and a use case may satisfy it in a way not named here. A disclaimer is a disclosure rather than a grounding mechanism and does not satisfy the second clause on its own; output disclosure to users is AIG-016. AIG-027 holds the behaviour required when the system is uncertain and AIG-008 the verification and validation gate.
Applicability (9 profiles)
The use cases, the acceptable rate and the measured rate on the deployer's own data are the deployer's. The grounding mechanism may be a provider feature or the deployer's own retrieval layer.
Art.15(1) holds a generative accuracy claim to the same consistency standard as any other performance property and Art.13(3) requires the performance metrics and the known foreseeable risks to health, safety and fundamental rights to be given to the deployer in the instructions for use. The measured hallucination rate and the threshold it is judged against therefore leave the organisation with the product. Neither article fixes a rate, which is why both mappings stay partial and why the documented threshold remains the provider's to set and to defend.
The use cases, the acceptable rate and the measured rate on the deployer's own data are the deployer's. The grounding mechanism may be a provider feature or the deployer's own retrieval layer.
Framework Mappings (17)
| AIS-10 | Output Validation | informative |
| EU-AI-Art.13.3 | Transparency — Mandatory Content of Instructions for Use | partial |
| EU-AI-Art.15.1 | Accuracy, Robustness and Cybersecurity — Performance Standards | partial |
| EU-AI-Art.15.4 | Accuracy, Robustness and Cybersecurity — Benchmarks and Measurement Methodologies | informative |
| EU-AI-Art.15.5 | Accuracy, Robustness and Cybersecurity — Declaration of Accuracy Levels and Metrics | informative |
| A.6.2.4 | AI system verification and validation | partial |
| AML.M0021 | Generative AI Guidelines | informative |
| AML.M0022 | Generative AI Model Alignment | informative |
| MG-4.1-002 | Post-Deployment AI System Monitoring | MG-4.1-002 | informative |
| MP-2.3-001 | Scientific Integrity and Testing Considerations | MP-2.3-001 | full |
| MP-2.3-003 | Scientific Integrity and Testing Considerations | MP-2.3-003 | full |
| MS-1.1-005 | AI Risk Measurement Approach Selection | MS-1.1-005 | informative |
| MS-2.5-003 | AI System Validity and Reliability | MS-2.5-003 | full |
| MS-2.5-005 | AI System Validity and Reliability | MS-2.5-005 | informative |
| MAP 2.2 | AI System Knowledge Limits Documentation | informative |
| MEASURE 2.5 | AI System Validity and Reliability | informative |
| LLM07 | Misinformation | full |
Evidence (2)
Serving configuration for each generative use case showing the grounding or verification mechanism active in the path, the documented acceptable hallucination rate and, for safety-critical use cases, the human verification step.
Example: LLM service configuration (AWS Bedrock Knowledge Bases + LangChain chain config): RAG enabled with source citation required, disclaimer appended to all responses not backed by retrieved sources (unverified_content_label: true), max_unverified_claims_per_response: 0 for medical use case, acceptable_hallucination_rate: <1% per monthly eval
Test: Request the serving configuration for each generative use case in scope. Verify: (1) at least one grounding or verification mechanism is active in the serving path and is reachable in the deployed configuration rather than only described, (2) the acceptable hallucination rate for the use case is recorded in or referenced by the configuration, (3) a use case documented as safety-critical has the human verification step configured before the output is acted on, (4) the configuration is version-controlled and carries a review date within the defined interval.
Hallucination evaluation report from testing showing measured hallucination rate for the LLM system against the documented acceptable threshold, using a defined evaluation dataset or benchmark.
Example: Hallucination Evaluation Report · Legal Research Assistant v2 (Confluence, 2026-02-01): evaluated on 500-question benchmark, hallucination rate 0.4% (threshold: <1%), RAG citation accuracy 97.8%, factual grounding score 0.94, result: PASS; evaluation dataset: internal legal QA benchmark v3
Test: Request the hallucination evaluation reports for the generative use cases in scope. Verify: (1) the evaluation dataset and methodology are named and the dataset is retrievable, (2) the measured rate is reported and compared against the documented acceptable rate for that use case, (3) an evaluation exists dated before deployment and a further evaluation within the defined interval, (4) a measured rate above the threshold carries a recorded remediation decision and a re-measurement result, (5) for a safety-critical use case the report also records the rate at which human verification changed the output.
Questions (2)
Is a measured hallucination rate recorded for each generative use case whose outputs users may rely on as factual?
The measurement is the control. A use case with grounding mechanisms in place but no measured rate against a documented threshold does not meet the first clause.
Which of the following are in place for your generative use cases that produce statements users may rely on as factual?
Every item applies to each use case in scope. The grounding mechanism may be retrieval with source citation, a post-processing verification step, an uncertainty score carried with the output or another mechanism that does the same job; the control does not require a named technique. A disclaimer on its own is a disclosure, not a grounding mechanism.