AI assurance / Technical risk guide
LLM and RAG risk.
Connect system failures to controls, tests and release decisions. A practical guide to evaluating retrieval-augmented generation, protecting access boundaries and producing evidence that GRC and MLOps teams can use together.
01 / Scope
Review the application and its permitted use.
A large language model (LLM) generates text; retrieval-augmented generation (RAG) supplies retrieved material as context for that generation. Retrieval may use keywords, vectors or a hybrid approach. The original RAG study by Lewis et al. (2020) demonstrated benefits on specific knowledge-intensive tasks. It does not establish the reliability or security of a different deployment.
A useful answer still depends on the right evidence reaching the model, the evidence being current and authorized, and the response using it correctly. A citation can point to a real document and still fail to support a claim. A grounded answer can faithfully repeat an incorrect source.
- GRC decisions
- Define permitted tasks, affected users, prohibited uses, risk ownership and the authority to accept residual risk. Identify applicable legal and sector requirements separately.
- MLOps decisions
- Identify the deployed model, retrieval configuration, corpus, identity path, tools and failure handling. Agree on observable tests and release criteria with the risk owner.
- Shared scope
- Record whether the system retrieves, answers, recommends or acts. A policy assistant and a tool-enabled assistant that changes customer records need different controls.
The controls and test patterns below are InfoSecured’s practical synthesis. NIST AI 600-1 is voluntary, cross-sector guidance; OWASP identifies security risks. Neither provides automatic regulatory compliance or an application certification.
02 / Architecture
Make each trust boundary explicit.
Start with ingestion: retain source provenance, effective dates, document versions and access-control metadata through parsing, chunking and indexing. Treat retrieved text, uploaded files and external responses as potentially hostile content.
- AuthenticateEstablish the caller and tenant from trusted identity infrastructure.
- Retrieve and authorizeApply source permissions; preserve them across search, reranking and context assembly.
- GenerateSupply bounded context with source identifiers. Separate evidence from instructions.
- Validate and deliverCheck the response and citation access; gate any tool action independently.
Across the path: enforce request budgets, protect caches and retain access-controlled diagnostic evidence.
Authorization belongs in application and data services. A prompt telling the model to hide confidential material is insufficient. In Microsoft’s security-filter pattern, document authorization depends on applying the filter to every query; hiding a metadata field from results is not a security mechanism.
For your implementation, derive filters from trusted identity, handle missing or expired permission data according to a documented deny policy, and check entitlement before any content reaches the model or user. Enforce equivalent boundaries for alternate retrievers, citation previews, session memory and caches. Invalidation must cover revoked access and deleted sources.
For tool-enabled systems, a generated action is a proposal. The execution service should independently validate the action, parameters and caller’s authority. Apply these controls wherever the application can execute actions.
03 / Control matrix
Test a failure, then retain evidence of the control.
This starting matrix links engineering work to an assurance decision. Assign a named owner and acceptance criterion to each applicable row. The tests are examples to adapt, rather than an exhaustive security assessment.
On narrow screens, scroll the table horizontally. Keyboard users can focus the table region and use the arrow keys.
| Failure and consequence | Control to implement | Test to run | Evidence to retain |
|---|---|---|---|
| Unreliable or poisoned sources Outdated policy or manipulated content drives an answer. | Approved ingestion; version and provenance records; quarantine and deletion propagation. | Insert a superseded policy, duplicate chunks and a tampered source. Check eligibility and removal from the final context. | Source owner, ingestion rules, corpus snapshot and failed-source disposition. |
| Provider and logging exposure Sensitive inputs reach an unapproved service or diagnostic store. | Approved data paths; minimization and redaction; verified provider processing, retention and training-use terms/settings. | Trace normal, retry and error paths with synthetic sensitive markers. Check destinations, telemetry redaction and log access. | Permitted data categories, provider terms/settings snapshot, outbound-path checks and log retention/access rules. |
| Permission leakage Another role or tenant’s material reaches context, output or citations. | Permission-aware retrieval and context assembly; entitlement-aware cache isolation and invalidation. | Use paired identities, cross-tenant queries and warm-cache requests; revoke access and test again. | Identity-policy reference, expected allow/deny cases, stage-level source IDs and leak findings. |
| Prompt injection User or source content attempts to redirect behavior or obtain secrets. | Separate instructions and source data; constrain capabilities and egress; validate outputs and tool requests. | Challenge user inputs and retrieved material with malicious instructions. Test answer manipulation and unauthorized actions under repeated runs. | Payload version, attack objective, achieved behavior, control settings and retest result. |
| Missing or misleading context Relevant evidence is absent, truncated or displaced. | Measure retrieval and final-context coverage; govern chunking, ranking, token budgets and freshness. | Test distractors, conflicting sources, long contexts and evidence placed at different positions. | Query labels, retrieved ranking, final context, relevant source versions and configuration. |
| Unsupported answer or citation A material claim is wrong, ungrounded or linked to irrelevant evidence. | Claim-level evaluation; verified source references; defined abstention and escalation behavior. | Include answerable and unanswerable questions. Inspect material claims against authoritative references and cited passages. | Rubric, expert judgments, disagreements, unsupported claims and escalation outcomes. |
| Unsafe output or excessive agency Generated content executes or an action exceeds authority. | Context-appropriate output encoding, schema validation, least-privilege tools and approval for consequential actions. | Challenge HTML rendering, structured fields and tool parameters. Verify denials and approval gates at execution. | Validation rules, tool scopes, authorization decisions and approval/action records. |
| Resource exhaustion or change failure Requests consume excessive cost or a new configuration silently regresses. | Token, time, concurrency and spend limits; bounded retries; monitored release and tested rollback. | Exercise oversized input, repeated calls, retrieval outages and rollback to a known configuration. | Load profile, latency/cost/error results, alert delivery and recovery exercise. |
Security references: OWASP LLM08:2025 for retrieval and embedding risks; LLM01:2025 for prompt injection; LLM05:2025 for output handling; and LLM06:2025 for excessive agency. No single prompt or filter makes these failures impossible.
04 / Evaluation
Measure retrieval, answers and security separately.
Build a versioned test set from the permitted tasks and plausible failures. Include difficult questions, source conflicts, stale documents, denied access, unsupported requests and degraded dependencies. Separate tuning data from a held-out acceptance set; document coverage gaps and possible test contamination.
Retrieval quality
Where queries have labeled relevant sources, define Recall@k = relevant items in the top k ÷ all relevant items in the eligible corpus. Fix whether “item” means document or chunk, deduplicate appropriately, and state how permissions define eligibility. Queries with no relevant eligible source need a separate no-answer test; their recall denominator is zero.
Measure relevant evidence surviving into the final model context as well as initial search results. A strong retriever can still lose the answer during reranking, truncation or context assembly.
Answer and citation quality
Score correctness against an authoritative reference, grounding against supplied evidence, citation support at the material-claim level, and citation coverage across claims needing evidence. State the rubric and denominator for each. Verify cited source identity, version and reader access.
Track correct abstention on unanswerable cases and unnecessary refusal on answerable ones. A system that refuses everything can look safe while being unusable.
Security and operating behavior
Report unauthorized disclosures and actions as explicit findings. Define attack success by an observed prohibited outcome; report attempts, successful attempts, repetitions, identities and conditions. A detection alert alone does not demonstrate that an attack was blocked.
Report latency percentiles, cost per request, timeout/error rates and budget violations under a stated load. Keep failures visible; do not calculate a flattering latency or quality score solely from successful requests.
The RAGAs paper (Es et al., 2024) illustrates why retrieval relevance, faithful use of context and generated-answer quality need separate evaluation. Automated judges can accelerate a review; our recommended practice is to calibrate them against blinded domain-expert labels, record disagreements and avoid letting a model grade its own release without independent checks.
Lost in the Middle (Liu et al., 2024) found position-sensitive performance in the models and tasks studied. Use that finding to motivate local tests of evidence order, context length and distractors. It does not establish how every current model behaves.
Acceptance criteria should precede the run. Agree on task-specific thresholds, mandatory security gates, sample sizes and review authority. Report counts and uncertainty alongside averages. Zero observed leaks in a finite test set is evidence about that run, not proof of zero production risk.
05 / Evidence
Keep enough context to investigate a result.
A dashboard score cannot explain which source, permission or configuration produced an answer. Link each evaluated request to the deployed configuration and a protected diagnostic record. The following is an illustrative record outline, not a required standard or a software API.
- Scope and identity
- Test/run ID; use case; pseudonymous caller/tenant reference; authorization-policy version; expected permission result.
- System configuration
- Provider and available model identifier; prompt hash/version; retriever, embedding and reranker versions; chunking settings; context/token limits; tool permissions.
- Source and retrieval
- Corpus/index snapshot; source IDs and versions; permission-decision references; ranked candidates and final context references.
- Outcome and judgment
- Protected response reference; cited passages; tool proposal and execution result; timings, token/cost data; rubric version; evaluator judgments; finding IDs.
- Decision and follow-up
- Owner; permitted release scope; unresolved findings; residual-risk decision; expiry or review trigger; required retests and rollback reference.
Use references or hashes where they support integrity and linkage; a hash does not show that the original content was accurate or lawful to use. Protect any retained prompts, responses and source excerpts with appropriate access, minimization and retention controls. Diagnostic logs can themselves expose sensitive data.
Document reconstruction limits: a hosted model alias may change, an exact provider build may be unavailable, and stochastic output may vary. Preserve what is available and distinguish configuration reconstruction from guaranteed identical replay.
06 / Release and operation
Turn the review into a bounded decision.
Agree on scope and gates.
Name the accountable risk owner, technical owner and release authority. Specify permitted users/tasks, mandatory controls, acceptance criteria and escalation routes before evaluation.
Evaluate the deployed path.
Test the integrated application with representative identities and data. Retain configurations, failures and expert judgments. Resolve release-blocking findings; record any permitted exception with compensating controls, an owner and an expiry.
Limit exposure and verify recovery.
Choose a staged rollout suited to the impact. Set stop conditions, request budgets and monitoring ownership. Exercise disablement or rollback, including how corpus and permission state remain valid.
Retest when the evidence no longer applies.
Assess changes to models, prompts, parsing, embedding, ranking, context limits, corpora, permissions, tools and use cases. Run focused regression tests for the affected boundary and broader checks when the change can alter integrated behavior.
NIST AI 600-1’s suggested actions address empirical capability evaluation (MS-2.3-002), source and citation verification (MS-2.5-003), and post-deployment monitoring (MG-4.1). These provide a useful reference for the evidence chain; our sequence above is an implementation pattern, not a claim of formal NIST alignment.
Connect wider obligations and third-party dependencies through Model Risk Support and Vendor AI Risk Evidence. When people review answers or approve actions, define their information, authority and escalation duties through Human Oversight.
07 / Worked scenario
A good answer can still fail the access boundary.
Fictional example / Policy-search assistant
An employee asks about an internal exception. The assistant produces an accurate, properly cited answer, but its response cache contains a restricted document retrieved earlier for another group.
- Engineering finding
- The cache key includes the question and model configuration, but omits the caller’s entitlement boundary. Retrieval filtering worked on the first request; a later cache hit bypassed it.
- Evidence to obtain
- Paired-identity traces for cache miss and cache hit; policy/source versions; source IDs in the delivered response; tests after permission revocation.
- Release decision
- Block the affected release path until the leak is corrected and independently retested. Assess existing exposure, restrict or disable the cache, and evaluate the incident under the organization’s procedures.
- Retest focus
- Prove that cache hits respect current entitlement and cannot return another tenant’s material. Cover response content, source titles and citation previews, plus invalidation after revocation.
An answer-quality average cannot close this finding. GRC needs the disposition and residual-risk decision; MLOps needs the failing path, corrective change and test evidence.
08 / Sources and limits
Use the sources for what they establish.
Primary sources underpin the concepts and risk categories on this page. Implementation advice, test examples and the worked scenario are InfoSecured’s synthesis. They need adaptation to the system, organization and jurisdiction.
- NIST AI 600-1 — Generative Artificial Intelligence Profile (2024)Voluntary AI RMF companion covering generative-AI risks and suggested risk-management actions. Relevant sections: 2; MS-2.3, MS-2.5; MG-4.1.
- OWASP LLM01:2025 — Prompt InjectionDirect and indirect instruction attacks; mitigations reduce risk without establishing foolproof prevention.
- OWASP LLM05:2025 — Improper Output HandlingRisks when generated output is passed to interpreters or downstream systems without appropriate validation and encoding.
- OWASP LLM06:2025 — Excessive AgencyRisks from excessive functionality, permissions or autonomy in applications that can act.
- OWASP LLM08:2025 — Vector and Embedding WeaknessesRetrieval-related threats including unauthorized access, cross-context leakage and poisoning. Relevant to vector-based deployments.
- Microsoft Learn — Security Filter PatternA concrete document-authorization implementation for Azure AI Search; provider-specific documentation, rather than a universal RAG security guarantee.
- Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (NeurIPS 2020)Foundational architecture and task-specific experimental results.
- Es et al. — RAGAs: Automated Evaluation of Retrieval Augmented Generation (EACL 2024)A research framework for evaluating distinct dimensions of RAG quality; not a sufficiency claim for automated assurance.
- Liu et al. — Lost in the Middle: How Language Models Use Long Contexts (TACL 2024)Experimental evidence about context-position sensitivity under the study’s models and tasks.
This is technical and governance guidance, not a jurisdiction-specific legal interpretation. Evidence supports a decision within the tested scope; it does not eliminate residual risk or replace specialist review where the use case requires it.
Put the review to work
Link tests, findings and decisions.
Use the AI Assurance Evidence Review Kit to organize the review record, then adapt the controls and evaluation cases on this page to the application you operate.