AI procurement & third-party risk
Vendor AI risk
Vendor AI risk is the risk introduced when an organization relies on an external AI product, model, data source, or service. The review should answer one practical question: can this product be used for this purpose, with these data, permissions, and dependencies, at an acceptable level of risk?
A supplier questionnaire is input to that decision—not the decision itself. The useful work is connecting vendor claims to evidence, testing the product in the intended workflow, assigning responsibilities, and planning for change or exit.
Start before the questionnaire
Define the purchase in operational terms.
AI vendor due diligence should be proportional to the use. A drafting assistant working from public material and an agent able to change customer records should not receive the same review.
What will the product draft, recommend, rank, predict, retrieve, or execute?
Who relies on the output, and who is affected when it is wrong?
Which records, prompts, files, logs, or derived stores may contain sensitive or regulated data?
Which systems, connectors, tools, and permissions can the product reach?
Which outcomes are unacceptable, hard to reverse, or time-critical?
Who can approve the use, restrict it, suspend it, and authorize restart?
Set review depth from consequence and exposure. Limited, reversible assistance may support a narrower review. Sensitive data, consequential recommendations, broad permissions, or operational dependence require stronger evidence and testing.
Third-party does not mean one party
Map what sits behind the product.
The company on the contract may depend on a model provider, hosting platform, external data, APIs, or tools. Shared upstream dependencies matter for security, continuity, change control, and fallback planning. [1]
- 01Your organizationUsers, data, permissions, workflowWhat you configure and permit
- 02Application supplierInterface, retrieval, integrationsWhat wraps the AI capability
- 03Model / service providerModel, hosted API, inference serviceWhat generates, scores, or predicts
- 04Supporting dependenciesHosting, data, libraries, external toolsWhat the service relies on
Ask three questions across the chain:
- Who receives or can access our information?
- Who can materially change the product’s behavior?
- Which shared dependency could interrupt both the primary service and a supposed fallback?
Due diligence that changes the decision
Request evidence you can examine.
Ask for evidence tied to the product, version, configuration, and intended use. A vendor response records a claim. The review should determine whether the claim is sufficiently supported for the decision you need to make.
| Decision question | Request from the supplier | Examine before approval |
|---|---|---|
| What exactly are we buying? | Product and model identifiers, enabled AI features, intended uses, known limitations, architecture, and material upstream dependencies. | Do the documents describe the service and settings you will actually receive? Record differences between the demonstration, trial, and proposed deployment. |
| Does it perform the task well enough? | Evaluation methods, task definitions, datasets, sample sizes, error breakdowns, test dates, and access to a relevant trial. | Do the evaluation conditions resemble your users, inputs, languages, and consequences? Look beyond a headline accuracy number to the errors that matter. |
| Where does our information go? | Data flows covering prompts, files, outputs, logs, support access, retention, deletion, training, and product improvement. | Do contractual commitments and technical configuration agree? Clarify backups and derived stores such as indexes or embeddings. |
| What can the product access or do? | Connector permissions, tenant-isolation design, authorization controls, security testing, and mechanisms to restrict or disable actions. | Can ordinary users retrieve restricted material or trigger unintended actions? Separate supplier controls from controls that depend on customer configuration. |
| How are changes and incidents handled? | Version policy, change notices, incident contacts, support commitments, diagnostics, and treatment of upstream-provider changes. | Can you detect a material change, assess its effect, and restrict use while it is reviewed? Confirm who receives notices and who can act on them. |
| Can we recover or leave? | Export formats, deletion procedures, continuity arrangements, termination terms, and migration support. | Can the business continue without the product? Test what can be exported, what must be rebuilt, and whether a fallback shares the same critical dependencies. |
Documentation
Use each artifact for the claim it can support.
Model cards can provide intended-use and evaluation context; dataset documentation can describe data creation, composition, use, and maintenance. [2] [3] Neither automatically establishes that the surrounding application’s permissions, retrieval, integrations, or customer configuration work correctly.
Independent reports
Read the scope before relying on the badge.
For an audit report, certificate, or test summary, check the assessed entity, service, period, exclusions, exceptions, and customer responsibilities. Record which specific vendor claims the report supports and which remain open.
Need a record structure for the evidence itself? See Vendor AI Risk Evidence. For evaluating claim strength, see AI assurance evidence.
Worked review
Turn supplier claims into checks.
Example: an internal assistant answers employee questions from company policies. It has read-only access and returns source references. The decision is whether to permit a limited deployment and under what conditions.
“Answers are grounded in your documents.”
Use clear questions, conflicting policy versions, missing answers, and misleading retrieved passages. Verify that citations actually support the answer and that uncertainty reaches a useful fallback.
Test questions, expected outcomes, cited passages, errors, configuration, and policy-owner adjudication.
“Users only see content they may access.”
Use test accounts with different permissions. Attempt restricted queries, change access, and repeat. Check answers, citations, previews, and conversation history for unintended disclosure.
Access settings, observed results, permission-change behavior, and ownership of any mismatch.
“Your data is not used for training.”
Trace the commitment through the purchased plan, enabled features, upstream providers, support access, retained logs, and relevant contract terms. Functional testing alone cannot establish internal data-use practices.
Applicable commitments, settings, data flows, corroborating assurance, and unresolved retention or secondary-use questions.
Pilot discipline
Test the claim before scaling the dependency.
Define success, critical failures, and comparison with the current process before the trial. Measure supported answers, appropriate abstentions, access-control failures, correction effort, and useful task completion. Keep counts and denominators visible. A limited sandbox can answer functionality questions while data-use or contractual questions are still being resolved. Research on AI functionality failures supports scrutinizing whether products work as claimed rather than assuming advertised benefits. [4]
Controls that survive procurement
Assign responsibility before deployment.
A supplier may operate the model while your organization controls users, data, integrations, and decisions made from its output. Write down the boundary so necessary controls do not fall between teams.
| Control area | Supplier | Customer | Shared / coordinated |
|---|---|---|---|
| Service & evidence | Specified behavior, relevant documentation, supported configurations | Approved use, acceptance criteria, local validation | Resolving evidence gaps and material limitations |
| Data & access | Service-side handling, isolation, support access | Permitted data, users, connectors, permissions | Investigating leakage, misuse, or access-control failures |
| Change & incidents | Notices, diagnostics, supplier response | Impact assessment, restriction, local incident response | Coordinated investigation and corrective action |
| Continuity & exit | Export, deletion, termination support | Fallback process, migration readiness, access removal | Transition, retained records, unresolved exceptions |
NIST’s Generative AI Profile connects third-party assessment, supplier agreements, incident coordination, monitoring, and fallback planning. [5] Treat those topics as issues to resolve with the relevant procurement, privacy, security, and legal specialists—not as generic model clauses.
The assessment has to end somewhere
Make a bounded decision.
Separate evidence received from evidence sufficient. An unanswered question about confidential-data access can matter more than a long list of satisfactory questionnaire responses.
Evidence supports the defined use, essential controls are in place, and an authorized owner accepts the remaining risk.
A narrower use or controlled pilot is supportable. Enforce excluded data, users, actions, or integrations and define what would justify expansion.
A material question remains open. Name the missing evidence, owner, and decision it blocks; keep pending review from becoming informal production use.
The product does not meet the need, an essential control is unavailable, or the remaining risk is unacceptable.
Decision record
Preserve the reasoning, not only the score.
Record the product and version; intended and excluded uses; data and action boundaries; evidence examined and dates; test outcomes; unresolved gaps; required controls and owners; decision-maker; conditions; and reassessment triggers.
Reassessment triggers
Review when the service or use changes.
- Model, retrieval, action-taking, or upstream-provider changes
- New population, language, dataset, decision, or integration
- Access incidents, repeated unsupported outputs, complaints, or rising correction effort
- Changed terms, reduced evidence, renewal, discontinuation, or support degradation
Exit readiness
Test whether the business can actually leave.
Exercise export, fallback operation, access removal, and migration with realistic volume and staffing. Identify common upstream dependencies before treating a second supplier as an independent fallback. Supply-chain risk management continues after onboarding. [6]
Sources & evidence notesPrimary guidance and peer-reviewed research used on this page
The decision questions, worked example, responsibility matrix, and decision outcomes are InfoSecured’s practical synthesis. The sources below support the underlying documentation, supply-chain, functionality, and third-party-risk concepts.
- UK National Cyber Security Centre and international partners (2023). Guidelines for secure AI system development: Secure development.
- Mitchell, M., et al. (2019). Model Cards for Model Reporting. FAT* ’19, 220–229.
- Gebru, T., et al. (2021). Datasheets for Datasets. Communications of the ACM, 64(12), 86–92.
- Raji, I. D., Kumar, I. E., Horowitz, A., & Selbst, A. D. (2022). The Fallacy of AI Functionality. FAccT ’22, 959–972.
- NIST (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, especially GOVERN 6.1–6.2.
- Boyens, J., et al. (2024). Cybersecurity Supply Chain Risk Management Practices for Systems and Organizations. NIST SP 800-161 Rev. 1, Update 1.