Skip to contentThe Observability LayerSearch

Flagship compendium · Section 11 of 24

Part III.F: Third-Party & Supply-Chain Controls

Nearly every "in-house" agent in financial services is a thin wrapper around someone else's frontier model, API, and, increasingly, tool servers. The institution owns the outcome; the vendor owns the weights. Supervisors have noticed: IOSCO's 2026 toolkit puts agentic AI and vendor concentration inside a securities supervisor's remit, with the pointed suggestion that where many firms lean on the same provider, "supervisors could consider assessing systemic vulnerabilities" (iosco.org FR/02/2026). The controls below treat the model provider as what it is: a material outsourcing arrangement with a mind of its own upstream.

TPR-01: Model-Provider Due Diligence and Third-Party AI Risk Classification

Objective. Subject every external model, API, and agent-component provider to due diligence proportionate to criticality before use.

Control. No third-party model or agent component enters production without a completed due-diligence assessment covering the risk areas IOSCO enumerates for third-party AI providers: "Inadequate due diligence. • Poor contract terms. • Inadequate monitoring of the third-party Provider. • Vendor concentration and dependency risk. • Service provider failures. • Lack of technical skills involved in procurement process." (iosco.org/library/pubdocs/pdf/ioscopd823.pdf). Each engagement receives a criticality tier driving the depth of TPR-02..08.

Implementation (financial enterprise). Fold model-provider assessment into the existing TPRM program with AI-specific fields: eval evidence (TPR-04), documentation artifacts (TPR-05), change-notification terms (TPR-03), access-security posture (TPR-08). Staff procurement with AI-literate technical reviewers: the IOSCO list names procurement skill gaps as a risk area in its own right. Scope diligence lifecycle-wide, as the toolkit does: it "aims to cover the full lifecycle of a particular AI system and to apply to all AI system types … to those using GenAI and emerging Agentic AI techniques."

Maturity. Baseline: questionnaire-based diligence with criticality tiering. Enhanced: technical verification of vendor claims; standing vendor scorecards. Frontier: continuous assurance feeds (vendor eval releases, incident disclosures) into a live vendor risk profile.

Ownership. 1st line: procurement and the sponsoring business. 2nd line: TPRM and model risk co-own the AI addendum. 3rd line: audits diligence completeness for critical vendors.

Evidence. Completed AI diligence records; criticality-tier rationale; procurement reviewer qualifications; refresh schedule adherence.

Mappings. NIST AI RMF Govern, Map; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: vendor model risk; DORA ICT third-party risk management aspect.

Sources. iosco.org/library/pubdocs/pdf/ioscopd823.pdf (T1): https://www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf

TPR-02: Vendor Concentration and Systemic Dependency Management

Objective. Measure, report, and bound the firm's dependency concentration on frontier-model providers.

Control. The firm maintains a concentration view of all AI workloads by underlying model provider, reported to the board risk committee, with defined tolerance thresholds. Supervisory basis: "in cases where multiple institutions rely on the same third-party AI providers, supervisors could consider assessing systemic vulnerabilities, recognizing that AI failures in systemically important firms or shared infrastructure could have cascading effects", and "larger, systemically important institutions may face heightened supervisory expectations even for medium or low-risk AI applications, given their potential market-wide impact" (iosco.org/library/pubdocs/pdf/ioscopd823.pdf).

Implementation (financial enterprise). Extend the DORA-style critical-ICT-provider register to model providers; tag each MAS-01 inventory entry with its full model dependency chain, including models behind vendor products. Set concentration appetite (maximum share of critical processes on one provider); breaches are risk-appetite exceptions. Prepare the examiner answer now: IOSCO flags "concentrations in and dependencies on third-party providers, adversarial attacks" among the risks frameworks "will likely need to evolve to address."

Maturity. Baseline: provider concentration report by business line. Enhanced: appetite thresholds with exception governance; dependency chains through vendor products. Frontier: sector-level view exchanged with supervisors; tested multi-provider substitution for critical flows (TPR-07).

Ownership. 1st line: business/technology owners report dependencies. 2nd line: risk aggregates, sets appetite, escalates breaches. 3rd line: audits completeness of the dependency map.

Evidence. Concentration dashboard; appetite statement and breach log; board pack extracts; supervisor correspondence.

Mappings. NIST AI RMF Govern, Map; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: aggregate model risk; DORA concentration-risk aspect (directly analogous).

Sources. iosco.org/library/pubdocs/pdf/ioscopd823.pdf (T1): https://www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf

TPR-03: Upstream Model and Weight Change Management

Objective. Ensure no upstream model change (version, weights, defaults, deprecation) reaches production agents without detection, re-validation, and approval.

Control. Every external model dependency is version-pinned where the provider allows; provider change events trigger a mandatory impact assessment and proportionate re-validation before the new version serves regulated processes. Silent upstream drift discovered in production is a recordable incident. IOSCO's lifecycle-wide supervisory scope (model development and testing, monitoring and controls, outsourcing: iosco.org/library/pubdocs/pdf/ioscopd823.pdf) implies the firm, not the vendor, answers for post-change behavior.

Implementation (financial enterprise). Contract for advance change notification and pinning (TPR-08); run a standing canary eval suite against pinned and candidate versions; gate promotion on eval parity plus control-stack re-baselining: evasion dynamics shift with model capability (arxiv.org/abs/2607.02514), so monitor efficacy is re-measured at every upgrade, not assumed. Provider deprecations are forced changes entering the same gate on an accelerated path. [practice guidance: not directly source-backed]: detect unannounced changes via daily behavioral fingerprinting (fixed probe prompts, distribution checks).

Maturity. Baseline: version pinning plus release-note monitoring. Enhanced: canary eval suite gating every version promotion. Frontier: automated fingerprint drift detection with auto-rollback to the pinned version.

Ownership. 1st line: platform engineering operates gates. 2nd line: model risk defines re-validation depth per materiality. 3rd line: audits change-event handling end-to-end.

Evidence. Pinning configuration; change-event log with impact assessments; canary results per promotion; incident records for silent drift.

Mappings. NIST AI RMF Manage, Measure; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: model change management and revalidation; DORA ICT change management aspect.

Sources. iosco.org/library/pubdocs/pdf/ioscopd823.pdf (T1), https://www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf; arxiv.org/abs/2607.02514 (T1), https://arxiv.org/abs/2607.02514

TPR-04: Vendor Evaluation Evidence: Demand, Verify, and Do Not Rely Solely on First-Party Claims

Objective. Base reliance on vendor models on evaluation evidence that is transparent, valid, and at least partly independent.

Control. For each critical model dependency, the firm obtains evaluation evidence sufficient to know what claims it supports, what system was tested, how the behavior was elicited, and how validity was checked: the transparency bar proposed in the OpenAI/METR/Apollo shared playbook (openai.com/index/trustworthy-third-party-evaluations-foundations, reported content, provisional claim), including whether results were adjusted for confounds such as reward hacking and sandbagging. First-party evidence alone is insufficient for critical dependencies: the position that in-house evaluation is not enough, and that robust third-party evaluation and flaw disclosure are required, is established in peer-reviewed literature (proceedings.mlr.press/v267/longpre25a.html; arxiv.org/abs/2503.16861).

Implementation (financial enterprise). Score vendor eval packs against a checklist: system-under-test identity (exact version, TPR-03), elicitation method, validity checks, evaluator independence and access tier (TPR-08), and coverage gaps, mapping work finds systematic gaps between first- and third-party evaluation coverage (arxiv.org/abs/2511.05613). Where the vendor proposes the standard it is measured against, treat the framing critically. Supplement with internal pre-deployment assurance for agents built on vendor models: ontology-grounded scenario testing beat persona-based testing on regulatory coverage (48.3% vs 33.1% across 1,800 scenarios spanning Fintech, Banking, Insurance and Healthcare), with a machine-verifiable trust certificate as the sign-off artifact (arxiv.org/abs/2606.04037).

Maturity. Baseline: vendor eval documentation collected and checklist-scored. Enhanced: independent third-party eval evidence required for critical models. Frontier: machine-verifiable assurance artifacts (trust-certificate pattern) attached to each deployment approval.

Ownership. 1st line: sponsoring teams collect evidence. 2nd line: model risk assesses sufficiency and owns the reliance decision. 3rd line: audits reliance decisions against policy.

Evidence. Scored eval-evidence checklists; independence assessment; internal assurance-run results; deployment approvals citing the evidence.

Mappings. NIST AI RMF Measure, Govern; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: validation of vendor models; DORA, testing aspect.

Sources. openai.com/index/trustworthy-third-party-evaluations-foundations (T2, provisional claim), https://openai.com/index/trustworthy-third-party-evaluations-foundations/; proceedings.mlr.press/v267/longpre25a.html (T2), https://proceedings.mlr.press/v267/longpre25a.html; arxiv.org/abs/2503.16861 (T2), https://arxiv.org/abs/2503.16861; arxiv.org/abs/2511.05613 (T2), https://arxiv.org/abs/2511.05613; arxiv.org/abs/2606.04037 (T1): https://arxiv.org/abs/2606.04037

TPR-05: Documentation Artifacts for Third-Party Models and Datasets

Objective. Require complete, current documentation (model cards, dataset documentation, system cards) for every third-party model and dataset before reliance.

Control. Third-party models and datasets are ingested only with documentation meeting a firm-defined minimum (intended use, training-data characterization, limitations, eval results, version). Missing or stale documentation blocks or conditions the reliance decision. the cited research documents that real-world documentation practices for third-party ML models and datasets are inconsistent (The State of Documentation Practices of Third-party Machine Learning Models and Datasets, arxiv.org/abs/2312.15058), so the control assumes gaps and verifies rather than trusts.

Implementation (financial enterprise). Store artifacts in the model inventory alongside SR 11-7 development evidence; diff them at every vendor version change (TPR-03); where a component (e.g., an embedded dataset) arrives undocumented, record documentation debt with a 2nd-line remediation or compensating-control decision. Feed stated limitations directly into use-restriction configuration.

Maturity. Baseline: documentation collected and completeness-checked at onboarding. Enhanced: documentation diffs on every version change; debt register with owners. Frontier: automated conformance checks of documentation claims against observed behavior.

Ownership. 1st line: onboarding teams collect and register artifacts. 2nd line: model risk sets the minimum standard and adjudicates debt. 3rd line: audits debt aging.

Evidence. Artifact repository extracts; completeness scores; documentation-debt register; use-restriction configs traceable to stated limitations.

Mappings. NIST AI RMF Map, Govern; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: developmental evidence for vendor models; DORA, third-party documentation aspect.

Sources. arxiv.org/abs/2312.15058 (T2): https://arxiv.org/abs/2312.15058

TPR-06: Third-Party Tool and MCP-Server Supply-Chain Security

Objective. Prevent third-party agent tools, MCP servers and equivalent tool endpoints, from becoming an unvetted execution and exfiltration channel.

Control. Third-party MCP servers and agent tools are onboarded through the same security gate as any software supply-chain dependency: provenance verification, permission scoping, and continuous monitoring, per the OWASP GenAI practical guide for securely using third-party MCP servers (genai.owasp.org, CheatSheet 1.0). No agent may connect to an MCP server absent from the approved registry.

Implementation (financial enterprise). Maintain an approved-tool registry linked to the MAS-01 inventory (which agent may call which server, with what scopes); pin tool-server versions and treat manifest changes as supply-chain events: a changed tool description is a changed attack surface; run third-party tools least-privileged with egress restrictions; log every invocation attributably (MAS-02). [practice guidance: not directly source-backed]: sandbox third-party tool execution away from systems of record; re-review on any declared-capability change.

Maturity. Baseline: approved registry plus scoped credentials per tool. Enhanced: version pinning, manifest-change detection, egress controls. Frontier: runtime behavioral monitoring of tool servers with automatic quarantine on anomaly.

Ownership. 1st line: platform engineering operates the registry and gates. 2nd line: information security sets the onboarding standard. 3rd line: audits registry-to-runtime conformance.

Evidence. Registry with scopes and versions; onboarding security assessments; manifest-change alerts and dispositions; invocation logs.

Mappings. NIST AI RMF Map, Manage; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: operating environment controls; DORA ICT supply-chain aspect.

Sources. genai.owasp.org/resource/cheatsheet-a-practical-guide-for-securely-using-third-party-mcp-servers-1-0 (T1): https://genai.owasp.org/resource/cheatsheet-a-practical-guide-for-securely-using-third-party-mcp-servers-1-0/

TPR-07: API Dependency Resilience, Degradation, and Exit

Objective. Ensure the failure, throttling, or withdrawal of an external model API degrades agent-dependent processes gracefully rather than catastrophically.

Control. Every critical agent workflow has a documented and tested response to provider unavailability: fail-safe halt, human fallback procedure, or substitution to an alternate approved model. "Service provider failures" is an enumerated third-party risk area in IOSCO's supervisory toolkit, and AI failures in "shared infrastructure could have cascading effects" (iosco.org/library/pubdocs/pdf/ioscopd823.pdf). Untested fallbacks do not count.

Implementation (financial enterprise). Classify agent workflows by outage tolerance; for low-tolerance flows, pre-approve an alternate provider or staffed non-AI procedure and rehearse the switch: substitution changes model behavior, so the alternate path needs its own validation and control re-baselining (TPR-03 gate applies). Include provider outages in operational-resilience exercises; wire rate-limit exhaustion into the same playbooks as hard outages. [practice guidance: not directly source-backed]: agents should emit a clean "cannot proceed" state rather than degrade silently on partial API failure.

Maturity. Baseline: documented fallback per critical workflow; annual test. Enhanced: pre-validated alternate provider for the most critical flows; outage scenarios in resilience exercises. Frontier: automated failover within tested control parameters, with post-failover monitoring intensification.

Ownership. 1st line: process owners maintain and test fallbacks. 2nd line: operational risk sets tolerance classes and reviews test results. 3rd line: audits test authenticity (did the fallback actually run?).

Evidence. Fallback runbooks; test/exercise reports; alternate-provider validation records; outage incident post-mortems.

Mappings. NIST AI RMF Manage; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: contingency around model failure; DORA operational-resilience testing aspect (directly analogous).

Sources. iosco.org/library/pubdocs/pdf/ioscopd823.pdf (T1): https://www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf

TPR-08: Contractual Controls, Secure Access Tiers, and the Third-Party Assurance Ecosystem

Objective. Encode the firm's AI supply-chain requirements in enforceable contract terms, and govern any access granted to third parties (evaluators, auditors, vendors) by explicit access-risk tiers.

Control. Contracts with model providers and AI vendors must cover, at minimum: change notification and version pinning (TPR-03), eval-evidence delivery (TPR-04), documentation obligations (TPR-05), availability/exit terms (TPR-07), audit and information rights, and data-use restrictions, "Poor contract terms" is an enumerated IOSCO third-party risk area (iosco.org/library/pubdocs/pdf/ioscopd823.pdf). Conversely, when the firm grants third parties access to its own models or agent internals, access is tiered by risk: RUSI treats evaluator access as a controlled artifact, proposing an Access–Risk Matrix from API-only access up to write access to model internals, the tier it identifies as highest-risk, with security risks "from intellectual property leakage to model compromise to exploitation by state-sponsored actors" that "remain poorly mapped and inadequately standardized" (rusi.org, May 2026; quotations reported by the press; primary wording not established here).

Implementation (financial enterprise). Maintain an AI contract-clause library with legal and TPRM; make the clauses gating conditions in procurement, not aspirations after signature. For inbound access (external auditors, validators, regulators' assessors), record the tier granted, threat scenarios considered, and compensating controls, and ask every external evaluator the RUSI question: what access tier did you have? Track the maturing assurance market: third-party audit ecosystem design (doi.org/10.1145/3514094.3534181), the UK trusted third-party AI assurance roadmap (gov.uk), so contracts can require recognized assurance forms as they standardize; structured-access research (governance.ai) informs safe researcher/evaluator access design.

Maturity. Baseline: standard AI clause set in all new contracts; access grants documented. Enhanced: clause compliance monitored (notifications actually arrive); access-risk matrix operational for all inbound access. Frontier: contracts require independent assurance artifacts as the external ecosystem matures; periodic re-negotiation driven by supervisory expectations.

Ownership. 1st line: procurement and vendor managers. 2nd line: legal, TPRM, and model risk co-own the clause standard and access-tier policy. 3rd line: audits clause presence and access-grant hygiene.

Evidence. Clause library and coverage reports; executed contracts vs. standard; inbound access-grant register with tiers; assurance artifacts received.

Mappings. NIST AI RMF Govern; ISO/IEC 42001 AIMS; EU AI Act, not sourced in the sources cited here; SR 11-7 aspect: vendor arrangements and contingency; DORA contractual-provisions aspect (directly analogous).

Sources. iosco.org/library/pubdocs/pdf/ioscopd823.pdf (T1), https://www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf; rusi.org/explore-our-research/publications/insights-papers/developing-framework-secure-third-party-access-frontier-ai (T2), https://www.rusi.org/explore-our-research/publications/insights-papers/developing-framework-secure-third-party-access-frontier-ai; doi.org/10.1145/3514094.3534181 (T1), https://doi.org/10.1145/3514094.3534181; gov.uk/government/publications/trusted-third-party-ai-assurance-roadmap/trusted-third-party-ai-assurance-roadmap (T1), https://www.gov.uk/government/publications/trusted-third-party-ai-assurance-roadmap/trusted-third-party-ai-assurance-roadmap; governance.ai/research-paper/structured-access-for-third-party-research-on-frontier-ai-models (T3): https://www.governance.ai/research-paper/structured-access-for-third-party-research-on-frontier-ai-models