Article

Spotlight on AI in financial services

Regulatory expectations in relation to artificial intelligence (AI) use in banking and in financial services

Spotlight on AI in financial services
Published Date
Sep 22, 2026

The last few months have involved a number of regulatory developments in relation to AI, but one particularly useful paper is a report published by the International Organization of Securities Commissions (IOSCO). This considers how financial services regulators can supervise AI use, providing banks and financial services firms with a useful snapshot of what they can or should do to satisfy expectations and manage regulatory risk.

This bulletin provides a summary of some of the key points from this paper, together with a checklist for firms that we have prepared drawing on the AI supervisory tool included in the IOSCO report.

1. Regulatory supervision of AI use

Following work by its Fintech Task Force (FTF), IOSCO has published a report titled “Supervisory Toolkit for AI Use in Capital Markets”.1 For the industry, this provides an insight into what supervisors are likely to be concerned about in relation to a firm’s use of AI, where they might focus attention, what questions they might ask, and what they are looking to see as a supervisor. It is therefore a valuable resource for a regulated firm’s internal team to map and mitigate potential regulatory risk and “get ahead” of regulatory expectations.

Regulatory goals

  • Risk-based and proportionate oversight of AI systems used by firms.
  • Maintaining trust in markets as AI adoption grows. To this end, the following principles are key: proportionality, accountability, transparency, and reliability.
  • AI systems should be deployed in a manner consistent with principles of fairness, explainability, transparency, interpretability, and operational resilience.

Regulators may assess risk levels using the following approach, which firms should bear in mind:

Factor 1—consider the nature and complexity of the relevant AI system

Example:

  • “More complex or difficult to interpret AI systems may limit the ability to identify unintended behaviour and to diagnose the root causes of incidents”.
  • “AI systems that process near real-time updates in data may not perform well if the data differs from that on which the system was trained”.
  • GenAI may generate incorrect outputs (hallucinations) that are convincing but wrong.

Factor 2—consider the level of human oversight of the AI system—i.e., if something goes wrong, will a human detect and be able to resolve it before actual harm arises

How to classify levels of human involvement:

  • Human-in-control: In this scenario, the AI system cannot act alone on any particular recommendation or output. The human uses or does not use the AI system’s recommendations or output at their discretion.
  • Human-in-the-loop: In this scenario, the AI system evaluates input and acts on a particular recommendation or output only if a human approves it.
  • Human-on-the-loop: In this scenario, the AI system evaluates input and acts on a particular recommendation or output unless a human disapproves.
  • Human-out-of-the-loop: In this scenario, the AI system evaluates input and acts on a particular recommendation or output without human involvement.

Factor 3—consider the potential implications for clients and the broader market—i.e., severity of harm that could be caused by an incident

  • Assessing the impact on clients—client/ investor protection issues (e.g., suitability, privacy, data protection and discrimination); volume and type of clients/investors potentially affected; and ability to address harm (e.g., whether there are effective redress mechanisms).
  • Assessing the impact on the market— (1) larger, systemically important institutions may face heightened supervisory expectations even for medium or low-risk AI applications, given their potential market-wide impact; (2) where a number of firms rely on the same third-party AI providers, consider systemic vulnerabilities and the potential for a failure in a systemically important firm or shared infrastructure having “cascading effects”. This may be referred to as concentration risk.
  • Assessing the impact on a firm—e.g., potential for AI systems to disrupt core business activities, cause a system to fail, or result in adverse operational, legal and/or reputational consequences.

In analysing AI use within a firm, supervisors may also consider:

  • Does the firm have a governance and risk management framework calibrated to the nature of use cases at the firm and their potential outcomes? This includes the necessary knowledge and resource for adequate testing, maintenance and monitoring.
  • Does the firm have measures around data quality?
  • Does the firm have the appropriate level of transparency for key stakeholders in relation to its use of AI?

Key risks

What are regulators particularly worried about in relation to the use of AI? 

  • The potential for concentrations in and dependencies on AI service providers.
  • Increasing complexity and opacity of certain AI systems. Example: “Due to their non-deterministic nature and technological complexity, GenAI systems may be difficult to predict, evaluate, understand, explain, and test.” IOSCO gives hallucination risk as an example (for more on this, see below).
  • The rapid and unpredictable evolution of AI technologies and market dynamics add to the complexities of supervisory risk assessment—in other words, there is a risk that regulators miss something because of how quickly things move in this area. 

Plus, an overlapping point—the continuing evolution of AI use cases: 

  • IOSCO gives an example of a recent development in frontier AI models which have highlighted autonomous capabilities, particularly the ability of such systems to independently discover, chain and exploit security vulnerabilities. Developments in such capabilities may place greater emphasis on continuous risk assessment, integrated cyber and AI governance, and shorter remediation cycles.

Areas of regulatory focus

Overarching focus: 

  • How AI use may impact (1) investor protection (2) market integrity and (3) financial stability.

Specific areas of focus:

  • Governance and oversight
  • General risk management
  • Outsourcing and third-party dependency
  • Model risk management (including development, testing and ongoing monitoring)
  • Cybersecurity
  • Malicious use
  • Data privacy/protection
  • Investment advice and suitability (where relevant)
  • Market risks
  • Disclosure and transparency
  • Recordkeeping
  • Data quality and bias
  • Ethical concerns
  • System reliability and business continuity

Other risks to note: 

  • People may over-rely on AI and fail to do proper checks or conduct adequate oversight (for more on this, see below).

Mitigating the risk of hallucination

Supervisors can look for the following: 

(a) Grounding and guard-railing techniques

Technical grounding and guard-railing measures can reinforce the factual accuracy of responses. These techniques include Retrieval-Augmented Generation (RAG), Chain of Verification (CoVe), and Multi Agent Debate, among other techniques”2—however, IOSCO notes these do not fully eliminate the risk of hallucination. 

(b) Human oversight

“Ensuring that a human overseer is part of the AI system at the appropriate point(s) in the AI system’s lifecycle is widely considered as an effective measure. However, it is pointed out that human overseers may be susceptible to automation bias, may tend to over-rely on AI, and may be unable to adequately oversee certain processes in an AI system.

Robust governance and risk management at the firm level should consider how human oversight can be effectively implemented, including whether the human overseer has the requisite capabilities and knowledge of the system’s design and limitations to exercise effective oversight and whether there is sufficient accountability for human oversight.” 

(c) Disclosure

IOSCO takes pains to confirm that firms are responsible for the financial products and services they provide, regardless of whether they use AI or some other technology. The primary responsibility for addressing hallucination risk therefore remains with the firm. But alongside steps to address this risk, appropriate disclosures and client education on AI risks can improve client and user awareness of the risks involving AI systems, including the risk of hallucination. 

Specific risks with Agentic AI

The complexities of Agentic AI compound the complexities of risk assessment and can introduce risks with significant consequences in financial markets: 

  • Agentic AI systems can have access to sensitive data, be exposed to malicious external content and/or have the ability to communicate externally. If not configured properly, this could result in data compromise, exfiltration, plus operational and cybersecurity issues.
  • “Vulnerabilities caused by poorly designed prompts, inadequate safeguards or insufficient testing and monitoring of agent interactions could lead to poor performance and failures in Agentic AI systems.”
  • As Agentic AI uses more components, the potential interplay between components could create a greater risk of “unexpected emergent behaviors or the potential for cascading impacts or failures across interconnected systems”.
  • Increased “complexity and opacity, coupled with increasingly automated workflows, can create the potential for unpredictable or unwanted behaviors”. E.g., “an AI agent may take undesired steps to pursue its goal, such as engaging in collusive behaviors with other components or systems, or its goal may become misaligned from the one for which it had been deployed”.
  • Additional challenges with detecting and addressing unwanted activity in Agentic AI systems can arise due to their complexity and opacity. This puts more pressure on governance, risk management and oversight. 

Mitigating risks in relation to GenAI

A distinction for GenAI is that system behavior cannot be fully specified or anticipated in advance of deployment, as opposed to traditional software, which behaves according to rules that developers specify (and therefore understand). This provides a stronger guarantee as to how the system will behave.

On the other hand, GenAI relies on underlying models which are trained on data, and learns tasks from data rather than explicit programming. This gives rise to a level of uncertainty, and GenAI systems lose many of the guarantees that make rules-based systems predictable. 

To mitigate these risks, firms can:

  • maintain comprehensive documentation
  • conduct formal and independent reviews
  • validate systems prior to and after deployment (e.g., using stress testing and scenario analysis)
  • conduct ongoing performance monitoring to detect issues such as drift, degradation or bias of underlying models
  • benchmark outputs against industry standards to assess quality and consistency.

2. Checklist

For a checklist we have prepared to reflect some of the key points from the supervisory toolkit in the IOSCO report, see Schedule 1.

3. Next steps

Through its Artificial Intelligence Working Group (AIWG), IOSCO will next conduct a review of emerging industry practices in relation to AI, focusing on disclosure, recordkeeping, reporting and governance. This is likely to provide a useful benchmark for global banks in particular, and is therefore something to watch out for.

In the interim, if any of the A&O Shearman team can assist you with questions or projects relating to AI, please do not hesitate to email your usual contact.

Schedule 1 checklist

Based on the supervisory toolkit suggested by IOSCO, a checklist for firms to use in practice is as follows:

AI governance and oversight

What should a firm do?

  • Strong board and senior management oversight and accountability:
    • The board should ensure it promotes a corporate culture that prioritises ethical, fair, and responsible AI use across the organisation.
    • It should ensure clear accountability mechanisms are in place in relation to AI and its risks.
  • Adequate reporting (MI) to board and senior management:
    • MI should be provided on areas such as the design, implementation and use of AI systems, governance and risk management, and measures to address investor protection and market integrity.
  • Documented internal AI governance and risk management framework and broader AI policies and procedures, including:
    • IT and data governance framework
    • human oversight policies and procedures for appropriate and proportionate human intervention/ interruption in the operation of the AI system— in particular, more human involvement may be warranted where an outcome can impact clients and the markets
    • a clear AI strategy and risk appetite
    • appropriate qualitative statements and quantitative measures or limits.
  • Ensure decision-making and oversight mechanisms are proportionate to the materiality and complexity of AI use cases.
  • Ensure the risk management framework:
    • encompasses the entire AI lifecycle, from design/development to deployment/ongoing operation to retirement
    • includes controls for use case design, as well as model, data and third-party risk
    • includes measures to promote transparency, reliability, robustness, resilience, fairness, security, safety and privacy
    • includes escalation protocols for AI-related incidents
    • includes a requirement to maintain comprehensive audit trails to enable a supervisory review as well as adequate reporting to senior management and the board
    • is regularly reviewed over time (i) bearing in mind the developing nature of the technology, the risks to which it may give rise, etc, (ii) factoring in the experience of the firm and the industry as this evolves over time, and (iii) factoring in risks such as concentrations in and dependencies on third-party providers, adversarial attacks and unintended market impacts.
  • Implement an AI inventory and classification system to use in relation to AI use cases—i.e., to calibrate and identify risk levels.
    • The inventory would capture key attributes—e.g., the AI system’s purpose and description, scope of use, data and model usage, and upstream and downstream dependencies.
  • Documented approval processes.
    • An appropriately senior individual or groups of individuals, with the relevant skill set and knowledge, to sign off on initial deployment and substantial updates of relevant AI technology.
  • A training and ongoing education programme for board, senior management and staff exercising control functions (“AI literacy”).
  • Strong understanding and knowledge of AI system design.
  • Clear and documented roles and responsibilities for developers, deployers and users.
  • For GenAI and Agentic AI, consider establishing pilot and experimentation frameworks with clear policies and procedures for pilots that are bound by time and user limits and often occur in a contained environment or with non-sensitive data.
  • Ensure Internal Audit conducts reviews concerning the use of AI as appropriate.

What records should a firm establish/maintain to show a regulator?

  • AI governance policies and procedures
  • AI risk management framework
  • AI inventory/registry including technical documentation/description of AI systems
  • Governance committee documents
  • Organisational charts and role/responsibility assignment matrices that delineate roles/responsibilities
  • Training programs, materials, and records relating to AI systems and usage
  • Human oversight policies and procedures for appropriate and proportionate human intervention/interruption in the operation of the AI system
  • Documentation of staff qualifications including certifications, competency assessments and continuing education plans

NB: This last point is key, as the report acknowledges (1) a lack of skilled professionals in AI and related fields (“talent shortages”), (2) concerns about the opacity of systems, and (3) the risk that firms rely too much on external vendors without a sufficient understanding of how a model or system really works (“reliance on third-party models”)—and if there is no understanding, how can there be sufficient oversight. The IOSCO report also helpfully notes as follows: “Some firms have set up AI Centers of Excellence to drive innovation, promote good practices, and build AI capabilities.”

Model risk management

What should a firm do?

  • Robust model testing including back testing, stress testing, testing for bias and model drift, and underlying data testing
  • Robust ongoing performance monitoring
  • Independent validation of model performance
  • Alert mechanisms for anomaly detection
  • Methodology for AI system suspension where anomalies are detected

What records should a firm establish/maintain to show a regulator?

  • Model validation and testing policies and procedures and reports, including pre-deployment, post-deployment and ongoing validation, monitoring and testing, including documentation and logs
  • Model performance monitoring policies and procedures and reporting, including performance and bias
  • Model change management policies and procedures and logs
  • Independent AI model validation and testing reports
  • Model performance anomaly detection alerts and processes

Investment advice and suitability (where relevant)

What should a firm do?

Robust and sufficient measures in place to: 

  • ensure the suitability of AI-generated recommendations (if any)
  • ensure adequate consideration of individual client circumstances
  • avoid bias in AI systems, including logic and prompts
  • avoid overstating the firm’s AI capabilities (“AI-Washing”). 

PLUS robust human oversight.

What records should a firm establish/maintain to show a regulator?

  • Client profile and input data
  • Suitability policies and procedures addressing client profile, including risk tolerance and investment objectives
  • Policies and procedures addressing the availability of product or service offerings for investment recommendations by an AI system
  • AI recommendation logs, including outputs and supporting data
  • Human review and override documentation
  • Conflict identification and management policies and procedures
  • Client complaints data
  • Evidence of client disclosures

Market risks

What should a firm do?

Document the following risks, and have adequate policies and procedures in place to mitigate them: 

  • AI use amplifying market volatility 
  • Flash crashes from AI-driven trading 
  • Herding behaviour/correlation/collusion 
  • Liquidity issues 
  • Other sources of systemic risk

What records should a firm establish/maintain to show a regulator?

  • Market risk policies and procedures
  • Stress testing policies and procedures
  • Circuit breaker or “kill switch” policies and procedures
  • Volatility and liquidity management plans
  • Emergency policies and procedures

NB: Business continuity considerations must dovetail with the approach on any relevant “kill switch”, so the firm can demonstrate how it will continue to function (and be compliant), and protect clients and their assets/interests, if any relevant “kill switch” was activated

System reliability and business continuity

What should a firm do?

Document the following risks, and have adequate policies and procedures in place to mitigate them: 

  • AI system failures 
  • Single points of failure 
  • Service disruptions

PLUS have in place: 

  • a robust and adequate approach to business continuity
  • backup and recovery procedures.

What records should a firm establish/maintain to show a regulator?

  • Operational resilience framework, encompassing both business and cybersecurity
  • Business continuity and disaster recovery planning and testing, including plans for AI service disruptions and system outages
  • Backup and recovery policies and procedures, including safeguards to protect client records and other sensitive information
  • System availability and performance monitoring, including related reporting and service level agreement compliance

Cybersecurity and data privacy/ protection

What should a firm do?

Document the following risks, and have adequate policies and procedures in place to mitigate them: 

  • AI system attacks/breaches
  • Client data exposure or exfiltration
  • Model theft or manipulation (poisoning)
  • Inadequate access controls
  • Other data leakage
  • Use of AI for advanced cyberattack—i.e., social engineering, deepfakes, identity theft

What records should a firm establish/maintain to show a regulator?

  • Cybersecurity policies and procedures and standards, including where relevant the cloud infrastructure associated with AI systems used
  • Penetration test framework
  • Identity and access management controls for AI systems and data
  • Incident response policies and procedures
  • Privacy impact and data privacy assessments
  • Security audit reports, logs and penetration test results
  • Continuous training of users

Outsourcing and third-party dependencies

What should a firm do?

Document the following risks, and have adequate policies and procedures in place to mitigate them: 

  • Third-party AI vendor risks, including data access and privacy risks, and cybersecurity risks
  • Inadequate due diligence
  • Poor contract terms
  • Inadequate monitoring of the third-party providers
  • Vendor concentration and dependency risk (both direct and indirect dependencies)
  • Service provider failures
  • Lack of technical skills involved in procurement process

Further checklist: 

  • Document policies and procedures on onboarding, and development and deployment controls for third-party AI use.
  • Ensure these are adequate for the risk materiality of the use case, system or model.
  • Ensure the firm gets notice of and assesses the impact of third-party AI updates and changes, and ensure contracts have clear expectations and responsibilities.
  • Ensure the firm has a right to audit third-party AI providers.
  • Ensure the firm’s outsourcing and third-party agreement inventory includes details of AI use and AI-related providers.
  • Ensure the firm’s control framework “scales” with the risk materiality of the third party’s AI use/applications.
  • Ensure the firm gets sufficient data on testing and sufficient technical details in relation to AI use.
  • Ensure the firm’s systems include due diligence (where appropriate) on fairness considerations in relation to vendor AI use.
  • Ensure the firm’s policies and procedures include supply chain risk assessments and validation, including reviews of model provenance, training data integrity, and other known vulnerabilities.
  • Ensure robust contingency plans are in place to address potential failures, unexpected behaviour of third-party AI, or discontinuing of support by vendors, particularly for third-party AI used in high-risk materiality use cases, systems or models.
  • Consider what approach to testing is appropriate, and also consider exit plans for normal and stressed exit scenarios.
  • Conduct an enhanced assessment where more complex or novel AI products and services are involved.

What records should a firm establish/maintain to show a regulator?

  • Policies and procedures on vendor selection, due diligence, and contract terms.
    • This should include policies and procedures on notice and exit provisions, service level requirements, and data protection obligations.
    • Vendor due diligence should also cover the vendor’s level of knowledge, expertise and experience on AI systems where contract terms are not tailored to the firm, an assessment of the risks associated with using the vendor’s services and the firm’s own processes to manage those risks.
  • Ongoing monitoring and oversight procedures, including performance metrics, issue escalation policies, and remedies for poor performance
  • Third-party validation and assessments of AI systems, including the identification of cross-market dependencies and potential single points of failures affecting multiple entities

Disclosure and transparency

What should a firm do?

  • Where relevant, make adequate AI disclosures to clients—e.g., inform clients when they are interacting with AI systems (e.g., chatbots, robo-advisors, or decision support tools).
  • Consider whether it is desirable or necessary to make disclosures about how relevant AI systems work, including material information about performance, limitations, and suitability.
  • Consider if consents are desirable or necessary (or opt out rights).
  • Where relevant, make adequate disclosure of the use of third-party AI services.
  • Ensure the firm does not make claims that are misleading or overstate its AI capabilities (“AI-Washing”).
  • Be aware of the need for transparency where AI systems have relevant limitations.
  • Ensure the firm’s governance and risk management practices ensure that AI systems are deployed in a manner consistent with principles of explainability, transparency and interpretability. 

“Transparency reflects the extent to which information about an AI system and its outputs is available to individuals interacting with such a system. Explainability refers to a representation of the mechanisms underlying AI systems’ operation, whereas interpretability refers to the meaning of AI systems’ output in the context of their designed functional purposes. Transparency can answer the question of “what happened” in the system. Explainability can answer the question of “how” a decision was made in the system. Interpretability can answer the question of “why” a decision was made by the system and its meaning or context to the user.”

What records should a firm establish/maintain to show a regulator?

  • Client agreements, including account information, acknowledgements, and marketing materials relating to AI usage
  • Disclosures outlining material risks and impacts on clients
  • Disclosure of incidents where adoption of AI systems has raised regulatory, ethical or legal issues
  • Client communications relating to AI usage, including policies and procedures for monitoring such communications for accuracy
  • Policies and procedures for reviewing and updating client disclosures to ensure AI-related disclosures are accurate and up-to-date

Recordkeeping and audit trail

What should a firm do?

  • Robust and sufficient AI system records
  • Documented ability to explain AI logic
  • Audit trails or logs with no omissions
  • Sufficient communications with supervisors
  • Sufficient documentation retained throughout the life cycle of any relevant AI system
  • Robust and sufficient AI system oversight throughout the lifecycle
  • Record and track modifications to AI systems over time and maintain a copy of the models used and modified over time (model versioning)

“Documentation and logs should be maintained throughout the different lifecycle stages: design, development, modification, deployment, ongoing monitoring and retirement. Such documentation and logs should facilitate the effective supervision and monitoring of the use of AI systems and support market participants’ accountability for the AI systems’ output, proper functioning throughout the AI system lifecycle, and the implementation of corrective actions when adverse outcomes occur. Any such negative outcomes should be recorded and root cause analysis performed to determine their cause and any remediation required. Supervisors may require expert specialists and additional technical tools to review these records and documentation pertaining to AI systems.”

What records should a firm establish/maintain to show a regulator?

  • Recordkeeping policies and procedures for AI systems and data, including data with a third-party or external provider
  • AI inventories
  • Recordkeeping of AI-generated outcomes and how AI systems generate outputs
  • Compliance policies and procedures for adherence to laws, regulatory requirements and internal standards
  • AI system decision and usage logs and audit trail documentation, including inputs, outputs and AI system logic
  • Regulatory filings and supporting documentation
  • Incident reporting
Footnotes:

1 www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf

2 “RAG: a mechanism that searches and retrieves curated knowledge bases or documents before generation of a response. CoVe: a technique designed to reduce hallucinations in large language models by introducing an additional verification step after the initial reasoning or generation process and forcing the model to verify its own draft before it responds. Multi-agent debate: Multiple LLMs or agents engage in mutual critique, discussion, and voting.”

Related capabilities