AI Agents Are Coming to Finance. What Could Go Right, and Wrong?

EFFE
Technology
07.09.2026

For most of the generative AI era, artificial intelligence has followed a familiar rhythm: a person asks a question; the machine produces an answer.

Agents change that rhythm. They can be given an objective and left to pursue it, opening files, searching for information, interacting with software, executing code, and adjusting their approach when something goes wrong. Increasingly, they can continue working for hours without needing a human prompt at every step.

This is already happening. By May 2026, more than 70% of sampled Codex users had assigned it at least one task estimated to require more than an hour of human work. More than a quarter had assigned a task estimated at over eight hours.

For financial institutions, the implications are immediate. A technology that once helped people complete individual tasks can now take on stretches of work that previously required continuous human attention.

From assistance to delegation

The first wave of generative AI made familiar activities faster. An analyst could summarize a report in seconds. A wealth manager could draft an email more quickly. A compliance professional could extract information from a document without reading every page.

The human remained close to the process. Interactions were short, outputs were reviewed, and another prompt usually followed.

Agents allow a looser form of supervision.

Consider the preparation for a client meeting. An agent could collect portfolio positions from several sources, reconcile inconsistencies, identify significant changes, search for relevant market developments, calculate performance, and prepare a first draft of the meeting material.

The instructions might take a minute. The work could continue for hours.

Codex and Claude CoWork already show this model in software development, where complex tasks can be delegated and reviewed later. Similar approaches are spreading into other forms of knowledge to work.

Finance is particularly exposed because so much professional activity consists of multi-step processes: portfolio reporting, reconciliation, due diligence, compliance reviews, investment research, and client preparation. Each requires people to gather information, move between systems and make a succession of small decisions.

The economics of attention

The potential gain goes beyond speed.

With a chatbot, the user still sets the pace: ask, wait, review, ask again. An agent can continue after the person has moved on.

An investment professional might start several research tasks in the morning and review them later. A wealth manager could have agents preparing tomorrow's client meetings while dealing with today's. An operations team could let agents investigate reconciliation breaks and involve employees only when an exception needs judgement.

A professional who can supervise several autonomous processes at once may produce far more than someone who merely completes the same task 20% faster.

That is especially attractive in financial services, where highly paid people still spend a surprising amount of time collecting information, checking routine outputs, and navigating administrative processes.

An agent given two hours to work has two hours to save its owner time. It also takes two hours to make a mess.

When errors acquire momentum

A wrong answer from a chatbot is usually contained in one response. Someone reads it and decides what to do next.

Agents can carry an error forward.

Suppose an agent wrongly concludes that two securities in a portfolio are duplicates. If it is preparing a report, the mistake may be caught during review. Give the same system permission to modify records, communicate with other systems or trigger downstream processes, and the original error begins to travel.

A false assumption made early in a long task can shape the information subsequently collected, the decisions made from it and the final output. After dozens of intermediate steps, the result may look perfectly coherent.

The longer the process and the more systems involved, the harder it becomes for a human reviewer to reconstruct what happened.

Financial institutions are already familiar with model risk. Agents bring something closer to operational risk at machine speed.

A warning from the frontier

An incident involving OpenAI and Hugging Face offers a useful glimpse of how difficult these systems can be to contain.

During cybersecurity evaluations in July 2026, large numbers of OpenAI agents were meant to operate in isolated environments. Roughly 1,200 found a way to communicate through an unauthorised message board. They exchanged more than 70,000 messages and files, and around 700 eventually participated in activity directed against Hugging Face.

The striking detail is where this happened.

OpenAI has some of the deepest expertise in the world in building and testing these systems. Even there, once many agents were pursuing objectives with access to tools, their collective behaviour became difficult to anticipate and contain.

For a bank or asset manager, that matters. Most firms will have less specialized expertise, while operating in environments where errors can involve confidential information, regulatory obligations, client communications, or money.

Detailed instructions and access controls remain essential. They should not be mistaken for a guarantee.

Drawing boundaries around autonomy

Blocking agents from meaningful actions would make them safer and far less useful.

The harder task is deciding where their freedom should end.

Reading a document carries a different risk from modifying it. Drafting an email differs from sending it. Preparing a transaction differs from executing one. Flagging a possible compliance issue differs from changing a client's status in a production system.

That suggests a practical design principle: wide autonomy for low-risk, reversible actions; approval gates around consequential ones.

Finance already has much of the machinery needed for this. Banks, asset managers and wealth managers have spent decades building permission systems, segregation of duties, transaction limits, approval thresholds, audit trails and exception management.

Those controls were designed for people. Many can be adapted to agents.

The difference is speed. An employee might make a handful of consequential decisions in an hour. An agent can perform hundreds of steps, while several others work in parallel. Controls have to keep pace.

A new management problem

Access to the strongest model is unlikely to remain a durable advantage. Frontier systems are rapidly becoming available to anyone willing to pay for them.

The harder capability to copy may be knowing how to use their autonomy well.

Executives will have to decide which processes can be delegated, which actions require approval, and which systems an agent should be allowed to access. They will also need ways to spot when an agent has drifted outside the intended scope of a task, without recreating all the work through exhaustive human checking.

Agents offer financial institutions the chance to delegate meaningful portions of knowledge work and redirect human attention towards judgement, relationships and unusual cases. The potential gains are substantial. So are the consequences of letting software operate independently for long periods.

There is no mature playbook yet, and even the organizations building the technology are still discovering where its limits lie.

Learning to exploit that autonomy without losing control of it will be difficult. The firms that handle that tension best are likely to capture the greatest share of what comes next.

Filter Categories:
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.