Banks are beginning to experiment with agentic AI in areas ranging from trading to Treasury, while humans still retain oversight over the most consequential decisions.
That makes this the right moment to ask a harder question: What happens when AI moves from helping a trader make a decision to having the authority to act on one?
Financial firms know how to test models. They measure accuracy, run back tests, simulate losses, monitor drift and impose risk limits. Those tools were built for a world in which a model produced a signal and a person, or a conventional system, decided what to do with it. Agentic AI changes that relationship.
The issue is no longer only whether a model is right. It is how much judgement we are prepared to delegate to it.
After years of trading futures markets, I have learned that some of the most important decisions are not about finding a trade. They are about rejecting one that looks perfectly reasonable on paper.
In markets, survival often depends less on knowing when to act than on knowing when not to.
A setup may have worked hundreds of times before. Yet a trader may step aside because liquidity has thinned, spreads are behaving unusually, price is moving too easily, or a familiar relationship between markets has suddenly broken down. That is judgement informed by context.
Some of that context can be coded, and more of it will become machine-readable as AI improves. But translating judgement into software creates a new problem – a system can be statistically sophisticated and still fail to recognise when the assumptions behind its decisions no longer apply.
Traditional model validation asks whether a system behaves as designed. For increasingly autonomous AI, the more important question is whether it should always be allowed to act as designed.
Consider a system instructed to reduce exposure when volatility rises. That is sensible risk management. But what happens when volatility is rising partly because other automated systems are reducing exposure at the same time?
Selling increases volatility. More systems respond. Liquidity providers become cautious. Risk limits tighten. What began as individually rational risk reduction can turn into a feedback loop.
Nothing has to malfunction. The model may have passed every test. The risk controls may have worked. Every institution may have acted rationally according to its own mandate.
The combined outcome can still be destabilising. That is why the next stage of AI risk management cannot focus only on whether models are accurate. It must also examine the authority given to them, the speed at which they can exercise it, and the consequences when many institutions automate similar forms of judgement.
This is not an argument against automation. Financial markets have benefited enormously from it. Electronic execution has lowered costs, increased speed and expanded access. But automating execution and automating judgement are different things.
Execution asks: How should this decision be carried out? Judgement asks: Should this decision be made at all?
That distinction matters more as AI systems gain the ability to plan, use tools, adapt their actions and complete increasingly complex tasks with less human intervention.
Regulators are beginning to focus on the problem. The Bank of England has noted that more autonomous AI systems are still used mainly for research, coding, surveillance and other lower-risk functions rather than widespread fully autonomous trading. It has also warned that more direct use of AI in trading and portfolio decisions could increase correlated behaviour.
The Bank, together with the Bank for International Settlements and other institutions, is already studying what happens when AI agents allocate capital and react to changing conditions in simulated markets.
That work matters because the financial industry still has a window in which to design safeguards before autonomous decision-making becomes deeply embedded in market activity.
Those safeguards should go beyond conventional model validation. A firm should certainly ask whether its AI system is accurate. It should also ask what authority the system has, how quickly it can use that authority, and what happens when markets move beyond the conditions on which it was tested.
Risk teams should stress-test successful models, not just failing ones.
What happens if the AI performs exactly as intended but liquidity disappears faster than expected? What happens if several institutions receive the same information and independently reach the same conclusion? What happens if an AI correctly identifies rising risk but its attempt to escape that risk contributes to the instability it detected?
These are not simply questions about model accuracy. They are questions about market structure and decision rights.
There is another danger: automation can create confidence faster than understanding. A trading model may have been tested against millions of observations. The statistics may be excellent. But the data still come from a world that already happened.
Markets change because participants change. A strategy becomes crowded. Liquidity migrates. Regulation alters incentives. New technology changes execution. Relationships that once appeared stable weaken.
The next crisis will not announce which historical regime it resembles. This is where human oversight still matters, although not necessarily in the way many firms imagine.
The purpose of human oversight should not be to compete with AI at processing information. Humans will lose that competition.
The human role is to question the frame around the information.
Why is this opportunity appearing now? What assumption is the system treating as permanent? Under what conditions should we deliberately do nothing?
Those questions become more important, not less, as machines become more capable.
Wall Street has spent decades learning how to automate execution. It is now beginning to automate judgement.
That transition may ultimately improve financial decision-making. But before firms give machines greater authority to act, they need to stress-test more than the model.
They need to stress-test the judgement they are delegating to it.