AI & governance
AI in banking: Why submission-ready results do not come from the model alone.
Around 500 investment bankers evaluated AI-generated work products from workflows close to real investment banking practice. **Not a single output** was deemed ready to submit without changes. Forty-one per cent required extensive revision; 27 per cent were unusable. The benchmark did not test simple chatbot replies, but Excel financial models, PowerPoint decks, PDF reports and Word memos, the artefacts that in practice go to supervisors or internal decision processes. A concise public write-up, including at The Decoder, summarises the results.
Clean text is not a working model
Many debates confuse plausible language with dependable work. A polished paragraph is not yet a reviewed file-based deliverable. A deck can look convincing quickly. A financial model still has to **calculate**, with traceable formulas, consistent assumptions and scenario capability. Where key figures are pasted as constants instead of being derived through formulas, usability for real banking work collapses, even when slides look coherent.
BankerToolBench: what the benchmark measures
BankerToolBench was developed by Handshake AI and McGill University as an open-source benchmark. Tasks reflect typical junior-banker work: searching data rooms, using market-data platforms, analysing SEC filings and producing **several linked files**. Scoring used rubrics defined by experienced bankers, on average many criteria per task, covering technical correctness, client readiness, adherence to instructions, traceability and consistency across files.
Where models fail, and why that matters for institutions
Failures were not limited to wording. Models broke down on formulas and code, on domain logic, on consistency across files, on aborted data pulls and sometimes on **fabricated numbers** presented as sourced. A test of this kind matters for banks because it does not measure whether AI sounds good. It measures whether AI can **deliver work** that holds up in a regulated, data-intensive setting.
Systems, not the chatbot
The problem is not the chatbot as such. It is the assumption that generative systems can produce professional banking work without tight steering, a clear data foundation and defined accountability. In financial services, approximation is not enough, neither for internal steering nor for anything that might later face clients or committees.
Client contact does not start in the meeting
Client contact does not start in the conversation. A memo, a pitch or a model becomes client-close as soon as it feeds a decision template. Errors are then concrete: they affect judgements, recommendations and transactions. Direct outward use remains the hardest case. There, technical accuracy, traceability, regulatory permissibility and internal accountability apply at once. A model can generate text. **It does not carry responsibility.**
Supervisors: FINMA, BIS and FSB
FINMA speaks to the gap between experiment and production. It expects supervised institutions to maintain governance and risk management, to inventory and classify AI use and to assign clear responsibilities. For externally sourced solutions it highlights particular challenges, transparency over data and methods used, and adequate due diligence on providers. That is not red tape. It is the frame that determines whether use stays controllable.
International bodies do not see risks only “in the model”. The Bank for International Settlements (BIS) stresses in financial-stability work that AI, without adequate controls and supervision, can amplify existing vulnerabilities. The Financial Stability Board (FSB) names third-party dependencies, cyber risk and challenges around model risk and governance, in line with what benchmarks such as BankerToolBench surface in concrete outputs.
Priority: controlled internal processes first
For banks, asset managers and family offices, the priority is sober: put AI first in **controlled internal** processes. There it can research, structure, summarise, compare and prepare. It can relieve specialists. It should not replace them prematurely, and not where outputs move toward clients or committees without human review.
The most productive uses are often below the surface: meeting prep, document analysis, internal knowledge search, quality checks, meeting notes and first drafts. Less visible, more robust.
The bottleneck is architecture, not model choice
The real bottleneck is architecture. Data sources must be vetted. Roles must be fixed. Outputs need review paths. Approvals must be documented. Escalations must be clear. Without that structure, AI produces speed without reliability.
BankerToolBench does not show that AI is useless in banking. It shows that **submission-ready work does not come from a model alone**. It comes from a system of data, process, control and professional accountability.
In financial services, the winner is not the tool that promises the most. It is the structure that delivers reproducibly correct outcomes.