---
title: "AI playbook for finance"
description: "The 2026 World Economic Forum AI playbook for financial services, read against the bank projects I have worked on. What six institutions actually shipped, the two numbers that draw the line for agentic finance, and why the ones who got it right never made the model the product."
seo_title: "AI in finance: what six institutions shipped"
seo_description: "71 percent of banking customers want an AI assistant, 82 percent want to approve its actions. What six institutions measured, and what they built instead."
slug: ai-playbook-for-finance
status: published
published_at: 2026-08-23
author: Arman Obosyan
author_url: https://sugra.systems/about
section: general
primary_keyword: world economic forum ai playbook financial services banks agentic ai
hero_image: /blog/images/posts/ai-playbook-for-finance-hero.jpg
hero_alt: "A great banking hall drawn as fine amber wireframe: vaulted arches, columns and rows of ledger desks all glowing inside, with a single thin amber thread running out through the doorway into empty darkness. Title: AI in Finance. 71 want it, 82 want to approve it."
og_image: /blog/images/posts/ai-playbook-for-finance-hero.jpg
tags:
  - ai
  - agents
  - finance
  - engineering
---

A former colleague from my banking years put the World Economic Forum and Accenture's **AI Playbook for Financial Services** on LinkedIn. It caught my attention, and I had a ten-hour flight ahead of me, so I read all forty-nine pages of it.

I read it arguing. By the last page the argument had turned into this post, and NVIDIA's **State of AI in Financial Services** kept pulling at me afterwards, because the playbook cites it constantly and the two do not always agree.

Between them they settle a disagreement I have had with banks for as long as I have worked with them. Everyone in this industry is spending on AI. Most of it still lands inside the institution before it reaches the customer. The numbers that prove it are sitting in both documents, and neither one says it in those words.

I came up through technology and then through finance, I have worked on projects like these from the inside, and I now run a company built on the parts of this problem most teams still leave unresolved. That is the lens. Here is what matters.

## The two numbers that draw the line

Everything else in the playbook sits on top of one pair of findings.

**71 percent** of banking customers worldwide say they would welcome an AI assistant inside their bank's app.

**82 percent** want to approve the agent's actions before they are executed.

![Two figures side by side: 71 percent of banking customers would welcome an AI assistant inside their bank's app, and 82 percent want to approve the agent's actions before they are executed](/blog/images/posts/ai-playbook-for-finance/what-customers-want.jpg)

Read together, the two numbers define the product constraint. Customer interest is not the primary obstacle. Unattended authority is. Customers are not rejecting the assistant. They are rejecting an open-ended mandate.

So the product is not an autonomous agent. It is an agent that explains what it intends to do, pauses at a consent boundary, and executes only within the mandate the customer grants. That decides more architecture than any model choice, and if you cannot produce the record of intent, you have nothing to put in front of the 82 percent.

## What actually shipped

The playbook's value is that it is not theoretical. Eight institutions put their numbers in it. Six are worth your time:

| Who | What they changed | Measured result |
|---|---|---|
| International Bank of Azerbaijan | Automated underwriting for micro-business borrowers with thin credit files | 21 percent portfolio yield, non-performing loans under 4 percent |
| Allianz Partners with Taktile | Specialized agents per step of the health claims journey, humans on exceptions | Claims processing from days to minutes, 1.5 times better fraud detection |
| Mastercard Ethoca | Enriching transaction descriptions that customers could not recognize | 15 percent higher accuracy, 20 percent more records, 87 percent less processing time, 92 percent lower cost per record |
| EBRD | 300 evaluation reports and 20,000 pages into one grounded, citation-backed corpus | At least 20 percent time saved, 66 percent more lessons-learnt reports, adopted by 10 other institutions |
| StockGro (India) | Multi-agent investment research for retail investors | Four hours of manual research compressed under 10 seconds, research priced 98 percent below the market |
| KBTG (Thailand) | Company-wide AI literacy and a sandbox where staff build their own agents | 200 ideas, 60 prototypes, 8 scaled, roughly 30,000 workdays saved |

One thing to hold on to before we go further: none of these was done for the sake of the AI. Each one is a specific hole that was previously filled with people, closed and then measured in the units that business already used. Yield. Non-performing loans. Cost per record. Workdays.

## The ones who got it right did not treat the model as the product

Line them up against a harder test than "did it work". Did it reach the customer, and did it leave behind a capability rather than a demonstration?

**Azerbaijan is the strongest case in the report**, and it is not close. ABB did not improve an existing process. They opened a segment that could not be served at all, because manually underwriting a micro loan costs more than the loan earns. They built a machine-learning scorecard on internal transactions, bureau data, and cross-bank turnover parsed out of PDF statements, which is the least glamorous work in this entire document. Then they put it inside the lending platform and let it decide, set the limit, price the risk and disburse, end to end. Twenty-one percent yield, and non-performing loans under four percent, which is the number that proves it was not recklessness dressed as inclusion.

There is no foundation model anywhere in that story. A market nobody could serve was opened by a scorecard and some data plumbing.

**Mastercard** fixed something the customer actually feels, and kept people on the exceptions with confidence scoring and AI judging agents behind them. They also mixed classical machine learning with generative models in one product instead of replacing everything that worked with an LLM. That is engineering discipline, and it is rarer than it should be.

**Allianz** built modular agents that can be reused across products and geographies without rebuilding the workflow. The reuse is the tell. A demo does not need to be modular.

**EBRD** deserves a mention for the opposite reason: it is small, cheap and unglamorous. Three hundred documents, twenty thousand pages, answers grounded only in official evaluation material, citations mandatory. It is exactly the use case that 56 percent of NVIDIA's respondents name as their first workflow for agents, and EBRD is the one that shipped it with the citations attached, which is the part most teams skip.

I am more skeptical about StockGro than the report is. The reach is real and the numbers are impressive, but "85 percent improvement in decision accuracy" for retail investment guidance is the least verifiable claim in the document. Accuracy against what benchmark?

And "more than 80 AI agents" describes scale, not evidence. We have measured that exact question on ourselves. Over six months of multi-model review we recorded how much each reader actually caught, and headcount turned out to be the weaker variable. In our operational, non-random sample, the same model at a higher effort setting was associated with roughly twice as many serious findings per review, and the pattern held across three model pairs, including one on a slice of easier changes where the difficulty of the work cannot explain the gap. Turning up the effort on a reader we already ran bought more than adding another reader did. [The full numbers are here](/blog/527-ai-code-reviews).

Eighty of anything is a fleet size. It says nothing about how carefully any one of them reads, and in our data diligence is the variable that moved the result.

Every manager already knows this version of the problem. Forty people who skim do not outperform ten who read properly, and hiring the forty-first is always easier than getting anyone to work more carefully, which is exactly why it keeps being the option that gets chosen. Agents inherit the rule unchanged. The only thing that changes is that the payroll arrives as a token bill.

The common thread in the four that convinced me is not that they avoided models. **They avoided making the model the product.** They built a decision, a data product, a workflow and a grounded corpus. The strongest case opened an entirely new customer segment with a scorecard.

## What I see in the field

I have worked on projects like these under NDA, so there will be no names and no detail that identifies anyone. The pattern is what matters, and the pattern repeats.

Banks are spending serious money on AI right now. Most of that spending stops inside the building.

That is not a cynical reading. It is what the surveys say about themselves. Ask NVIDIA's respondents what AI actually improved and the answer is operational efficiency at 52 percent, employee productivity at 48 percent, and customer experience third at **37 percent**. Split it by segment and it gets sharper.

![Grouped bar chart of what AI improved by segment. Overall: operational efficiency 52 percent, employee productivity 48, customer experience 37. Fintech: 40, 53, 40. Consumer finance: 59, 43, 57. Capital markets: 53, 53, and customer experience only 27](/blog/images/posts/ai-playbook-for-finance/improved-by-segment.jpg)

Customer experience is simultaneously the most common place to deploy AI, at 42 percent, and not the best return, where document processing wins at 32 percent against customer experience at 30. In capital markets it is not close: the two internal measures sit at 53 percent each, and the customer gets 27.

The Forum says the same thing in politer language. Most institutions prioritized what it calls "no regret" investments: internal process improvement, copilots, basic training. Its own maturity model puts customer experience in the second stage, and most of the industry is standing in the first one, next to the words "back office" and "cost reduction". It is the same divide [we surveyed across the whole industry earlier this year](/blog/state-of-ai-2026), arriving now with a sector attached.

The reason is not mysterious. Inside the perimeter, a win is easy to measure and a failure hurts nobody outside the building. The customer surface is brand risk, regulatory exposure, and those 82 percent who want to approve every action. So it gets deferred, and it keeps appearing in strategy decks as the priority it is not.

## Why they build their own models

The other thing I keep meeting is the institution that decides to build its own model.

Ask why, and the answers rarely survive a second question. Two of them are honest. Data residency and compliance is a real constraint, and 36 percent of NVIDIA's respondents rate it among the most important factors when running inference. Unit economics at volume is the other: when your token bill is large enough, owning beats renting.

The rest comes down to status. A proprietary model is the most visible thing an institution can point at, it photographs beautifully in a board deck, and it carries an unspoken claim that we are the smartest people in this room. In my experience it is usually built out of not knowing what else to do.

The problem is that it answers a question nobody asked. Look at what those same institutions report is actually blocking them: **data issues at 40 percent**, up from 33 the year before and now the single biggest challenge. **People who can manage and monitor agents, 33 percent.** **Output that drifts and cannot be predicted, 34 percent.**

A model of your own fixes none of those three. It is the most expensive available way to avoid the unglamorous work.

One number makes the point better than my opinion does. Asked how important open source is to their AI strategy, 48 percent of management said very or extremely important, against 35 percent of the practitioners who would actually have to run it. When the executive floor is more radical than the engineers, the decision is usually not being made on engineering grounds.

And the playbook contains its own rebuttal. Candidly, which builds financial guidance for student borrowers, reports that frontier language models fail roughly 60 percent of realistic financial analyst tasks, and that 43 percent of their answers to student loan questions are wrong or misleading. Their fix was not a bigger model. It was encoding expert reasoning as auditable, human-authored modules with deterministic calculation underneath. The report's conclusion, in its own words, is that accuracy in regulated domains comes from domain expertise, not model scale.

This is the same failure we have measured on live data ourselves: models are [confidently wrong, and search does not fix it](/blog/ai-confidently-wrong-live-data).

## What a bank is actually buying

There is a third thing institutions buy instead of solving the problem, and it is the most interesting of the three, because unlike a home-grown model it is frequently the right call.

When a large bank runs its AI on Microsoft Foundry or AWS Bedrock rather than renting accelerators directly, it is not buying compute alone. Compute is the cheap part. It is buying one model catalogue, access control wired into the identity system it already runs, a private network path, logging, policy, monitoring, and a single supplier that has already survived procurement, legal and the risk committee. That is a paid layer of organizational convenience, and for a regulated institution the convenience is real.

The price of it is visible.

![Bar chart of the on-demand price of one H200 GPU-hour on 22 August 2026: Foundry Managed Compute 10.60 dollars in the cheapest US regions and 13.78 dollars in Sweden and the UK, RunPod secure tier 4.59 dollars, Nebius 4.50 dollars](/blog/images/posts/ai-playbook-for-finance/h200-price.jpg)

A hundred hours costs $1,060 against $459 or $450, so roughly two and a third times the price for the same accelerator model at on-demand list prices.

Here is the part I got wrong before I read the meter. That premium is not the managed service. Inside Azure the plain virtual machine costs the same per card, to the cent, and for a small model the managed service is far cheaper, because the machine makes you rent all eight cards while Foundry sells you one.

So the bank is not paying a premium for the managed layer. It is paying to be inside that perimeter at all, and that is a defensible purchase. Microsoft's own customer stories show the logic working: Bradesco reports an 83 percent resolution rate in digital customer service with technology costs down more than 30 percent, and Commerzbank's assistant handles more than 30,000 customer conversations a month and resolves three quarters without a human. Both are vendor-published rather than audited, and the second is a fair counter-example to my own complaint, because it reaches the customer.

The waste starts later, when convenience hardens into architecture. The pilot ends, the load becomes predictable, and the managed rate keeps being paid on autopilot. A small model gets an H100 because that was the template. An endpoint runs around the clock because nobody turned it off. Cost per useful answer is never calculated, and no alternative is tested, because a new supplier means another audit. The risk premium quietly becomes a bureaucracy premium.

The industry knows this and has started moving. In NVIDIA's numbers, pure cloud fell from 57 percent to 42 while hybrid rose from 26 to 47, and total cost of ownership is now named by 34 percent as a decisive factor in where inference runs. Measure convenience by all means. Just do not stop measuring utilization, cost per result, and what leaving would cost.

## Where the two reports contradict each other

Read them side by side and two contradictions fall out.

**The first is about data.** The Forum's entire architecture rests on a data foundation: structured and unstructured enterprise data turned into governed, reusable products so that agents can reason over something trustworthy. Meanwhile, in NVIDIA's numbers, attention to data processing has fallen three years running, 42 to 41 to 35 percent, while data problems climbed to first place at 40 percent.

![Two panels on the same scale. Left, what respondents pay attention to across three annual surveys: data analytics rises 56 to 57 to 68 percent, generative AI rises 40 to 52 to 61, data processing falls 42 to 41 to 35. Right, what blocks them: data-related issues rise from 33 to 40 percent while not having enough data to train on falls from 49 to 31 to 16](/blog/images/posts/ai-playbook-for-finance/attention-vs-pain.jpg)

Attention is going down while the pain is going up. That gap is where the next few years of expensive failure will come from.

Underneath it is a subtler shift worth noticing. The complaint that there is not enough data to train on collapsed from 49 percent in 2023 to 31 in 2024 to **16 percent** in 2025. Institutions solved volume. What they did not solve is provenance, residency and the fact that the data is scattered. The problem changed shape and a good deal of spending is still aimed at the old one.

**The second is about agents.** 42 percent of respondents say they are using or assessing agentic AI. Strip out the assessing and **21 percent have actually deployed any**. Another 18 percent say next year.

What is stopping the rest is not ambition. The top challenge with agents is performance reliability, meaning accuracy drift and unpredictable results, at 34 percent, followed immediately by not having anyone who can manage and monitor them, at 33. In other words the hard part is not building an agent. It is knowing whether what it just produced is right.

## What I recognize from our own work

Which brings me to the part where I stop being a reader.

The playbook lists, among its mitigations for model risk, "deploying multi-agent systems to review the accuracy of outputs and verify them against defined standards and other reliable data sources". We have been doing exactly that for six months and writing down every verdict, because [this company is built by agents that review and block each other's work](/blog/ai-replace-engineering-teams). Across 527 recorded reviews over 227 changes, **one in four changes carried a serious defect after being written, tested and believed finished**. Not draft code: changes whose author considered them done. If you are going to put agents in front of regulated decisions, that is the number I would want you to sit with.

The playbook also asks for explainability and human oversight on credit scoring and similar decisions. The operational version of that is less abstract than it sounds. Screen a company by name against a sanctions corpus and you get ten candidates and a job for a human. Screen the tax identifier and you get an answer. [The difference between those two calls is the whole product](/blog/sugra-entity). It is what makes a decision auditable instead of merely automated.

On natural language becoming the working interface, we did the small version of that on this site: search and an Ask AI mode that read the blog and the platform pages together, so the answer comes with the page it came from.

And on the AI colleague the reports keep describing, I have one. At SkyTel we built Ana Lytics, an AI employee for analytics, deliberately narrow: she knows the Communications Commission's regulation and the national statistics data, she is grounded through retrieval rather than memory, and she shows up as a positive line in the P&L, because she closes a domain of expertise that used to mean pulling skilled people off their real work. [That story is here](/blog/brave-new-world-95).

Sugra itself is intelligence infrastructure rather than a finance product, but finance is where the requirements first got concrete: provenance on every value, a citation that resolves, live data rather than a model's recollection of it. Those are not features we invented. They are the things the playbook's data layer is asking for, [written down as a product](/blog/platform-intro).

## What I would do on Monday

Five things, and none of them require a new strategy document.

1. **Pick one customer-visible process and finish it.** Not a pilot. The reports are full of institutions with impressive internal metrics and nothing the customer would notice.
2. **Buy the model, build the domain.** Candidly's numbers are the argument. Your advantage is your data and your rules, not your weights.
3. **Budget for verification, not just generation.** Reliability of output is the number one complaint about agents, and verification has to be independent of the generation path. Depending on the decision that means a different model family, deterministic controls, source-backed rules or human approval. Whichever you pick, depth of checking beats the number of checkers.
4. **Fix provenance before autonomy.** Data problems are the largest and fastest-growing obstacle, and volume is no longer the issue. Where each value came from is.
5. **Design for the 82 percent.** Build the consent boundary and the record of intent first. It is the difference between an assistant customers want and one they switch off.

## Sources and method

The playbook is [The AI Playbook for Financial Services](https://www.weforum.org/publications/the-ai-playbook-for-financial-services/), World Economic Forum in collaboration with Accenture, June 2026, drawing on more than 150 senior leaders across more than 100 organizations plus roundtables in Hong Kong, New York, London and Singapore.

The survey is NVIDIA's [State of AI in Financial Services: 2026 Trends](https://www.nvidia.com/en-us/industries/finance/ai-financial-services-report/), its sixth annual edition, fielded from August to September 2025 with 839 respondents split evenly between management and AI practitioners. One caveat that report states plainly and that you should carry with you: the sample was drawn from NVIDIA's own distribution lists and social channels. These are people already close to the technology, so read the adoption figures as the leading edge of the industry rather than its average.

Every chart here is rebuilt from the underlying values rather than reproduced, and each was checked against the source page it came from.

GPU rates were read on 22 August 2026 from the Azure Retail Prices API for Foundry Managed Compute and for the equivalent virtual machines, and from the published on-demand pages of RunPod and Nebius. The Bradesco and Commerzbank figures come from Microsoft's own customer stories, which are marketing material rather than audited disclosure, and are labelled as such above.

Everything attributed to our own work comes from our records, and the review numbers are published with their method attached.
