I hosted two executive roundtables on generative AI and agentic systems in financial services this year, and what struck me wasn’t just the pace of change. It was where the hard problems are moving towards.
If you’re a tech leader responsible for AI outcomes (or carrying part of the load), that shift matters. Financial services is a high-pressure testbed: strong incentives to move, tight regulation, and very real consequences when things go wrong. The patterns showing up here are early indicators of what most industries will hit next. I’m sharing them because hearing the same constraints from different organisations, months apart, gives you a clearer read on what’s changing and what to do about it.
In June, the room was still living in the messy middle: that stretch between promising pilots and something repeatable. The blockers weren’t imagination; they were foundations: getting internal generative AI platforms to scale reliably, getting copilots trusted enough to stick, and making governance plus safe data access work in a regulated environment. Cost and risk were already front of mind, but mostly through a scaling lens: orchestration, compliance, security and data integration are where programs get heavy, not in the model itself.
By November, the baseline hadn’t vanished, but the ambition had expanded and the hard problems had moved up the stack. Instead of trading thin use cases, leaders were talking about re-imagining whole value pools, then slicing small and shipping toward a clear destination. And with that ambition comes the next constraint. As systems move from answering to doing, the hardest questions stop being technical. They become organisational: who owns an agent’s behaviour, where accountability sits, and whether teams and customers are ready to trust autonomy. The tech is quick; the trust and operating model work is not.
Bigger bets and the reality of running them
At our most recent roundtable the conversation had more altitude. The focus had shifted from thin, tidy use cases to the big, stubborn domains where generative AI will actually earn its keep. The nuance mattered: going after a domain doesn’t mean boiling the ocean. You set the destination, then ship small, auditable slices toward it.
Financial services isn’t just experimenting anymore. It’s starting to lead. Even with regulation as a constant brake, larger institutions are already a year or two into domain-level programs, while smaller players are compressing multi-year transformations into months. That mix of scale, urgency and scar tissue is making the sector a proving ground for what “serious” generative AI looks like.
A practical pattern is enabling that ambition. Leaders are using a one-page working-backwards brief, a “future press release” / PRFAQ-style artefact made popular by Amazon, to align a coalition around a clear customer outcome at close to zero cost. If the outcome doesn’t feel compelling on paper, you’ve saved expensive learning. If it does, you’ve got a north star and a team before you write code. Because these briefs are cheap to create, teams can write several and let the strongest ideas pull a coalition around them. That’s prioritisation at speed.
What this means for you
-
- Pick a domain outcome first, then slice delivery toward it.
- Use working-backwards briefs to prioritise quickly and build sponsorship early.
- Expect the hardest work to sit in trust, ownership and adoption, not in the model.
That escalation to bigger bets also surfaced a new kind of trepidation. Use cases feel safe because they’re bounded. Domains feel risky because they’re organisationally messy. The group’s answer wasn’t to retreat into smallness, but to accept there’s no settled best practice yet. Get clear on direction, learn fast and keep moving to the next safe place.
Agents are easy, trust is hard
Once the room got into the agentic end of the spectrum, the tone sharpened. You can prototype an agent in hours. Making it production-safe in financial services takes weeks. Most effort goes into the trust envelope: guardrails, validation datasets, monitoring and observability And it has to exist before anything goes near customers.
A concrete edge of that trust work is advice and exploitation risk. Users will try to push an agent over the line, asking it to ignore guardrails or slip into recommendations. The system has to detect that drift, stay factual and refuse cleanly even when the user insists it’s fine. That refusal behaviour, plus the test sets and proofs behind it, is where the real engineering time goes.
Ways of working are changing too
Teams already building these systems also shared a pragmatic lesson: more agents doesn’t equal more capability. Early multi-agent setups ballooned into brittle, confusing systems. Over time, people simplified: fewer agents, clearer responsibilities, less latency, fewer failure modes.
Another thread that got louder was how quickly the shape of work is changing. Prototyping is dramatically cheaper and faster. The lines between product, design and engineering are blurring. People are building their own helpers. At the same time, high-level process redesign still needs strong human expertise to steer it. Your operating model has to assume faster cycles and more cross-disciplinary building.
Ownership is the real bottleneck
And then we hit the real November inflection: ownership. As systems move from answering to doing, accountability becomes the constraint. The group returned to a simple idea: an agent is like a staff member whose seniority increases as trust grows. Someone has to own it, sign off on it and be accountable for decisions, especially when automation crosses organisational boundaries. If ownership is fuzzy, progress stalls even when the technology is ready.
Finally, autonomy isn’t just about internal confidence. Customers may not be ready for what you’re ready to ship. One example shared in the room was technically strong, but landed as “scary” when shown to users. The takeaway wasn’t to slow down; it was to involve customers early enough to test readiness alongside value.
So if June was about making scale possible, November was about what happens when you try to run generative AI at domain scale: bigger ambition, heavier trust work, shifting roles and operating model questions moving to centre stage. The constraint has moved from “can we build it?” to “can we run it safely, with clear ownership, at domain depth?”
Where this leaves you
If you strip both roundtables back, they leave tech leaders with two hard questions that don’t go away:
-
- what are we really prioritising?
- what kind of platform are we building to support it?
On prioritisation, the rooms were pretty clear: stop treating use cases like a shopping list. The work is to pick one or two domains that matter over a ten-year horizon, write down the future you’re aiming at, and then let that vision compete.
In financial services, those domains usually aren’t the shiny new features. They’re the journeys and processes everyone has learned to live with: lending and onboarding flows that haven’t moved in years, claims processes stitched together through workarounds, fraud and scam handling that still leans heavily on humans. They’re big, boring and commercially critical, which is exactly why they’re where generative AI can earn its keep.
A working-backwards brief is cheap. You can write several for different domains, see which ones attract energy and sponsorship, and use that signal to decide where to spend real money and talent. If you get that call wrong, everything downstream is an expensive distraction.
On platforms, the lesson was that “platform” is the shared scaffolding that makes each new thing easier than the last. Orchestration, guardrails, validation data, monitoring, safe routes to production, patterns for human-in-the-loop: that’s the platform. It’s rarely glamorous. But it’s why some organisations can stand up a new agent in days, while others are still treating every experiment as a one-off.
So if you’re feeling the constraint move in your own organisation, that’s not a bad sign. It’s a signal that you’re past the first wave of experiments and into the real work: choosing the long-neglected domains that actually drive value, and building the boring foundations that let you keep shipping safely.
Get those two right: what you’re aiming at, and what you’re building on. And the rest of the mess starts to look a lot more manageable.
And if you’d like a thinking partner on that journey, from writing the first working-backwards brief, to choosing the right domain bet, to putting a real platform and trust envelope underneath it, this is the work we do every day. We’re always happy to roll up our sleeves and help you figure out what to do next.