Re-shaping (public sector) user centred design for an AI assisted world
New tools may shape a new process for user centred design
TL:DR
The rise of AI assistance across all aspects of digital teams requires us to question many of the processes and practices we’ve championed over the years in the public service (and elsewhere).
For those not keen on a long read, here is the TL;DR.
Things I am no longer precious about:
A front-loaded ‘discovery’ phase for research
Dedicated research headcount per team
Things I am now/continue to be precious about:
Teams exploring (with customer data) multiple approaches to solving customer problems.
Research as a team sport
Exposure Hours
Research focussing on people who are actually using the service
Close collaboration with behavioural analytics
Developing and communicating a shared context of users
A few notes before you get too annoyed
I use the terms user and customer, product and service almost interchangeably throughout this essay mostly as a way of acknowledging that both exist without having to repeatedly list them. Bear with.
I’m also aware that although the future is here, it is not evenly distributed and for many teams working within the public sector and elsewhere, AI assisted work is often not even on the radar yet. Think of this much more as what is coming rather than what you probably need to react to right now.
If you’re working in the private sector though I do hope you’re having a good think through most of this already.
I would love to hear if your thinking and experience aligns with this or if you have another perspective.
Introduction
I recently had the opportunity to read Ben Welby’s The Future of (Public Sector) Product Management in a Vibe-Coded World where he outlines a possible shift in how public service teams work in response to the rise of vibe-coding, or what Ben more accurately describes as AI-assisted delivery.
This interests me for a few reasons.
Firstly, I was involved in government service delivery prior to the rise of AI both in the UK (as part of the Government Digital Service (GDS)) and later within the Australian Government.
Secondly, I am right now wrestling with the implications of AI and LLMs in particular for my particular specialist area: helping teams to understand their users and those users’ needs, and to design services that make meeting those needs as easy and efficient as possible for both the user and the organisation delivering the service (otherwise known as user research).
It is early days on this journey and none of us knows for sure what the ‘right’ answer will be. Much of what Ben sets out I agree with, but there are two main areas I want to dive into to offer a different or perhaps extended speculation, based on my own set of experiences and expertise.
The first is related to the stages of service delivery, and the second is what research in an AI-assisted team might look like.
Building is cheap. Shared understanding isn’t.
Many AI commentators have noted that as building software gets cheaper, judgement becomes more valuable. This judgement directs a team to build a particular thing in a particular way with the belief that it will meet the market (or meet user needs).
This judgement operates at all levels of decision making. From product strategy down to the details of the interaction design.
Right now, and for the foreseeable future, AI works extremely well to assist the team with building whatever is decided and documented at an accelerated pace. It is arguably less good at working out what should be built and how it should be designed to optimise for good user experience.
I think this is probably fine though. Strong opinions on what to build seems to not be something in short supply amongst us humans. Most teams don’t get spun up until well after someone has somehow determined, to a greater or lesser degree of detail, what the solution is. Which tends to happen well before the discovery phase kicks off.
Dealing with these preconceptions is generally an undocumented part of the User Centred Design (UCD) process.
Interrogating the phased approach
Decades ago, when I was a baby UX Researcher, there was a rule of thumb we used to talk about to evidence the importance of testing ideas as early as possible in the design process. It went like this: making a change will cost you:
$1 on paper,
$10 in a prototype, or
$100 in live code.
The moral to the story was to make all your mistakes as early as possible because getting it wrong once the code was written was expensive.
The UK Government Service Manual and Service Standard and the proscribed phases of the agile project - Discovery, Alpha, Beta and Live - are at least in part informed by this truth, with the goal of maximising the likelihood of delivering a service that works for users whilst minimising the cost of the service to the taxpayer.
In his essay, Ben re-commits to this approach. He writes:
For more than a decade, the rhythm of Discovery, Alpha, Beta and Live has structured how digital public services come to life in the UK. That rhythm remains vital: civic scaffolding that helps us learn safely, stage by stage, on the path from curiosity to confidence.
In a vibe-coded world, where ideas can be prototyped as deployed code before a meeting ends, the meaning of those stages deepens. They are not hurdles to clear, but lenses for understanding: of users, of systems, and of what it takes to deliver a safe, sustainable service.
I think this is probably where Ben and I diverge most in our educated guesses about the future. Mostly because I’ve spent the last few months coming to terms with the thought that it is likely time to stop front-loading research by default. Otherwise known as the Discovery Phase.
It has always been difficult to get teams to commit to a true discovery. One where we genuinely hold our hands up and say we do not know what the answer is, or are completely open to what we might learn and how that might shape what service we deliver in response.
Almost always someone senior pretty much knows that they’re expecting to be delivered. And the team knows that’s expected of them too. And everyone is pretty keen get stuck into building it and making it real.
Now, sometimes this is a problem because the originating idea is not good. But plenty of other times, there are people who really do have a great intuition for how to combine opportunities in design and technology with a good enough understanding of the problem that users are encountering.
If we’re honest, much of the reason GOV.UK exists is down to the latter case, and we’re all the better for it.
In my recent review of the Bureau of Meteorology website redesign , I saw research artefact after artefact that represented weeks and months of discovery research work (and probably hundreds of thousands of dollars) that contained insights that would realistically have surfaced from an afternoon spent in a workshop with the right people making educated guesses. And making a list of the few areas they genuinely felt were high risk unknowns.
Anyone who has worked in large scale government digital services knows the BoM project isn’t unique. This isn’t an unusual example.
Meanwhile, these days, someone could often bang out a pretty decent working prototype in an afternoon.
This isn’t completely generalisable. There is a big difference between helping people check the weather forecast and many of the other much more specialised or high stakes activities that government digital services must support. There are times when a sensible team would make a call that some upfront discovery work is the responsible thing to do. But rather than requiring a full discovery in advance as a rule, perhaps we let teams decide if and when discovery is required. Or whether, instead, the most productive thing is to gain consensus on what we think we know about the user and get cracking exploring ways that might best meet their needs.
Assuming that tokens don’t become entirely unaffordable, and that AI assisted software development does remain relatively inexpensive - does it still make sense to require all teams to pass through a significant up front UX Research Discovery barrier? This is no longer the hill I’d die on.
Alpha/Beta Forever
If we’re open to dropping up front Discovery from the project lifecycle, then what might a good AI-assisted team process look like?
For me, it’s something like a permanent alpha/beta phase. Continuous Discovery and Beta forever - where Alpha means we routinely explore and test multiple design solutions and Beta means we remain in continuous iteration and improvement. We learn continuously from real people using our real service, and we are never ‘done’.
In theory, we know that the Live ‘stage’ was meant to be ‘keep iterating and learning’ but in reality we know it really means ‘Done’. The team rolls off and almost every one rolls onto something new.
This can be fairly productive when you’re in a ‘transformation’ mindset - blowing up old stuff and replacing it with something more modern. But as our digital capabilities mature, the goal should be to move beyond ‘project’ and ‘transformation’ thinking and towards continuous improvement of products and services becoming the new business as usual.
Alpha vibes
A documented and required process exists to protect us from our own worst instincts.
In Alpha, the instinct we most need to guard against is falling in love with the first idea. Alpha for me is all about multivariate concepting.
It has always been tremendously easy to fall in love with the first solution we land on. Now with AI Assistance, it becomes insanely easy to bring that first idea into reality and for the team to only ever seriously consider a single approach to solving the problem for customers.
This introduces several layers of risk to the team.
The first risk is that you could have come up with a solution that was substantially better. I’ve worked with teams who are proud of the KPIs they are achieving and that they continue to improve. But what if they could have been 10x better. We’ll never know, because we didn’t ever really test a range of alternative solutions.
The second risk is related. By exploring and testing only one idea (or small variations of one main idea), you reduce your surface area for learning.
Build just one pretty good solution and suddenly the full scope of data and insights you can now achieve, both from research as well as from behavioural analytics, are constrained by the edges of that one instantiation of the solution. Build more possible solutions, learn more.
Building different options is now so cheap and so fast, and processing data from usage and customer feedback is also now very low cost. It seems foolish, to me, to unnecessarily limit the surface area of your learning. And this should also allow you to consider building and learning from wildly different potential solutions - not just variations on a theme.
So that, to me, is the alpha vibe we want to keep perpetually in our teams. Compelling ourselves to keep exploring substantially different ways to solve problems for our users and learning from them. It’s not something we tend to love to do naturally - its hard and distracting and feels like it takes us away from the finish line - but it makes our teams smarter, and the experiences we create better. So it stays in the process and we must hold on to the discipline.
The thing about Beta
The annoying truth is that there is no real finish line. Nothing is ever truly done, because even if the backlog ever gets completed (we wish), the world changes - technology, policy, society - and our products and services must adapt accordingly.
The idea of a perpetual Beta recognises that fact - that we are never really done, but rather in a Business As Usual (BAU) state of continuously assessing what needs to be improved on this part of the services and whether that is more or less pressing than in other parts of the organisation such that we can justify assigning a team to do the work.
The idea of a perpetual Beta also focuses us on getting our ideas into the hands of real users sooner rather than later and learning from that real life usage (or avoidance of use) giving us the highest quality data points available. And to continue to focus more on that real life usage and much less on prototypes and concept or usability testing of prototypes.
Throughout the existence of the Service Standard in both the UK and Australian public service we tended to think about Beta as a race to Live. Or, Done. The product, design and engineering decisions are typically driven by the need to get this work finished (including opening it up progressively to real users for learning), so we can wrap it up and move onto the next project.
The vast majority of researchers’ (and designers’) work in the traditional Beta phase involves working on prototypes that are tested by people who could theoretically be users. Not actual users. Because for the majority of the Beta phase, no one is actually using the new service ‘in real life’.
And once the service turns ‘Live’, chances are there is little or no team around to learn from real usage either. So the team’s ability to learn from actual customer usage and to apply it has been traditionally very limited.
This seems a shame, given the highest quality insight comes from actual customers using the service. There is a risk that accelerating development further reduces the time spent with real customers. I’m hoping for the opposite. I hope that it encourages the team to get things into customers hands sooner, and in turn, focuses the research effort on understanding what is working and not and why with the experiences (plural because of the Alpha vibe), building a shared context of understanding our customers and their needs and experiences, and prioritising improvements that better meet user needs.
Will people actually use it? How will they use it? Where will they make mistakes? What will they interpret differently to what we expected? Where do they need extra support? All of these questions are much more easily and reliably answered in the context of real use.
Too often once a service is ‘launched’, the team moves into a mode of relying much more heavily on analytics for feedback but little is done to ensure that the interpretation of those metrics is reliable, or whether perhaps we’re seeing what we want to see, or reflections of our own assumptions and biases in the data.
Active, multi-modal research focusing on these users will help sharpen insights from all the data points that become available once the service is in the hands of real users.
Continuous discovery
And, as discussed, Discovery doesn’t go away, but rather than being a phase at the outset of a project, it becomes a mindset and a method that imbues the research the team does continuously, and then also deeply whenever it is required.
You can imagine a backlog of research questions that the team continues to seek to answer, integrated into different research methodologies alongside analytics, with confidence in the answers changing over time. And also bigger questions that emerge that the team feel they need to deeply understand to help them feel confident to make decisions.
How does this scale?
The downside to not having projects that ‘end’ when they go live (which they shouldn’t of course, but let’s be real) is that there is no clear point where you can pull the team off and move them onto another priority. For better thoughts than mine on Organisational Design and Operating Models that might support this kind of approach I’d refer you to Kate Tarling, author of The Service Organisation.
The work of research in the AI-assisted team
Despite some otherwise radical changes, I believe the research work to be done in the team remains surprisingly consistent with what good practice looked like a decade ago.
Research is still a team sport
First and foremost, research as a team sport is more important than ever.
In a time when people can make a decision and build it so rapidly, everyone’s gut feel needs to be finely tuned to their users’ needs. And there is no better way to do this than through exposure hours. Everyone in the team watching real people use the actual service, at least two hours every six weeks. It is arguably more vital now than ever before. And it’s not something we can or should outsource to AI.
In staying close to the data, we build the context and judgement that informs our interpretation of the data the service generates, and the insights and recommendations that AI generates from its analysis.
You can choose how and when to get your hands dirty with the data - either in facilitating the research or getting actively and manually involved in the analysis and manipulating the qualitative data. Do not allow the team to get arms length from the research data by outsourcing it wholesale to AI. That is exactly how judgement is lost and the working relationship with AI becomes unreliable.
The team should be doing a continuous mix of discovery and evaluative research (again, something we’ve advocated for and done for many years). Some of this will be human facilitated, increasingly, and as the tools continue to improve, more of it will be AI facilitated. Research data and behavioural analytics data should be informing each other continuously.
Creating a shared understanding of our users
And the other big and essential job that needs doing is the collation, synthesis and communication of the shared context of the user and user needs.
One of the biggest risks for teams who lean heavily on AI assisted research is the tendency towards rapid but fragmented research studies. Every individual doing just the work they need to answer their individual problem but often without consideration of the bigger picture.
Doing the work to describe the shared context of our team’s understanding of our users/customers is critical so that we don’t fall into the trap of one person holding a trunk and thinking its a snake, another holding a leg and thinking its a tree, and no one realising that we are dealing with an elephant.
The mitigation to this risk is the curation and communication of a shared understanding of our customers or users amongst the team. Constantine Papas calls this work creation of The Frame.
The frame is the organization’s accumulated, actively maintained model of its users.
Not what it has studied. What it currently believes, based on the best available evidence, about who its users are, what they’re trying to do, what motivates them, what creates friction for them, how they make decisions, and where the product fits or doesn’t fit into their lives.
The creation of this shared understanding, The Frame, requires work and therefore it requires someone to own, be accountable for, and be resourced to invest in creating, updating, and communicating this understanding. It makes sense to me that this person is a researcher. The identification of this as significant and important work that requires investment is the important part.
Papas says:
Real ownership requires the frame to be in someone’s actual job description with protected time attached. Not “part of their responsibilities alongside everything else.” Protected time. The kind that doesn’t get cannibalized when a PM needs something by Thursday, which is always.
It also requires the frame steward to have the organizational permission to call for discovery work without filing a request and waiting for prioritization approval.
Ongoing opportunities for continous discovery would assist the ‘frame steward’ be proactive on getting priority discovery work actioned.
Fractional Leadership + Multi-disciplined Team Members
The public sector in the UK and to a certain degree in Australia, has benefited enormously from investment in training and development of user research capability and embedding user research talent in many teams. I suspect, however, we will see fewer and fewer teams where each discipline, including research, is represented by an individual human. And probably rightfully so.
Our AI Assisted teams of the future will likely be smaller teams. And within these team we’ll see a different kind of collaboration. Collaboration that has been supercharged thanks to AI assistance. Teams where everyone could do some design, everyone could produce some code, and everyone could take on some research activities.
To a point, of course, because each of these crafts also requires judgement that comes from experience and expertise. And without that judgement, we can all do a lot of AI-assisted dumb stuff.
All of this sharing of responsibilities may feel uncomfortable right now, but like letting go of front-loaded discovery, accepting that research is a job to be done in a team and maybe done by a specialist or distributed across the team is the reality of what behaving natively in this new technological landscape will look like.
What we will definitely need is fractional specialist leadership.
These are people who can work across many teams and bring their research (or their design or their technical) expertise to help teams be confident they are executing that discipline effectively in that team’s context. That the team’s judgement about their customers is strong, and that the team has considered the risks and opportunities of different ways they could approach the work.
One of the risks of working with AI assistance is that people can always show up to the meeting with something that looks like research. But AI’s capability still outweighs its reliability. Human judgement is required to ensure the interpretation can be relied on, that it hasn’t hallucinated data or verbatims. That it hasn’t drifted from reality.
But we don’t just acquire judgement by virtue of being human. That judgement is born from observation and experience. Whether you’re a software engineer or a UX researcher, you’ve watched things fail over and over to the point where you have strong intuition for whether something has a high likelihood of success or not. You’ve hopefully been close enough to the data to know whether the claims being made ring true or not.
Judgement comes from time spent in your field of expertise — but it also comes from time spent in your domain. The problems you’re trying to solve for that customer. In order to work safely with AI, you need teams who have developed judgement about the subject domain, who can collaborate with the AI to instruct it well and course-correct it when required. And who will be able to call it out when, inevitably, something unreliable surfaces.
On synthetic data and data processing
I think Ben does a pretty cracking job of describing the pace and opportunities that AI brings to user research in AI Assisted teams.
His view that synthetic data will be enormously useful to help ‘test’ research scripts and surveys so we can refine them and be sure they provide the insights we require is one I hear echoed by many of experienced researchers who are embracing AI workflows.
We do need to take the warnings about the shortcomings of synthetic data seriously though. There will be plenty who will no doubt look at synthetic data, shrug and say ‘good enough’ We ‘just looking for a signal, the data doesn’t have to be perfect’ - every researcher will have heard that plenty of times even before AI was everpresent.
Evidence to date on the actual benefits of using synthetic data in research to inform product and design decisions is not encouraging. I’d go so far as to suggest it looks quite discouraging.
Our friend, Mr Papas, who is both prescient and remarkably prolific when it comes to writing about UX Research in this AI assisted era, shared a thorough review of the largest systematic review of synthetic participants ever conducted and the finding was unambiguous: synthetic users don’t work in the way we’d like.
But, precisely because of how believable they can be, the use of synthetic users can instead have a negative impact on a team’s ability to empathise and understand their users.
The believability of LLM participants is excellent. Experts in one study couldn’t distinguish synthetic personas from human-generated ones on surface inspection. The text is clean, structured, elaborate.
And shallow, and stereotypical, and perspectiveless, and occasionally fabricated. The review’s assessment: believability “may actually be more detrimental than beneficial by lending false credibility to misleading conclusions.”....
… Trust in LLM-generated personas can impede the ability to empathize and understand others. It reduces the depth of mental engagement. You don’t lean into understanding someone when you think the machine already did it for you.
I think this is the biggest challenge for AI-assisted teams: acquiring, maintaining and applying judgement that is based on the reality of our users and customers.
Using AI to remove the friction of recruiting, observing and analysing the output of research interactions is extremely tempting, but we flirt with a risky vicious cycle emerging. One where the synthetic data produces insights that look and feel like they might be true. And so you feel less inclined to interact with real users, who are slower, more difficult and messier. And so your judgement is impaired. And so the synthetic data looks ever more real.
AI assisted research has its place. There are huge opportunities, especially when it comes to pulling together large data sets and customer feedback from diverse sources and making sense of it. But allowing the team to automate and thereby outsource this activity wholesale is sufficiently tempting to humans that I believe we need to put safeguards in place. Which makes it something we should build into the rules of how we operate and our process.
Which, for me, just means keeping exposure hours for the whole team in the program forevermore.
Beware vibes without people
Ben closes his essay saying “The task is to use speed to get closer to reality, not further from it.” Nailed it.
Building with code has never been cheaper. But creating a shared understanding of who your users are, what they actually need, where and why the service is failing them develops only through sustained, repeated contact with real people in real contexts.
For teams to become and remain user centric, someone has to do the work to create the big picture, keep it updated, and make sure everyone in the team remembers the elephant when we’re moving fast and grabbing just what we need. And we need to stay close enough to real people and the data they produce so we know something real when we see it. And we won’t be fooled by something convincing and convenient but ultimately untrue.
It’s not the time to be arguing to keep things the way they’ve been for the last decade. Some of what was rightfully precious then won’t continue to earns its keep in our AI Assisted future.
But some things remain critical. Exposure hours, multivariate exploration, the shared understanding (the frame), the discipline of staying genuinely close to real people.
This is how we move fast but stay honest.


