top of page

When the Backlog Writes Itself: What MCP and Agentic AI Actually Mean for Delivery

Hi everyone, hope you're doing well?


I wanted to share something I've been thinking about for quite a while now. It started when I began building an Azure DevOps MCP server that takes the usual mix of discovery outputs, transcripts, workshop notes, requirements documents, process maps, and turns them into a structured backlog.


The technology itself is interesting, but what really stuck with me was what building it taught me about agentic AI. Not the hype. Not the LinkedIn version. The practical reality of where it genuinely helps, and where it absolutely shouldn't be left to its own devices.


So grab a brew and let's dive in.


The bottleneck was never the thinking


For years, I assumed the hardest part of turning discovery into delivery was the analysis.

You know the scenario:


* A workshop transcript

* A process diagram

* Half-finished requirements

* A few stakeholder notes

* Contradictory comments from three different teams


Making sense of that takes skill.

But if I'm honest, that wasn't the part that frustrated me.


The real pain came afterwards.


Creating the Epic.

Then the Feature.

Then the User Story.

Then the acceptance criteria.

Then doing it again.

And again.


Especially when you need to rework half of it when the scope changes next week.


That's not analysis. It's administration.

Important administration, yes. But administration nonetheless.


Note: I mean, I say pain, but part of me loves admin. But even this can get too repetitive for my ADHD dopamine seeking brain.


For a long time we've accepted that as the cost of maintaining a governed backlog. It was simply part of the job.


What agentic AI changes, when used properly, is that this cost becomes optional. You can hand the repetitive work to an agent and keep your own time and energy for the bits that genuinely need human judgement.


That's the real opportunity.


Why MCP finally got my attention


I'll be honest, I largely ignored Model Context Protocol when it first appeared.


Another acronym.

Another standard.

Another thing we're all apparently supposed to care about overnight.


But MCP has turned out to be one of the more important developments in the AI space. The easiest description I've heard is that it's the USB-C of AI. Before MCP, if you wanted an AI agent to interact with a business system, you built a bespoke integration. Every connection was different. MCP standardises that interaction. You expose tools through an MCP server and any MCP-aware client can discover and use them.


In my case, those tools include things like:


* Create Epic

* Create Feature

* Create User Story

* Generate Acceptance Criteria

* Process Discovery Transcript

* Create Tasks

* Extract Risks and Assumptions


What surprised me was how naturally this fit into the Microsoft ecosystem I already spend most of my time in.


The setup is actually pretty simple:


* Copilot studio provides the user-facing experience

* Power Automate handles document movement behind the scenes

* Entra ID manages authentication and access

* Azure Container Apps hosts the server


And the hosting cost still feels ridiculous.

Around £6 a month.


A user uploads a transcript.

The agent analyses it.

It creates an Epic → Feature → Story → Task hierarchy.

It drafts acceptance criteria.

It applies prioritisation.

It builds the backlog.


Then it simply reports back:

Backlog created: 4 epics, 12 features, 38 stories, 95 tasks.


The activity somebody was putting off for an entire afternoon happens while they're making a coffee.


That's no longer a chatbot answering questions.

That's an agent completing a piece of work.


The bit people keep missing: governance


This is where I think a lot of AI conversations go off track. If my server simply generated work items and pushed them into Azure DevOps, I'd have created a very efficient way to generate a very large amount of rubbish.


A backlog is only valuable if people trust it.

If someone has to manually audit every AI-generated story before using it, you've not solved the problem. You've just moved it somewhere else. Or in the words of one Prime Minister “kicked the can down the road”.


That's why the most important part of what I built isn't the generation layer.


It's the enrichment layer.


Every story is assessed before a human reviews it.


The system adds:


* Confidence scoring

* Quality measurements

* Clarity indicators

* Completeness checks

* Testability assessments

* Consistency checks

* T-shirt sizing estimates

* Dependency analysis

* Definition of Done suggestions


Most importantly, it identifies what's missing.

If a workshop never discussed reporting requirements, it doesn't hallucinate them.

It flags the gap.


That's a very deliberate design choice, the enrichment isn't another AI opinion, it's determijistic. If you run it twice against the same input and you get the same result. The goal isn't to replace the analyst's judgement. It is to focus their attention.


Think of it as a backlog health dashboard.


A quick way to identify:

* Ambiguous stories

* High-risk items

* Missing information

* Weak requirements

* Areas requiring further discovery


The AI handles the repetitive labour and the human remains accountable for the decisions.

That's the line I'm increasingly convinced we need to draw, the tool also reflects how discovery actually works in consulting.


Transcripts aren't truth, they're evidence and for want of a much better phrase (sorry Mr/Mrs Client), straight from the horses mouth.

They're one input among many. Where information is incomplete, the system creates placeholder discovery stories, captures fit-gap areas, records provenance, and automatically extracts RRAID items:


* Risks

* Requirements

* Assumptions

* Issues

* Dependencies


It's not pretending to know everything, it's just helping surface what still needs a conversation.


What this means for Business Analysts


YOU ARE NOT OUT OF A JOB


Every time agentic AI comes up, I can almost feel the same question hanging in the room.

"Is this coming for my job?"

Having spent quite a lot of time building one of these systems, my honest answer is:

No.


But it is coming for some of the worst parts of your job. A mentor once said to me, “If you don’t start using AI in your everyday workflow you are going to get left behind. AI will either happen with you or to you. You want the former!”


And that's probably a good thing.

The value of a great Business Analyst was never their ability to type user stories into Azure DevOps.


Their value comes from:

* Running effective workshops

* Understanding stakeholder needs

* Spotting hidden assumptions

* Challenging unclear thinking

* Managing competing priorities

* Defining what success looks like

None of that disappears.

If anything, it becomes more important.


As the administrative burden reduces, the essential human skills become more valuable and the more important basis for long term mutual beneficial client <> consultant relationships.


What changes is the shape of the role.


You move from being:


Author of every work item

to

Editor and governor of an AI-generated backlog


Your leverage increases.


The billable time you spend with the client is with the client, building that relationship, not having to take notes. Active listening, empathising and understanding their pain points.


One analyst can oversee significantly more discovery and delivery activity than before.


The premium shifts towards skills machines struggle with:


* Judgement

* Context

* Experience

* Facilitation

* Communication

* Stakeholder management


Especially stakeholder management.


Let's be honest, no AI wants to have that awkward conversation everyone has been avoiding for three months.


Where I've landed


Building this has changed the way I talk about AI.

I no longer ask:


"Can AI do the whole job?"


Because that's usually the wrong question.

It's the sort of thinking that produces impressive demonstrations and terrifying production environments.


The better question is:


“What work can I safely hand over, and what judgement should remain human?”


Get that balance right and agentic AI becomes a genuine force multiplier, get it wrong and you've simply automated the creation of technical debt.


MCP is making this practical today using tools many of us already have access to.


But MCP isn't really the story.

It's just the plumbing.


The real challenge is deciding which parts of our craft are worth automating and which parts are too important to give away. Which parts require a human.


That's the conversation I think we should be having.


If you're experimenting with Copilot Studio, MCP servers, or agentic delivery workflows, I'd genuinely love to hear how you're approaching that balance.


Drop me a message or leave a comment.


Or contact me - jon@jondoesflow.com


Jon

Comments


Subscribe Form

©2019 by Jon Does Flow. Proudly created with Wix.com

bottom of page