Skip to content
← Insights

MCP just grew up. Here is everything that changed and why it matters.

What MCP is, how it relates to the APIs you already understand, and a full breakdown of the most significant update to the protocol since it launched.

By Tochii Achebe11 min read

Every so often, a piece of infrastructure quietly reaches the point where it stops being an experiment and starts being the plumbing everyone builds on.

MCP just crossed that line.

On July 28th, Anthropic and the wider MCP community shipped the most significant update to the Model Context Protocol since it launched in November 2024. Across the four official SDKs, MCP is now seeing close to half a billion downloads a month. The TypeScript and Python SDKs have each individually crossed one billion total downloads. This is no longer a promising standard. It is production infrastructure that a huge share of the AI industry now depends on.

This week, we cover what changed, why it changed, and what it means for anyone building with AI agents.

But first, for anyone joining partway through this series, let us make sure the foundation is solid.


What MCP actually is

MCP stands for Model Context Protocol. Anthropic created it and open sourced it in November 2024 as a way to solve a very specific problem: AI models are only as useful as what they can access.

A brilliant AI model that cannot read your documents, query your database, or check your calendar is a brilliant model with nothing to work with. Before MCP, every AI provider and every company building an AI product had to write custom integration code for every single tool they wanted their AI to use. Want Claude to search your internal wiki? Custom code. Want it to also check Slack? More custom code. Want that to work with GPT as well? Write it all over again.

MCP replaced that mess with a single, open standard. One protocol that any AI model can speak, and any tool, database, or service can speak back to. Build the connection once, and it works with any AI system that supports MCP.

Contextualising MCP with APIs

We covered APIs in detail earlier in this series, so let us connect the dots directly.

An API is a set of rules that lets one piece of software ask another piece of software to do something and get a response back. When your banking app shows your balance, it is calling an API behind the scenes. APIs are how software has talked to other software for decades.

MCP is not a replacement for APIs. It sits on top of them.

Think of it this way. An API is a conversation between two specific pieces of software that were built to understand each other. Someone had to write code on both sides, in advance, agreeing on exactly what requests look like and what responses mean. That works well when you know in advance which two systems are talking.

MCP solves a different problem. It gives an AI agent a standard way to discover what tools are available, understand what each one does, and call them correctly, without a developer having to hand-write that integration in advance. Under the hood, an MCP server is very often just a wrapper around one or more existing APIs. MCP is the universal adapter. The API is still the underlying connection being made.

If HTTP is the language that lets any browser talk to any website, MCP is quickly becoming the language that lets any AI agent talk to any tool.


Why this update matters: from local tool to internet-scale protocol

To understand why the July 28th release is such a big deal, you need to understand the problem it solves.

MCP was originally designed with a simple mental model in mind: an AI assistant running on your laptop, talking to a small tool running as a local process, in a single continuous session. That design worked beautifully for that use case. It included a stateful handshake, where the client and server introduce themselves, agree on a shared session, and then keep that session alive for the whole conversation, tracked via a session ID.

Then something predictable happened. Developers loved MCP so much that they started deploying it everywhere: in the cloud, behind enterprise load balancers, serving thousands of simultaneous users, not just one person on one laptop.

And that stateful, session-based design, which worked perfectly for a single local connection, became a serious bottleneck at scale. If a session is tied to one specific server instance, you cannot simply spread traffic across a pool of interchangeable servers behind a standard load balancer. The server has to somehow remember which instance is holding which conversation. That requires expensive workarounds: sticky sessions, shared state stores, and infrastructure complexity that should not be necessary for what is fundamentally a request and response system.

The 2026-07-28 specification exists to fix exactly this. As lead maintainer David Soria Parra put it, this is probably the most substantial change to the specification since authorization was added, and it comes with breaking changes.

Update one: MCP is now stateless

This is the headline change, and it is a genuine architectural shift.

The old initialize and initialized handshake, along with the Mcp-Session-Id header that tracked a conversation across multiple requests, has been formally retired. Every request now carries everything it needs to be understood on its own: the protocol version, the client’s identity, and its capabilities, all self-contained in the message itself.

The practical effect is significant. Any request can now land on any server instance behind an ordinary round-robin load balancer, with no shared memory and no sticky sessions required. If your server does genuinely need to remember something between calls, the new pattern is to mint an explicit handle from a tool and have the model pass that handle back as a normal argument on the next call. The state is visible and travels with the data, rather than being hidden inside the plumbing of the transport layer.

For anyone running MCP servers at real scale, this single change removes an entire category of infrastructure headache.

Update two: Multi Round-Trip Requests replace the held-open connection

Sometimes a tool needs something from the user mid-task. A confirmation before it deletes something. A missing piece of information it was not given. Previously, this required the server to keep a connection open and reach back out to the client, which is exactly the kind of always-on, bidirectional stream that does not play nicely with a stateless, load-balanced world.

Multi Round-Trip Requests, or MRTR, solves this differently. Instead of holding the line open, the server simply responds to the original call with a special result saying, in effect, I need more information before I can finish. The client gathers what is needed and retries the same call with the answers attached.

Supabase’s Head of Product gave a clean, concrete example of why this matters: their MCP server runs statelessly, so it previously could not ask a user to confirm the cost of a new project, or confirm a query that would delete data, before acting. MRTR makes that kind of safety confirmation possible without breaking the stateless model.

Update three: header-based routing

Every request now carries its method and tool name in dedicated HTTP headers, Mcp-Method and Mcp-Name, rather than having that information buried inside the JSON body.

This sounds like a small technical detail. It is not. It means a gateway, a rate limiter, or a firewall sitting in front of your MCP servers can now make routing, security, and metering decisions just by reading standard HTTP headers, without needing to parse and inspect every request body. For enterprises running MCP behind serious infrastructure, this is exactly the kind of detail that determines whether a protocol is genuinely production-ready.

Update four: list results can now be cached

When an AI agent connects to an MCP server, one of the first things it typically does is ask what tools are available. Previously, that catalogue had to be fetched fresh, every time, from every connection.

Now, responses to tools/list, prompts/list, resources/list, and resources/read carry explicit cache hints and a time-to-live value. Clients can cache a server’s tool catalogue intelligently instead of re-fetching it constantly, which reduces load on the server and keeps things faster and cheaper for everyone involved.

Update five: authorisation gets meaningfully harder to get wrong

The MCP team has been clear that authorisation is where most implementers spend the bulk of their integration time, and this release invests heavily in tightening it up.

Issuer validation. Authorisation servers must now return an issuer parameter that clients validate before redeeming a code, closing a class of vulnerability where a malicious server could impersonate a legitimate one.

Fixing a long-standing CLI headache. Desktop and command-line clients can now correctly register as native applications during setup, which fixes a frustrating, common error where local tools were rejected for using a localhost redirect.

Credentials are now bound to their issuer. A set of credentials minted by one authorisation server can no longer be reused against a different one.

A formal move away from Dynamic Client Registration. In favour of Client ID Metadata Documents, a more standards-aligned approach to how clients identify themselves. The old method still works for now but is on a deprecation path.

Update six: Tasks becomes an official extension

Not every piece of agentic work finishes in a second. Some tasks are genuinely long-running: a deep research job, a large document analysis, a multi-step workflow that takes minutes rather than milliseconds.

Tasks, which shipped as an experimental feature in the previous release, has now graduated into a formal extension with a redesigned lifecycle built specifically for a stateless world. A server can respond to a tool call with a task handle immediately, and the client then checks on progress, requests updates, or cancels the task using that handle, rather than everyone waiting on an open connection for however long the work takes.

AWS contributed this extension and now supports it natively in Amazon Bedrock AgentCore, specifically framing it as the foundation for reliable, long-running enterprise agents.

Update seven: MCP Apps, interactive interfaces inside the protocol

This is one of the more visually interesting additions. MCP Apps lets a server ship an actual interactive HTML interface that the AI client can render safely inside a sandboxed frame.

Tools declare what their interface looks like in advance, so the client application can pre-fetch it, cache it, and security-review it before anything is ever shown to a user or allowed to run. Crucially, any action a person takes inside that rendered interface still travels back through the exact same audit and consent path as a normal, direct tool call. You are not opening a security side-door by adding a richer visual experience.

Figma’s VP of Engineering pointed directly at this as a meaningful unlock, since it lets AI-generated output land directly inside their design canvas as something a team can explore and refine together, rather than as a static, disconnected output.

What was deprecated, and what that actually means for you

Several older features have been formally marked as deprecated: the original Roots, Sampling, and Logging capabilities, the legacy Dynamic Client Registration flow, and the older HTTP plus Server-Sent-Events transport that MCP originally launched with.

None of this breaks overnight. The project has committed to a formal deprecation policy with a minimum twelve month window before anything deprecated is actually removed. If you have an existing MCP integration today, it will keep working. But the clear signal from the maintainers is that new implementations should be built against the new patterns from this point forward, not the old ones.


Why this matters beyond the changelog

It would be easy to read all of this as a list of technical housekeeping items. That would miss the actual significance of what just happened.

Every major cloud and AI infrastructure provider showed up to comment on this release. AWS is running Tasks natively in Bedrock AgentCore. Google Cloud is calling it a massive leap forward for enterprise scalability. Cloudflare shipped day-zero support in their Agents SDK, running MCP servers directly inside Cloudflare Workers. Microsoft is building it into Foundry as the standard their platform scales integrations on. Honeycomb reports that agents already account for nearly 20 percent of their monthly interactive queries.

This is what a piece of infrastructure looks like right before it becomes invisible in the way that HTTP is invisible. Nobody thinks about the fact that HTTP exists when they open a website. In two or three years, MCP is on a clear trajectory to become exactly that kind of assumed, load-bearing layer beneath every serious AI product.

What this means for you, practically

If you are building or specifying MCP servers, the practical guidance is straightforward. Support the extensions gracefully, even the ones you have not adopted yet, so your client does not break when it encounters a server using MCP Apps or Tasks. If you are starting a new integration today, build it against the new stateless core from the outset rather than the older session-based pattern, since that is clearly where the whole ecosystem is heading.

If you are a product leader or founder evaluating whether to invest engineering time in MCP integrations for your company, this release is a strong signal in favour of doing so. The protocol has just addressed the single biggest technical objection that enterprise teams raised about running it at scale. The infrastructure providers are not experimenting anymore. They are building it in as a default.


Reflection for the week

We have tracked MCP across several articles in this series now: what it is, how it compares to APIs and SDKs, and how to build one from scratch. This update is the moment the protocol stopped behaving like a promising new idea and started behaving like real infrastructure.

Real infrastructure is not exciting in the way a flashy new AI model announcement is exciting. It is exciting in a quieter, more durable way. It is the plumbing that everything else gets to take for granted.

The people who understand this layer now, while it is still being actively discussed and shaped, are the ones who will be building the products that simply work when everyone else is still fighting session management bugs.

The question worth sitting with this week:

If you are building or planning to build with AI agents, does your current architecture assume state that this update just made unnecessary? And what would change if it did not?


Help shape future Learn with Tochii articles
I’m putting together future guides on practical ways professionals can use AI at work, in school, and in everyday life.

What’s one AI question you’ve always wanted answered?

Your responses will help decide what I write about next.

Take the 2-minute survey: https://bit.ly/LWTsurvey


Want to understand the infrastructure behind AI agents at this level of depth?

MCP, agent architecture, and how to build production-grade AI systems are core parts of what we teach in Amakora Group’s AI programmes.

Explore our courses: amakoragroup.com/programs

Apply for the fellowship: amakoragroup.com/apply

Until next week,

Tochii

Founder, Learn with Tochii | Amakora Group

Inspire · Educate · Empower

📧 contact@tochukwuachebe.com

🌐 https://www.tochukwuachebe.com

First published on Learn with Tochii on Substack. Subscribe to get new essays by email.

More insights

All articles

We also train teams to do this well.

See the programmes