mcp

MCP went the HTTP way: what changed in the 2026 protocol

MCP's 2026-07-28 revision removes sessions and makes each request self-contained. Here is what changed and how to use it.

Kirti Rathore··8 min read

Your MCP server works in development. Then you put three replicas behind a load balancer.

The first tool call reaches replica A. The next reaches replica C, which knows nothing about the session created by A. Now you need sticky routing or a shared session store.

The 2026-07-28 revision of the Model Context Protocol removes that problem from the protocol itself.

There is no session establishment handshake and no Mcp-Session-Id. Each request is self-contained.

One terminology note before we continue: "MCP v1" and "MCP v2" are convenient labels, but the protocol is officially versioned by date. This article compares the legacy era through 2025-11-25 with the modern era beginning at 2026-07-28. SDK package versions are a separate thing.

The old model: initialize, then talk

Legacy MCP began with an initialize request. The client and server negotiated the protocol version and capabilities, then the client sent notifications/initialized. Over Streamable HTTP, a server could also issue an Mcp-Session-Id that the client returned on later requests.

Diagram of legacy MCP architecture, session flow, JSON-RPC messages, primitives, and security boundaries
Legacy MCP through 2025-11-25: initialization and optional protocol sessions provide context across calls.

That model was useful for interactive, bidirectional clients. It also tied later messages to context created earlier in the connection.

At small scale, the coupling is easy to miss. At production scale when you must scale out your servers, it creates questions that infrastructure must answer:

  • Which replica owns this session?
  • What happens when that replica is lost or restarts?
  • Is session data shared, retained, or cleaned up?
  • How long should the server retain it?

The new stateless model

Modern MCP makes a request self-describing. The protocol version and client capabilities travel in request _meta; clients should also include their identity.

POST /mcp HTTP/1.1
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search
Authorization: Bearer ...
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "search",
    "arguments": { "query": "checkout timeout" },
    "_meta": {
      "io.modelcontextprotocol/protocolVersion": "2026-07-28",
      "io.modelcontextprotocol/clientInfo": {
        "name": "incident-agent",
        "version": "1.0.0"
      },
      "io.modelcontextprotocol/clientCapabilities": {}
    }
  }
}

A client that wants to inspect the server first can call server/discover insted of tools/call, but discovery is optional. It is no longer a mandatory state-creating handshake.

Diagram of modern stateless MCP architecture, request flow, JSON-RPC messages, primitives, and security boundaries
Modern MCP from 2026-07-28: self-describing requests, optional discovery, explicit state, and no protocol session.

Here is the practical protocol diff:

  • Startup: legacy MCP uses initialize; modern MCP has no mandatory handshake.
  • Protocol session: the optional Mcp-Session-Id is gone.
  • Discovery: initialization-based discovery becomes the optional server/discover call.
  • Multi Round-Trip Requests: New feature added in MCPv2. We'll discuss this later.
  • Caching: now has first-class support. List responses include explicit ttlMs and cacheScope hints.

Flexible enough to allow stateful workflows

Suppose a debugging tool needs to preserve an investigation across calls. The server can create an explicit handle:

{
  "investigation_id": "inv_447",
  "status": "collecting_evidence"
}

The agent passes that handle into the next tool call:

{
  "name": "inspect_deployment",
  "arguments": {
    "investigation_id": "inv_447",
    "deployment_id": "deploy_92"
  }
}

The investigation may live in a database. The important change is that its identity and lifecycle are explicit. It can survive a client restart, move between agents, expire under its own policy, and be authorized independently of the HTTP connection.

This is the same separation that makes the web scalable. HTTP is stateless, but web applications still have accounts, carts, jobs, and databases. The state belongs to the application, not to an implicit transport conversation.

Multi Round-Trip Requests (MRTR)

MRTR was added to solve one obvious problem: what if a tool starts work and then needs user approval or another missing input?

Here is how this flow can now be handled:

client -> server: tools/call(restart_service)
client <- server: resultType=input_required

client asks the user for approval

client -> any server replica: retry tools/call with inputResponses
client <- server: resultType=complete

The server does not keep the original worker suspended while it waits. It returns what it needs, and the client retries the operation with the answer.

This improves resilience, but it also can introduce side-effects.

A state-modifying tool such as restart_service or create_incident should accept an idempotency key so a network retry does not perform the action twice.

Why switch to the new MCP protocol

Scaling becomes simpler. Requests can use round-robin routing without pinning a client to one replica or sharing protocol-session memory.

Gateways gain useful context. Mcp-Method and Mcp-Name let a gateway distinguish tools/list from tools/call, or read_logs from restart_service, without parsing an arbitrary JSON body. That enables per-tool authorization, rate limits, metering, and audit rules.

Discovery becomes cacheable. Tool, prompt, and resource listings include freshness and cache-scope hints. Deterministic ordering also keeps prompt caches stable when the catalog has not changed.

For teams building AI tools for developers, this is the larger design lesson: make state, authority, cost, and retries visible to the caller.

Takeaways

The 2026-07-28 revision is a breaking protocol change, but its architecture is easier to reason about:

MCP request state     -> travels with the request
application state     -> explicit handle or durable store
authorization state   -> identity and policy system
long-running work     -> task or extension

It does less on your behalf. That is the point.

The protocol no longer pretends that a connection is the right home for a browser, a workflow, or a custom protocol.

That makes MCP a better foundation for the kind of agent systems we are building at Modulo AI: disposable compute around explicit, durable work.

The official references are the 2026-07-28 release announcement, the modern Streamable HTTP specification, and the accepted proposals for stateless MCP and session removal.