A Foundry hosted session is not a conversation
This week's Microsoft Foundry primaries split three lifetimes that teams keep collapsing: WebSocket voice, crash-resilient background work, and user identity versus the VM sandbox. The session ID is the filesystem, not the chat.
Jeff
Windows & Microsoft Ecosystem

Microsoft's Agent Framework blog published a sentence on 22 September 2026 that should sit on the first page of every hosted-agent runbook: a Foundry hosted session is not a conversation. A conversation is message and tool-call history. A hosted session is sandbox compute and persisted files. Two days later, Tina Schuchman's Azure Blog put voice agents and long-running resilience on the same Foundry foundation as the GPT-6 family. If you treat those as one "session" knob, you will reconnect a voice socket to the wrong filesystem, or assume a crash replayed a workflow it only reentered.
This post is a field guide from Microsoft primaries fetched on 27 September 2026. It is not a Foundry tenant we exercised this morning. The claims below are from the Azure Blog, Microsoft Learn, and the Agent Framework DevBlog — plus a grep of this site's content/blog tree, which has no prior deep-dive on invocations_ws, resilient_background, or the two isolation controls.
What actually moved this week
Schuchman's 24 September post is the umbrella. The Learn pages are the contract.
| Surface | Status in the primaries | Operator takeaway |
|---|---|---|
| GPT-6 Sol and GPT-6 Luna, Claude Opus 5.5 in Foundry | Available in Foundry this week (Azure Blog, 24 Sep) | Model choice is continuous; evaluate against your traces, do not freeze architecture on one model |
| Voice agents | Public preview as a Foundry Agent Service agent type (Azure Blog); hosted path is invocations_ws (Learn) |
Same container can also expose /responses and /invocations |
| Long-running resilience | Public preview (Azure Blog + Learn) | Background work and crash recovery are different opt-ins |
| Insights in Foundry | Public preview (Azure Blog) | Production traces → recurring issues, not a dashboard you already had |
| Rubric evaluator, traces-to-dataset, Agent Optimizer | "Generally available later this month" (Azure Blog, 24 Sep) | Do not write GA into a runbook yet |
| AI Gateway tier in APIM | Preview support coming in October (Azure Blog) | Hub-and-spoke model governance, not a replacement for today's APIM integration |
| Hosted-agent isolation APIs | 22 Sep Agent Framework DevBlog | User identity and agent_session_id are independent |
Xia Song's 22 September Copilot post says Claude Opus 5.5 and GPT-6 Sol are rolling out across Word, Excel, PowerPoint, Chat, Cowork, and Copilot Studio, with Work IQ still grounding responses in org files, meetings, chats, and business data. Availability varies by license, access, and region.
Naomi Moneypenny's 22 September Foundry models post is the routing table: start demanding work on GPT-6 Astra; Sol for general production agents; Luna for high-volume extraction, summarization, and routing. Global Standard short-context input list prices in that post: Astra $10, Sol $2, Luna $0.10 per 1M tokens. Standard is listed across 28 Global regions plus US and EU Data Zones; Provisioned Throughput for Astra and Sol; Priority Processing for Sol.
Voice is a protocol, not a toggle
Build a voice agent with hosted agents is the how-to that matches the Azure Blog's "voice becomes a native part of the agent foundation" claim. Hosted agents support real-time voice through invocations_ws. The container exposes GET /invocations_ws with a WebSocket upgrade. The public URL is:
wss://<account>.services.ai.azure.com/api/projects/<project>/agents/<agent>/endpoint/protocols/invocations_ws?api-version=v1&agent_session_id=<session-id>
Facts from that Learn page that change design, not marketing copy:
- The platform does not parse, transform, or buffer frames at the application layer. Text and binary frames are raw-byte relay. 20 ms PCM at 16 kHz mono is about 640 bytes. The proxy rejects frames over 1 MB with close code
1009. - Callers present a Microsoft Entra bearer token on
Authorizationduring the upgrade. APIM and Agent Service validate it. The container does not see that header. Do not accept anauthorizationquery parameter. - Individual WebSocket connections are capped at about 30 minutes. Platform drain sends close code
1001(going away). Clients must reconnect with the sameagent_session_id. The sandbox persists; the platform does not replay missed frames. - Session resolution inside the container:
FOUNDRY_AGENT_SESSION_ID, else the query parameter, else generate a UUID. - Declare the protocol on the agent version with
ProtocolVersionRecord(protocol=AgentEndpointProtocol.INVOCATIONS_WS, version="2.0.0"). You can declare it alone or withresponsesandinvocations. - Voice recommendation: at least 1 vCPU / 2 GiB; sandboxes go up to 2 vCPU / 4 GiB.
- Validated sample frameworks: Microsoft Voice Live, Pipecat, LiveKit Agents. Hosted agents do not provide a managed WebRTC media service, TURN, or SFU. Use
invocations_wsas authenticated signaling if you run WebRTC yourself. - Idle timeout for the sandbox is 2–60 minutes, 15-minute default on this page's limits table (hosted-agents concepts also documents 5–60 minutes, same default). Compute deprovisions; session state persists.
- Python handler surface:
@app.ws_handleronInvocationsAgentServerHost, same host as@app.invocation_handlerfor HTTP/invocations.
Telephony is a bridge, not a Foundry SIP stack. Learn says use a provider such as Azure Communication Services or Twilio to bridge PSTN onto invocations_ws.
If you already shipped GPT-transcribe / GPT-live-transcribe as a speech endpoint, that is a different product surface. This week's voice agent is your container on the hosted-agent protocol, with Entra on the upgrade and session identity on the query string.
Crash recovery is off until you opt in
Resilience for long-running hosted agents (preview) splits three capabilities that product language often mashes together:
| Capability | What it provides | What it does not provide |
|---|---|---|
| Background execution | Work continues after the original HTTP connection closes; clients poll or reconnect | Recovery after the process that owns the work stops |
| Resilient execution | Durable work identity, persisted input, lease-based recovery, handler reentry | Automatic preservation of every intermediate app state or side effect |
| Stream replay | Retained events from a cursor for reconnecting clients | A checkpoint of the agent's internal workflow |
Recovery reenters the handler from the beginning. It is not deterministic replay. It does not restore local variables. Your handler must use durable checkpoints or watermarks.
The how-to Recover long-running work after a crash is explicit that crash recovery is off by default:
from azure.ai.agentserver.responses import ResponsesAgentServerHost, ResponsesServerOptions
app = ResponsesAgentServerHost(
options=ResponsesServerOptions(resilient_background=True),
)
Recovery applies only to Responses that are stored and background: store=true and background=true. Without the opt-in, a crashed background response is marked failed with error.code="server_error" and the handler is not reinvoked. Foreground (background=false) responses are always marked failed on crash because the client connection is already gone.
On the Invocations / task path, declaring @task is not enough. Call set_resilient_tasks_enabled(True) before host startup. On reentry, branch on context.is_recovery (Responses) or ctx.entry_mode == "recovered" (tasks). Fence non-idempotent side effects by flushing a watermark before the email or charge, then clearing it after commit. If shutdown cannot finish, exit_for_recovery() leaves the record in progress instead of writing a terminal state.
Learn marks this preview: no SLA, not recommended for production workloads. That is the honest ship status.
Two isolation knobs, not one
Roger Barreto and Tao Chen's 22 September Agent Framework post is the missing control-plane diagram.
| Control | Represents | Choices |
|---|---|---|
| User isolation | Whose data may be used | Caller's Microsoft Entra identity, or a delegated identity from a trusted middle tier |
| Foundry hosted session isolation | Where code and files live | Session created on first request, or an agent_session_id the application supplies |
The delegated identity is independent of the session ID. Separate user conversations can share a sandbox or get separate sandboxes. In a pooled design, a middle tier maps users onto a bounded set of session IDs, sends the delegated identity on each request, and Foundry keeps response chains private by user while the sandbox filesystem remains shared. Application-owned files, database rows, and caches then partition on both agent_session_id and user identity.
.NET stores the delegated identity on AgentSession (CreateFoundryHostedAgentSessionAsync(..., userIdentity:)). Python currently forwards x-ms-user-identity on each call. Agent Framework does not create the remote Foundry hosted session; it stores the ID and sends it. Hosted Agents themselves are GA; the AgentServer SDKs and Agent Framework Foundry hosting packages for .NET and Python are still pre-release — the DevBlog says so in the same post.
Hosted agents concepts matches that split: session ID is $HOME and /files; conversation ID is message history. Idle timeout 5–60 minutes, default 15. Permanent session delete after 30 days of inactivity. Disk budget up to 20 GiB at 1 vCPU or larger, about 20% reserved for the platform. Scaling is per session, not per replica. CPU and memory on the agent version describe one sandbox; oversizing multiplies by concurrency.
How the three lifetimes compose
Put the three primaries on one operator table.
| If you need… | Bind | Do not assume |
|---|---|---|
| Microphone in, speech out | invocations_ws + agent_session_id on the upgrade |
The 30-minute socket is the session. Reconnect; sandbox stays; frames do not replay |
| A research job that outlives the HTTP client | Responses background=true and store=true |
Background mode recovers a crashed process. It does not, unless resilient_background=True |
| A crash that must resume | Opt-in + watermarks / context.is_recovery |
Recovery restores your in-memory plan. It reenters from the start |
| Per-user private chat on a shared pool | Delegated user identity and a chosen agent_session_id |
One ID does both jobs |
| Files uploaded before the first model call | Application-created session via the project client | First-request auto-session is early enough for pre-upload |
Voice reconnects and crash recovery both hang on agent_session_id, for opposite reasons. Voice uses it so the same sandbox is still there after 1001. Resilience uses durable work identity so a later process can reclaim a lease. Mixing those without a watermark is how you double-send an email after a voice reconnect that also restarted a handler.
What to do this week
- Name three IDs: conversation (history),
agent_session_id(filesystem + VM), work/response id (durable attempt). If a diagram has one box labeled "session," redraw it. - Voice: declare
invocations_ws2.0.0, size ≥ 1 vCPU / 2 GiB, reconnect on1001with the same session id. Do not put API keys in the container. - Long jobs:
resilient_backgroundonly for stored background Responses, orset_resilient_tasks_enabled(True)before startup. Add a recovered-entry branch and a side-effect fence. Keep the preview disclaimer. - Middle-tier auth: use
x-ms-user-identity/userIdentity; pick pooling vs per-user sandboxes as a cost decision. Billing is CPU+memory of active sessions.
Preview stays preview. Hosted Agents GA does not promote the voice WebSocket path, crash recovery, or the hosting SDK packages.
Sources
- Tina Schuchman, Ship agents faster with expanded model choice, voice agents, and continuous optimization, Azure Blog, 24 September 2026
- Naomi Moneypenny, GPT-6 Astra, Sol, and Luna for production agents in Microsoft Foundry, Azure Blog, 22 September 2026
- Xia Song, More Models, One Copilot, 22 September 2026
- Roger Barreto and Tao Chen, Foundry hosted agent isolation with Microsoft Agent Framework, 22 September 2026
- Microsoft Learn: Build a voice agent with hosted agents
- Microsoft Learn: Resilience for long-running hosted agents (preview)
- Microsoft Learn: Recover long-running work after a crash (preview)
- Microsoft Learn: Hosted agents in Foundry Agent Service
Which of those three IDs is your current hosted agent actually binding on reconnect — conversation, sandbox, or durable work — and what happens to a side effect if the answer is "we used one string for all three"?