Human-in-the-Loop Approvals in Microsoft Foundry: Pausing Long-Running Agents Indefinitely Without Losing State
DEV Community

Human-in-the-Loop Approvals in Microsoft Foundry: Pausing Long-Running Agents Indefinitely Without Losing State

Human-in-the-Loop Approvals in Microsoft Foundry: Pausing Long-Running Agents Indefinitely Without Losing State Day 16 of Foundry 100 Days / 100 Blogs Table of Contents - The Problem: Agents That Need a Human, Not a Retry - Why This Matters - Core Concepts: Task IDs, Entry Modes, and Suspension - Architecture: Where the Approval State Actually Lives - Implementing an Approval Turn Step by Step - Wiring the Human Side: Notifications and Resume Triggers - Framework Interrupts: LangGraph and Microsoft Agent Framework - A Real-World Scenario: Expense Approval With Escalation - Production Considerations - Security Considerations - Performance and Scale - Cost Considerations - Common Mistakes and Pitfalls - Alternatives and Trade-offs - Practical Recommendations - Conclusion - References The Problem: Agents That Need a Human, Not a Retry Most agent failure modes are things you engineer around: a tool call times out, a model hallucinates a malformed JSON payload, a rate limit trips. You retry, you fence the side effect, you fall back to a smaller model. Microsoft Foundry's resilient task subsystem - leases, checkpoints, entry_mode recovery - exists precisely to make those failures invisible to the user (see Day 1 of this series, on crash-resilient long-running agents). But there's a category of "interruption" that isn't a failure at all: the agent is working correctly, and it has simply reached a point where it is not allowed to keep going without a person saying yes. A finance agent that wants to submit a $1,200 expense reimbursement. A DevOps agent that wants to run terraform apply against production. A support agent that wants to issue a refund above a threshold. In every one of these cases, the "right" behavior is not to fail, retry, or guess - it's to stop, ask, and wait, for however long it takes a human to look at a Slack message, a ticket queue, or an approval inbox. That could be ninety seconds. It could be three days if the approver is on vacation. Most agent runtimes handle this badly. If your agent is a synchronous request/response call sitting behind an HTTP connection, you cannot hold that connection open for three days. If you fake it with polling and an external state machine (a Durable Function, a Temporal workflow, a hand-rolled Postgres table with a status column), you've now built and are maintaining a second orchestration layer outside the agent runtime, with its own failure modes, and the agent's own conversational state (tool call history, prior turns, checkpoints) lives somewhere else entirely from the approval state. Reconciling the two after a crash is exactly the kind of glue code nobody wants to own. Microsoft Foundry's Agent Service takes a different position: human-in-the-loop approval isn't bolted onto the long-running agent primitives as a separate feature - it's a natural consequence of how multi_turn_task chains already work. If you understood Day 1's resilience model, you already have 80% of the mental model for how approvals work. This article is the other 20%: how to actually build the pause point, how to drive it from an application, how it survives a crash mid-pause, and where the sharp edges are in production. Why This Matters Enterprise AI adoption keeps running into the same wall: organizations are comfortable letting an agent draft an action but not execute it unattended, especially for anything touching money, infrastructure, or customer-facing communication. Every serious agent framework has converged on some notion of "approval gate" - LangGraph has interrupt() , Microsoft Agent Framework has RequestInfoEvent and ApprovalRequiredAIFunction (which we touched on in Day 7's workflow migration piece), CrewAI has human input tools. What's different about Foundry's approach is that it doesn't treat the approval pause as a special-cased control-flow primitive bolted on top of an ephemeral request handler. It treats it as an ordinary suspended state of a durable, server-tracked task - the same durability substrate used for crash recovery. That matters for three concrete reasons developers should care about: - You don't need a separate orchestration system. The task_id that identifies your approval chain is the same identity used for lease-based crash recovery. There's no second source of truth to keep synchronized. - The wait has no artificial ceiling. Because the chain is durable and not tied to a live process or open connection, "wait for approval" can mean seconds or it can mean a week over a holiday, with identical code. - It composes with the framework layer. If you're already building on LangGraph or Microsoft Agent Framework over the Responses protocol, the approval interrupt is just another checkpoint boundary that resilient_background=True already knows how to persist and rehydrate. Getting this pattern right is the difference between an agent that enterprises trust with consequential actions and one that gets restricted to read-only, "suggest but don't act" duty forever. Core Concepts: Task IDs, Entry Modes, and Suspension To build an approval step correctly you need to be precise about four concepts in the AgentServer SDK (azure-ai-agentserver-core ≥ 2.0.0 for Python, Azure.AI.AgentServer.Core ≥ 1.0.0-beta.28 for .NET, both currently preview surfaces subject to change): @multi_turn_task is a decorator that turns an async handler into a durable conversation chain. Unlike a one-shot @task (input in, output out, done), a multi-turn task doesn't terminate when the handler returns - it transitions into a suspended state and stays alive under a single task_id until either a new turn arrives or you explicitly delete it. task_id is the durable work identity that scopes the whole chain. It's caller-chosen, not server-generated, which is the detail that makes human-in-the-loop possible: your application decides the identity up front ("exp-42" for an expense report, a conversation thread ID, a ticket number), and every subsequent turn - including the human's reply, arriving possibly days later from a completely different process - reenters the same chain by reusing that same string. entry_mode on TaskContext tells your handler why it's being invoked right now. There are three values: - fresh - first execution for this(task_id, input_id) pair. - resumed - a subsequent turn on an existing chain (this is what fires when the human's decision comes back in). - recovered - the container crashed mid-attempt in a previous lifetime and the framework is re-invoking the same attempt from persisted input, without your explicit involvement. This three-way split is the crux of the whole pattern. Your handler branches on entry_mode to decide whether it's starting fresh, picking up a human decision, or being silently retried after an infrastructure hiccup. Critically, resumed and recovered are different things: resumed is an intentional new turn (the human replied), while recovered is the framework protecting you from a crash that happened before your fresh or resumed turn even finished. Entry modes govern how the framework re-enters your handler: fresh for the first execution, resumed for a genuine next turn (human decision or scheduled check), and recovered when a crash interrupts an attempt before it finishes. ctx.metadata is small, durable key-value state attached to the task that survives the suspension. The documentation is explicit that this should hold only small references - an expense ID, a step counter - not full payloads. The full request history and generated artifacts belong in your own storage or a FoundryStateStore -backed checkpoint (see Day 1 for the checkpoint/watermark pattern in depth). The diagram below shows how these pieces fit together end to end. The suspended chain lives in the durable state store, not in a live process - the container that handled turn 1 can exit entirely before a human ever replies. Architecture: Where the Approval State Actually Lives It's worth being explicit about what's happening at the infrastructure level, because "it just suspends" hides a few architectural decisions that matter once you're debugging a stuck task in production. When a @multi_turn_task handler returns without raising, the TaskManager - a server-side component that Foundry's Agent Service constructs when you call set_resilient_tasks_enabled(True) before host startup - writes the chain's current state to the durable state store and transitions its status to Suspended . This is not an in-memory pause. The container that handled turn 1 can be killed, scaled to zero, or replaced by a new revision entirely, and turn 2 can be served by a completely different container instance, because nothing about the suspension depends on process memory. The only thing that has to survive is the record in the state store and (if you're using framework-level checkpointing) the serialized framework state you wrote there yourself. This has a direct, useful consequence: an approval wait is not a container-hours cost. A suspended multi_turn_task isn't a thread blocked on input() , and it isn't a container kept warm waiting for a callback. The container that ran turn 1 can exit completely. Whatever compute picks up turn 2 - hours or days later - is a fresh invocation against the same task_id , and the framework's job is purely to route it to the right handler with the right persisted context, not to keep a process alive across the gap. The TaskStatus enum reflects this lifecycle explicitly: Pending → InProgress → Suspended → Completed . A chain sitting in Suspended is a durable database row (conceptually), not a live process. When you eventually call await approve.delete("exp-42") , you're deleting that durable record - worth noting because the framework does not garbage-collect suspended multi-turn chains automatically the way it cleans up completed one-shot @task records. If your approval chains never get an explicit resolution (the approver never replies, the ticket gets abandoned), you will accumula

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.