async-deep-agents: The End of the Blocking Agent

How an ASGI transport layer and a clever state-machine middleware are turning monolithic AI scripts into truly non-blocking, background microservices.

8 min read • View on GitHub • More from langchain-ai

A massive train station where a single large train blocks the main track above, while a network of pneumatic tubes routes hundreds of capsules efficiently below.
Synchronous agents block the main thread like a stalled train. Asynchronous execution routes tasks through parallel, non-blocking channels.

A reference implementation for non-blocking, background agent orchestration on LangSmith Deployments using Deep Agents.

hntrl, Primary Contributor · langchain-ai/async-deep-agents
Key Takeaways

The UX Bottleneck of Sequential AI

Sequential processing is the original bottleneck of modern AI orchestration. When a supervisor agent delegates a complex research task, it traditionally waits for the response. The user thread is blocked. The application feels sluggish and unresponsive. We are effectively forcing parallel computing architectures to behave like single-threaded scripts.

The Micro-Agent Architecture

The langchain-ai/async-deep-agents repository introduces a profound shift by treating subagents as asynchronous background microservices. Built as a polyglot monorepo with strict parity between Python and TypeScript, it uses LangGraph to define a supervisor that delegates to specialized worker nodes. Crucially, the supervisor invokes these nodes using an ASGI transport layer.

By omitting the explicit URL in the agent specification, the system routes traffic in-memory between co-deployed graphs. This eliminates HTTP latency entirely while maintaining a modular architecture where researchers and coders operate in isolation.

The Completion Notifier Middleware intercepts the end of a background job and pushes a notification back to the main thread.

Portrait of hntrl

The Push Notification Hack

The technical climax of this architecture is the CompletionNotifierMiddleware. In a standard asynchronous setup, a supervisor must poll the worker to check its status. This middleware implements a webhook-style push notification instead.

When a subagent completes its task, the after_agent hook triggers. It uses the LangGraph SDK to create a new run on the parent's thread. This effectively wakes up the supervisor by injecting a message into its conversation history, informing it that the background task is complete.

async def after_agent(self, run: Run, client: AsyncLangGraphClient):
    if run.status == "success":
        await client.runs.create(
            thread_id=parent_thread_id,
            assistant_id="supervisor",
            input={"messages": [{"role": "system", "content": "Background task complete."}]}
        )

Context as the New Memory

This decoupling offers a massive secondary benefit in the form of context window preservation. By isolating long-horizon tasks on separate threads, the supervisor's context remains pristine.

The supervisor never ingests the raw intermediate steps of a deep research or coding task. It only receives the final high-level summary. This prevents the dreaded lost-in-the-middle phenomenon that plagues dense LLM prompts.

A close-up of a vintage telegraph relay switch connected by a taut wire to a bright signal lamp on a separate control panel.
The middleware acts as a mechanical relay, bridging two isolated contexts to trigger an alert only when the job is done.

Orchestration vs. Roleplay

This low-level state machine approach stands in stark contrast to higher-level frameworks. Production applications require strict lifecycle management, including the ability to launch, monitor, steer, and explicitly cancel background jobs.

FeatureAsync Deep AgentsCrewAIAutoGen
Execution ModelNon-blocking background jobsSequential roleplayConversational turns
State PersistenceGranular graph state updatesIn-memory sessionMessage history
Primary Use CaseLong-horizon autonomous tasksCollaborative writingMulti-agent chat simulations