async-deep-agents: The End of the Blocking Agent
How an ASGI transport layer and a clever state-machine middleware are turning monolithic AI scripts into truly non-blocking, background microservices.

A reference implementation for non-blocking, background agent orchestration on LangSmith Deployments using Deep Agents.
- Async Deep Agents shift multi-agent orchestration from sequential blocking tasks to background microservice execution.
- An ASGI transport layer enables co-deployed graphs to communicate in-process without HTTP latency overhead.
- A custom middleware hacks the LangGraph lifecycle to inject push notifications directly into a parent thread when a subagent finishes.
- Isolating long-horizon tasks on separate threads prevents supervisor context bloat and the lost-in-the-middle phenomenon.
The UX Bottleneck of Sequential AI
Sequential processing is the original bottleneck of modern AI orchestration. When a supervisor agent delegates a complex research task, it traditionally waits for the response. The user thread is blocked. The application feels sluggish and unresponsive. We are effectively forcing parallel computing architectures to behave like single-threaded scripts.
The Micro-Agent Architecture
The langchain-ai/async-deep-agents repository introduces a profound shift by treating subagents as asynchronous background microservices. Built as a polyglot monorepo with strict parity between Python and TypeScript, it uses LangGraph to define a supervisor that delegates to specialized worker nodes. Crucially, the supervisor invokes these nodes using an ASGI transport layer.
By omitting the explicit URL in the agent specification, the system routes traffic in-memory between co-deployed graphs. This eliminates HTTP latency entirely while maintaining a modular architecture where researchers and coders operate in isolation.
The Push Notification Hack
The technical climax of this architecture is the CompletionNotifierMiddleware. In a standard asynchronous setup, a supervisor must poll the worker to check its status. This middleware implements a webhook-style push notification instead.
When a subagent completes its task, the after_agent hook triggers. It uses the LangGraph SDK to create a new run on the parent's thread. This effectively wakes up the supervisor by injecting a message into its conversation history, informing it that the background task is complete.
async def after_agent(self, run: Run, client: AsyncLangGraphClient):
if run.status == "success":
await client.runs.create(
thread_id=parent_thread_id,
assistant_id="supervisor",
input={"messages": [{"role": "system", "content": "Background task complete."}]}
)
Context as the New Memory
This decoupling offers a massive secondary benefit in the form of context window preservation. By isolating long-horizon tasks on separate threads, the supervisor's context remains pristine.
The supervisor never ingests the raw intermediate steps of a deep research or coding task. It only receives the final high-level summary. This prevents the dreaded lost-in-the-middle phenomenon that plagues dense LLM prompts.
Orchestration vs. Roleplay
This low-level state machine approach stands in stark contrast to higher-level frameworks. Production applications require strict lifecycle management, including the ability to launch, monitor, steer, and explicitly cancel background jobs.
| Feature | Async Deep Agents | CrewAI | AutoGen |
|---|---|---|---|
| Execution Model | Non-blocking background jobs | Sequential roleplay | Conversational turns |
| State Persistence | Granular graph state updates | In-memory session | Message history |
| Primary Use Case | Long-horizon autonomous tasks | Collaborative writing | Multi-agent chat simulations |