call-me-skill: The CLI Agent That Calls You Back

Breaking the terminal tether with asynchronous voice handoffs and persistent memory loops.

• View on GitHub • More from zarazhangrui

A vintage rotary phone connected to a computer terminal by a glowing fiber-optic cable. This illustrates the concept of bridging legacy voice communication with modern asynchronous AI terminal tasks.
The call-me-skill bridges the gap between CLI execution and human voice interaction.

Key Takeaways

The End of Terminal Babysitting

The Agentic Era has a user experience problem. Developers are forced to babysit long-running LLM tasks like complex refactors or database migrations. The agent might need a creative decision or a permission gate at any moment. This creates a "Terminal Tether" where the user cannot safely step away.

Imagine this: you kick off a long-running task in your terminal and step away to grab coffee. A few minutes later, your phone rings — it's Claude, telling you the task is complete and asking what to do next. Science fiction? Nope. It's real, thanks to the CallMe plugin for Claude Code.

The call-me-skill repository solves this by turning a CLI agent into a literal caller. It transforms the developer from a passive monitor into an active stakeholder. When a task requires input, the agent pauses its work and calls your phone.

Architecture of an Outbound Thought

The project relies entirely on Shell scripts to orchestrate the handoff. This design choice makes the skill ultra-portable and "agent-readable." AI agents like OpenClaw or Claude Code can easily inspect and modify the Bash execution logic on the fly.

How the asynchronous CLI process bridges to a synchronous phone call.

The execution begins with trigger-call.sh. The agent initiates an outbound call via the Retell AI API. Because a phone call is inherently synchronous but terminal tasks are asynchronous, the system uses a pull architecture. The poll-transcript.sh script loops quietly in the background, dipping into the API to check if the user has hung up and the transcript is ready.

A close-up of a mechanical gear system dipping a small bucket into a pool of water, representing a polling loop checking for data.
The polling loop abstracts the complexity of webhooks, allowing the agent to run strictly locally.

Beyond Notifications: The Memory Loop

Basic notification scripts suffer from amnesia. Every phone call feels like a first encounter. This repository introduces sync-memory.sh to solve the Groundhog Day problem.

After a call concludes, the agent summarizes the transcript. It extracts user preferences and operational decisions, pushing them back into the Retell Knowledge Base. The next time the agent calls, it already knows how you prefer to handle specific edge cases.

The memory synchronization loop ensures the agent learns from each conversation.

The Latency Trade-off

The voice agent landscape is divided by latency. Realtime WebSockets offer sub-second response times, mimicking human interruption. REST-based pipelines introduce a two to four second delay.

For content generation and deep technical handoffs, a slight delay is actually a feature. It provides the LLM with crucial reasoning time to parse complex logic before speaking, ensuring higher quality output over conversational speed.

ProjectArchitectureLatencyPrimary Use Case
call-me-skillREST / Polling~2-4 secondsAsync developer handoffs.
supercallOpenAI Realtime API<1 secondFluid conversational chat.
LiveKit-AgentSIP Trunking / WebRTC~1-2 secondsEnterprise inbound support.