synbad: The Lie of OpenAI Compatibility
Synbad is a diagnostic tool that proves why your coding agent’s LLM integration keeps breaking on edge cases.
- Synbad shifts the focus of LLM evaluation from prompt accuracy to strict protocol compliance.
- A built-in diagnostic proxy captures raw JSON payloads to diagnose parsing failures in the wild.
- Custom streaming aggregators reconstruct fragmented tool calls and reasoning tokens to expose provider-level bugs.
The Protocol Facade
Every modern inference provider claims to offer an OpenAI-compatible API. This promise is often a facade. While basic chat completions work flawlessly, the illusion shatters when developers introduce complex tool calling or reasoning tokens.
Coding agents rely on perfect JSON formatting and predictable stream chunking to execute functions. When a provider splits a tool name across multiple network packets or malforms a reasoning block, the agent crashes. Synbad was built to catch these exact structural failures.
We never quantize our models, and we run a battery of tool-call and reasoning-parsing tests against our models to make sure they work well in coding agents using our open-source project Synbad.
Stitching the Stream Together
Most LLM bugs occur during streaming. Synbad operates a custom streaming aggregator in its core engine. It manually reconstructs tool calls and reasoning content from fragmented chunks to verify that the provider is sending valid deltas.
The framework uses a dedicated buffer to manage state across these chunks. If a provider splits a JSON argument abruptly, Synbad stitches the payload back together before running its custom assertion library. This ensures that the test fails only if the underlying protocol is actually broken.
Normalizing the Thinkers
The rise of reasoning models like DeepSeek-R1 has introduced a new frontier of API fragmentation. Providers disagree on how to expose the hidden thinking tokens. Some use a dedicated reasoning content field, while others overload existing structures.
export function getReasoning(msg: { reasoning_content?: string, reasoning?: string }) {
return msg.reasoning_content || msg.reasoning;
}
Synbad abstracts this chaos with a multi-spec reasoning handler. It acts as a normalization layer, allowing developers to write a single test that runs consistently across disparate provider implementations.
A Wireshark for LLMs
The most powerful feature of Synbad is its diagnostic proxy. It acts as a Man-in-the-Middle between a user application and the LLM provider. This allows developers to log raw JSON request bodies seamlessly.
When an agent fails in production, developers can capture the exact payload that caused the crash. They can then feed this payload directly into Synbad to create a reproducible test case. It bridges the gap between chaotic real-world usage and structured evaluation.
Evaluation with Purpose
Traditional evaluation frameworks focus heavily on model accuracy and retrieval metrics. Synbad carves out a highly specific niche. It ensures that models can function reliably within coding agent frontends by validating structural integrity.
| Feature | Synbad | Traditional Evals (e.g., DeepEval) |
|---|---|---|
| Primary Focus | Protocol compliance and structure | RAG accuracy and knowledge retrieval |
| Target Workload | Agentic tool calling | Q&A and document summarization |
| Debugging Tools | Built-in MITM HTTP proxy | Synthetic data generation pipelines |
By holding providers accountable for technical precision, Synbad provides a necessary layer of trust. It proves that true compatibility requires more than just a matching endpoint URL.