synbad: The Lie of OpenAI Compatibility

Synbad is a diagnostic tool that proves why your coding agent’s LLM integration keeps breaking on edge cases.

8 min read • View on GitHub • More from synthetic-lab

An inspector examining a series of identical shipping containers, finding subtle cracks and malformations in the standardization.
Synbad inspects the structural integrity of LLM responses, ensuring strict compliance with API protocols.
Key Takeaways

The Protocol Facade

Every modern inference provider claims to offer an OpenAI-compatible API. This promise is often a facade. While basic chat completions work flawlessly, the illusion shatters when developers introduce complex tool calling or reasoning tokens.

Coding agents rely on perfect JSON formatting and predictable stream chunking to execute functions. When a provider splits a tool name across multiple network packets or malforms a reasoning block, the agent crashes. Synbad was built to catch these exact structural failures.

We never quantize our models, and we run a battery of tool-call and reasoning-parsing tests against our models to make sure they work well in coding agents using our open-source project Synbad.

Synthetic.new, Project Team · Synthetic.new ♥️ OpenCode

Stitching the Stream Together

Most LLM bugs occur during streaming. Synbad operates a custom streaming aggregator in its core engine. It manually reconstructs tool calls and reasoning content from fragmented chunks to verify that the provider is sending valid deltas.

The framework uses a dedicated buffer to manage state across these chunks. If a provider splits a JSON argument abruptly, Synbad stitches the payload back together before running its custom assertion library. This ensures that the test fails only if the underlying protocol is actually broken.

Synbad's aggregator reconstructing fragmented stream chunks into valid JSON payloads.

Normalizing the Thinkers

The rise of reasoning models like DeepSeek-R1 has introduced a new frontier of API fragmentation. Providers disagree on how to expose the hidden thinking tokens. Some use a dedicated reasoning content field, while others overload existing structures.

export function getReasoning(msg: { reasoning_content?: string, reasoning?: string }) {
  return msg.reasoning_content || msg.reasoning;
}

Synbad abstracts this chaos with a multi-spec reasoning handler. It acts as a normalization layer, allowing developers to write a single test that runs consistently across disparate provider implementations.

A Wireshark for LLMs

The most powerful feature of Synbad is its diagnostic proxy. It acts as a Man-in-the-Middle between a user application and the LLM provider. This allows developers to log raw JSON request bodies seamlessly.

When an agent fails in production, developers can capture the exact payload that caused the crash. They can then feed this payload directly into Synbad to create a reproducible test case. It bridges the gap between chaotic real-world usage and structured evaluation.

A mechanical net catching data packets flowing along a wire between two large machines.
The diagnostic proxy intercepts and logs payloads for reproducible debugging.

Evaluation with Purpose

Traditional evaluation frameworks focus heavily on model accuracy and retrieval metrics. Synbad carves out a highly specific niche. It ensures that models can function reliably within coding agent frontends by validating structural integrity.

FeatureSynbadTraditional Evals (e.g., DeepEval)
Primary FocusProtocol compliance and structureRAG accuracy and knowledge retrieval
Target WorkloadAgentic tool callingQ&A and document summarization
Debugging ToolsBuilt-in MITM HTTP proxySynthetic data generation pipelines

By holding providers accountable for technical precision, Synbad provides a necessary layer of trust. It proves that true compatibility requires more than just a matching endpoint URL.