pr_sync_agent: The Repo That Reviews Itself While It Syncs
A two-agent Python pipeline that turns GitHub PRs into structured AI reviews, then mirrors approved file changes across repositories without a local clone.
- pr_sync_agent treats GitHub as the primary system of record and mutates repositories file by file through the API, so it can run without a local clone.
- The project splits review and replication into separate agents, which makes the control loop easier to reason about than a single monolithic bot.
- Its most interesting implementation trick is the OpenAI-shaped client pointed at Gemini, which lets the review stack swap model providers without rewriting the orchestration.
- The API-only approach is elegant for ephemeral CI and mirrored repos, but it also makes scale and rate limits part of the design conversation.
A Sync Bot That Never Clones
The surprising thing about pr_sync_agent is not that it uses an LLM. It is that the sync path itself is built around the GitHub API instead of git clone. Files are created, updated, and removed one at a time in the destination repo, which turns GitHub into the place where the work happens, not just where the work is stored.
That makes the project feel less like a script and more like an API-native control plane for mirrored codebases. It is a strong fit for ephemeral CI, where runners are disposable, credentials are narrow, and a full checkout is often more baggage than value.
for file in pr.get_files():
if file.status == "removed":
repo.delete_file(path=file.filename, sha=file.sha, branch=target_branch)
elif file.status == "modified":
repo.update_file(path=file.filename, content=file.patch, sha=file.sha, branch=target_branch)
elif file.status == "added":
repo.create_file(path=file.filename, content=file.patch, branch=target_branch)
Two Agents, One Loop
The repo is built as two separate agents, and that separation matters. One agent reads the PR diff and writes a structured review. The other agent mirrors the change set into a target repository and opens a linked PR there. That means review and replication do not compete for the same logic path.
| Traditional clone-based pipeline | pr_sync_agent |
|---|---|
| Checkout repo, run shell sync, push changes | Call GitHub APIs file by file |
| Needs local git state and often SSH credentials | Works with API tokens and ephemeral runners |
| Review is usually a separate manual or CI step | Review becomes part of the same control loop |
| Easy to batch large file sets | Cleaner for smaller, observable sync operations |
| Fast to reason about in a shell | Cleaner to compose across services |
The sync agent handles three file states: added, modified, and removed. That sounds ordinary until you notice the implied tradeoff. Instead of one bulk push, the system expresses every change as an explicit API action, which is easier to audit but can become chatty when the sync grows large.
Why the GitHub API Matters
This design removes a lot of runner baggage. There is no working tree to maintain, no local merge base to manage, and no SSH dependency to provision. For teams that maintain downstream mirrors or shadow repos, that is a real simplification.
The cost is just as real. API commits can be verbose, and rate limits become part of the architecture instead of an implementation detail. The approach is strongest when the goal is repeatable, observable sync, not bulk history rewriting or heavyweight repository surgery.
That tradeoff is easy to miss if you only look at the happy path. But the project is honest about its shape: it favors portability and clarity over raw throughput.
The OpenAI-to-Gemini Bridge
One of the smartest implementation details in the repo is the model bridge. The code uses LangChain's OpenAI client shape, but points it at Google's OpenAI-compatible Gemini endpoint. In practice, that means the orchestration layer can stay stable while the provider behind it changes.
llm = ChatOpenAI(
model="gemini-3-flash-preview",
openai_api_base="https://generativelanguage.googleapis.com/v1beta/openai/",
openai_api_key=os.environ["GOOGLE_API_KEY"],
)
That is less a model story than an interface story. The value is not just that Gemini can read long diffs. It is that the surrounding system treats the model as a swappable component, which is exactly how you want an automation pipeline to age.
The Operator’s Console
The Streamlit app gives the repo a second personality. It is not only a CI automation package. It is also a small control room for testing prompts, inspecting behavior, and running an LLM sandbox before anything is posted back to GitHub.
That matters because sync and review systems fail in subtle ways. A sandbox makes it easier to see whether the prompt is too blunt, whether the diff is too noisy, or whether the agent should be dry-run only. In other words, the UI is there to keep the automation legible.
The project also uses a straightforward configuration pattern: environment variables first, then Streamlit secrets. That keeps the same code usable in local development, GitHub Actions, and hosted Streamlit deployments without forking the setup story.
What This Replaces, and What It Cannot
| Old pattern | This repo |
|---|---|
| Clone, patch, push | API-only file mutation |
| One bot per task | Review agent plus sync agent |
| Opaque CI shell steps | Visible PR and file-level actions |
| Single-provider coupling | OpenAI-shaped wrapper with Gemini behind it |
| Broadly flexible, often messy | Narrower, cleaner, and easier to audit |
The best way to read pr_sync_agent is as a challenge to old CI habits. It replaces clone-heavy glue with a GitHub-native loop, then adds an automated reviewer so the sync itself can be checked as it moves. That is a neat shape for teams that live in mirrored repositories and want the machinery to stay visible.
It is not a universal answer. Large syncs will stress the API, and the project still looks like early-stage software rather than hardened infrastructure. But as a pattern, it is compelling: GitHub becomes the transport, the audit trail, and the review surface all at once.