robbalian/claude-tax-filing Makes Claude Remember a Tax Return
A Python-backed Claude Skill that reads tax PDFs, persists working state on disk, discovers form fields, fills official forms, and verifies the result before it trusts the output.

Skills are not just a single `.md` file anymore. They can also include scripts, code snippets, and example files, which makes them much more powerful.
- The repo’s real breakthrough is not tax automation, but durable agent design that keeps memory, calculations, and verification on disk.
- SKILL.md acts like a control plane, while Python scripts handle the exact PDF work the model should not improvise.
- Field discovery and closed-loop verification matter more than prompt cleverness in a workflow where forms punish guesswork.
- Compared with a plain prompt or manual forms, the skill trades polish for a more auditable and failure-resistant process.
Tax filing is a terrible place to fake confidence. The inputs are messy, the forms are brittle, and the cost of a wrong field is not abstract. That is why robbalian/claude-tax-filing is interesting: it shows how to make Claude behave less like a chat window and more like a durable workflow.
The real innovation is not tax logic. It is memory.
The repo’s core move is simple and surprisingly strong. It does not ask Claude to hold an entire return in its context window. Instead, it forces the important state onto disk, then teaches the model to treat those files as the source of truth.
Robert Balian built a Skill, not just a prompt
That distinction matters. This is not a one-off prompt stuffed into a README. It is a Claude Code Skill with instructions, local scripts, and a working directory, which means the model is being wrapped in software discipline instead of being trusted to improvise its way through a high-stakes task.
That quote captures the repo’s philosophy. The skill file is important, but it is only the control layer. The real reliability comes from the scripts and the files that Claude can read, write, and revisit after the conversation has moved on.
SKILL.md is the control plane
The research notes point to a strict operating model inside SKILL.md. It enforces context budget rules, pushes work into a local work/ directory, and tells Claude when to use Python helpers instead of trying to reason through PDFs in its own context. That is the difference between a flashy demo and an actual system.
skills/tax-filing/
SKILL.md
scripts/
discover_fields.py
fill_forms.py
verify_filled.py
work/
Read that tree as an architecture diagram in miniature. SKILL.md decides the workflow. The scripts do the brittle work. The work/ directory preserves the state that should survive compaction, retries, and human review.
The PDF work is split into discovery, filling, and verification
This is where the repo gets very practical. Government PDFs are not friendly data entry forms. They can hide fields in XFA or deeply nested AcroForm structures, which is why generic LLM prompting is the wrong tool for the job. The repo uses targeted Python utilities to discover field IDs, map them compactly, and fill them deterministically.
The useful detail is not that the scripts are clever. It is that they are narrow. discover_fields.py maps the form, fill_forms.py writes the data, and verify_filled.py checks the result. Claude is asked to orchestrate, not to hallucinate field-level precision.
Verification is what makes the system credible
That closed loop is the real reliability story. The repo does not stop at a filled PDF. It reads the output back, compares it to an expected file, and makes the model prove it did the job before the result is treated as final. In high-friction workflows, that is the difference between automation and theater.
Compared with TurboTax, fillable forms, and a plain prompt
| Approach | How it works | Memory / state | Verification | Best fit |
|---|---|---|---|---|
| robbalian/claude-tax-filing | Claude orchestrates local Python helpers that read docs, fill PDFs, and audit the result. | Persistent files in a work directory. | Scripted read-back against expected values. | Power users who want auditable automation over a brittle workflow. |
| TurboTax-style software | A guided interview hides the complexity behind a product UI. | State lives inside the product. | Built-in validation and product rules. | Broad consumer filing with the least friction. |
| IRS Free File Fillable Forms | Users enter data directly into official-style forms. | State is mostly in the user’s head. | Minimal guidance and limited guardrails. | People who already know the forms and want a free path. |
| Plain LLM prompt | A chat model is asked to reason through the task ad hoc. | Only the context window unless you add more. | None unless the user builds it manually. | Quick questions, not a high-stakes filing workflow. |
The table points to the real trade-off. Guided software is polished but closed. Fillable forms are free but unforgiving. A plain prompt is flexible but too forgetful. This repo sits in the awkward middle, where you give up slickness in exchange for a system that can actually survive the task.
What this repo suggests about agent software
The broader lesson is bigger than taxes. If an LLM workflow is long, brittle, or regulated, the model should not be the only place where important state lives. Put memory on disk. Put precision in scripts. Make verification mandatory. That is how you turn a clever assistant into something closer to software.
Seen that way, robbalian/claude-tax-filing is less about filing taxes and more about durable agent design. It shows a practical pattern for high-stakes work: keep the model inside a narrow operating envelope, let deterministic tools do the exacting tasks, and require proof before you trust the output.