robbalian/claude-tax-filing Makes Claude Remember a Tax Return

A Python-backed Claude Skill that reads tax PDFs, persists working state on disk, discovers form fields, fills official forms, and verifies the result before it trusts the output.

8 min read • View on GitHub • More from robbalian

A tax desk where paper documents flow into a compact mechanical workstation beside a laptop and a filing cabinet. It explains the article’s main idea: the system moves memory, calculation, and verification outside the model so the workflow can survive forgetting.
The trick is not better small talk with a model. The trick is giving it durable state, local tools, and a final check before output.

Skills are not just a single `.md` file anymore. They can also include scripts, code snippets, and example files, which makes them much more powerful.

robbalian, Project Author/Maintainer · Repository: robbalian/claude-tax-filing
Key Takeaways

Tax filing is a terrible place to fake confidence. The inputs are messy, the forms are brittle, and the cost of a wrong field is not abstract. That is why robbalian/claude-tax-filing is interesting: it shows how to make Claude behave less like a chat window and more like a durable workflow.

The real innovation is not tax logic. It is memory.

The repo’s core move is simple and surprisingly strong. It does not ask Claude to hold an entire return in its context window. Instead, it forces the important state onto disk, then teaches the model to treat those files as the source of truth.

The workflow is built around external memory. When Claude compacts, the files stay.

Robert Balian built a Skill, not just a prompt

That distinction matters. This is not a one-off prompt stuffed into a README. It is a Claude Code Skill with instructions, local scripts, and a working directory, which means the model is being wrapped in software discipline instead of being trusted to improvise its way through a high-stakes task.

A WSJ-style hedcut portrait of the repository owner rendered in black ink on white. It identifies the maintainer behind the Skill and signals that this is a single-author system built around a clear point of view.

That quote captures the repo’s philosophy. The skill file is important, but it is only the control layer. The real reliability comes from the scripts and the files that Claude can read, write, and revisit after the conversation has moved on.

SKILL.md is the control plane

The research notes point to a strict operating model inside SKILL.md. It enforces context budget rules, pushes work into a local work/ directory, and tells Claude when to use Python helpers instead of trying to reason through PDFs in its own context. That is the difference between a flashy demo and an actual system.

skills/tax-filing/
  SKILL.md
  scripts/
    discover_fields.py
    fill_forms.py
    verify_filled.py
  work/

Read that tree as an architecture diagram in miniature. SKILL.md decides the workflow. The scripts do the brittle work. The work/ directory preserves the state that should survive compaction, retries, and human review.

The PDF work is split into discovery, filling, and verification

This is where the repo gets very practical. Government PDFs are not friendly data entry forms. They can hide fields in XFA or deeply nested AcroForm structures, which is why generic LLM prompting is the wrong tool for the job. The repo uses targeted Python utilities to discover field IDs, map them compactly, and fill them deterministically.

A close-up of a magnified tax form field being inspected like a machine part. The image explains why PDF field discovery needs dedicated tooling: the model must learn the hidden structure before it can fill anything reliably.
Government forms are not just documents. They are structured objects with hidden IDs, nested fields, and traps for guesswork.

The useful detail is not that the scripts are clever. It is that they are narrow. discover_fields.py maps the form, fill_forms.py writes the data, and verify_filled.py checks the result. Claude is asked to orchestrate, not to hallucinate field-level precision.

Verification is what makes the system credible

A desktop scene where a completed tax form is checked against a reference sheet under a loupe. It shows that the workflow does not trust the first output. It reads the result back and compares it to what should have been written.
The final step is not generation. It is proof that the generated PDF matches expected values.

That closed loop is the real reliability story. The repo does not stop at a filled PDF. It reads the output back, compares it to an expected file, and makes the model prove it did the job before the result is treated as final. In high-friction workflows, that is the difference between automation and theater.

Compared with TurboTax, fillable forms, and a plain prompt

ApproachHow it worksMemory / stateVerificationBest fit
robbalian/claude-tax-filingClaude orchestrates local Python helpers that read docs, fill PDFs, and audit the result.Persistent files in a work directory.Scripted read-back against expected values.Power users who want auditable automation over a brittle workflow.
TurboTax-style softwareA guided interview hides the complexity behind a product UI.State lives inside the product.Built-in validation and product rules.Broad consumer filing with the least friction.
IRS Free File Fillable FormsUsers enter data directly into official-style forms.State is mostly in the user’s head.Minimal guidance and limited guardrails.People who already know the forms and want a free path.
Plain LLM promptA chat model is asked to reason through the task ad hoc.Only the context window unless you add more.None unless the user builds it manually.Quick questions, not a high-stakes filing workflow.

The table points to the real trade-off. Guided software is polished but closed. Fillable forms are free but unforgiving. A plain prompt is flexible but too forgetful. This repo sits in the awkward middle, where you give up slickness in exchange for a system that can actually survive the task.

What this repo suggests about agent software

The broader lesson is bigger than taxes. If an LLM workflow is long, brittle, or regulated, the model should not be the only place where important state lives. Put memory on disk. Put precision in scripts. Make verification mandatory. That is how you turn a clever assistant into something closer to software.

Seen that way, robbalian/claude-tax-filing is less about filing taxes and more about durable agent design. It shows a practical pattern for high-stakes work: keep the model inside a narrow operating envelope, let deterministic tools do the exacting tasks, and require proof before you trust the output.