SCORE: The Repo That Lets You Watch an AI Rewrite Scientific Code

A Google Research artifact that turns code generation into a visible refinement loop. The real product is not the model output. It is the audit trail.

8 min read · google-research/score

A browser window stretches across a white field, showing a long trail of code diffs that feel like an evidence file. It explains that SCORE is less about a final script and more about the visible sequence of revisions that produced it.
The repo reads like an audit log, not a single release artifact.
Key Takeaways

SCORE is interesting because the repo makes the model's improvement process visible. Instead of hiding revisions inside a notebook or prompt log, it publishes the diffs. That turns code generation into a forensic artifact, which is exactly what scientific software needs when trust matters.

Code associated with the paper An AI system to help scientists write expert-level empirical software

google-research/score repository description, Source Code Repository · google-research/score

The Repo Is Mostly a Record of Change

The most revealing files live under docs/, where thousands of HTML diff pages preserve individual revisions. Each page shows an old state, a new state, and just enough browser logic to make the edit readable. The result feels less like a library and more like an archive of attempts.

const codes = JSON.parse(`{
  "old_index": 976.0,
  "old_code": "...",
  "new_code": "..."
}`);

populateDiffs(codes);

Why Scientific Code Needs Visible Iteration

Scientific code is not judged only by whether it runs. It has to respect constraints, handle corner cases, and still produce useful results under a budget. SCORE is built for that world, where a cleaner revision is more valuable than a flashy one-shot answer.

This is not an officially supported Google product.

google-research/score README, README · google-research/score README

What SCORE Actually Generates

Inside the diffs, the generated Python looks like empirical software, not toy output. You see model version bumps, constrained configuration lists, and forecasting logic that gets rewritten when a previous choice is too brittle. The repo's point is not that AI can write code. It is that AI can be made to revise code under explicit pressure.

A close-up desk scene shows a pen, a terminal window, and a checklist as a rough draft becomes a tighter implementation. It explains SCORE's loop of critique, revision, and constraint, which is the engine behind the repository.
The system treats improvement as a sequence of explicit edits.
MAX_CONFIGS = 8

config_list = [
    {"name": "last_value_naive"},
    {"name": "additive_seasonal"},
    {"name": "hybrid_decomposition"},
]

MODEL_VERSION = 11
# revised after critique
MODEL_VERSION = 12

A visible correction loop is what makes the system legible to researchers, not just useful to the model.

How the Browser Diff Viewer Makes the System Legible

The browser layer matters because it changes how you read the system. A plain diff would show edits. SCORE's viewer turns those edits into a navigable sequence, which makes the model's behavior inspectable by someone who is not inside the training loop.

function populateDiffs(codes) {
  const oldCode = codes.old_code;
  const newCode = codes.new_code;
  renderDiff(oldCode, newCode);
}

SCORE vs. General-Purpose Code Assistants

General coding copilots optimize for speed and breadth. SCORE optimizes for traceability and domain fit. That makes it narrower, but also more honest about the job in front of it.

One side of the scene shows a sealed black box that spits out a single script. The other side shows a transparent staircase of revisions with each step exposed. It explains why SCORE is easier to trust in scientific work than one-shot code generation.
Transparency is the feature, not the decoration.
General code assistantSCORE
Optimizes for quick output and broad usefulnessOptimizes for traceable revision in scientific software
Usually shows the final answer, not the path to itShows the edit path as the main artifact
Fits many coding tasksFits constrained empirical and forecasting work
Measures success by usefulness at the promptMeasures success by the quality of the improvement trail

That is the real lesson in this repo. AI coding tools do not all have to chase the same kind of helpfulness. SCORE suggests a more specialized bargain: if the work is scientific, show the steps, expose the constraints, and make the revision history part of the result.