SCORE: The Repo That Lets You Watch an AI Rewrite Scientific Code
A Google Research artifact that turns code generation into a visible refinement loop. The real product is not the model output. It is the audit trail.
- SCORE treats AI code generation as an auditable revision trail, so the history of improvement matters as much as the final script.
- The repository is built for scientific software, where robustness, constraints, and reproducibility matter more than a single clever answer.
- Its browser diff viewer is not decoration. It is the mechanism that makes model iteration legible to researchers.
- Compared with general-purpose coding assistants, SCORE is narrower but more honest about the kind of work expert empirical software requires.
SCORE is interesting because the repo makes the model's improvement process visible. Instead of hiding revisions inside a notebook or prompt log, it publishes the diffs. That turns code generation into a forensic artifact, which is exactly what scientific software needs when trust matters.
Code associated with the paper An AI system to help scientists write expert-level empirical software
The Repo Is Mostly a Record of Change
The most revealing files live under docs/, where thousands of HTML diff pages preserve individual revisions. Each page shows an old state, a new state, and just enough browser logic to make the edit readable. The result feels less like a library and more like an archive of attempts.
const codes = JSON.parse(`{
"old_index": 976.0,
"old_code": "...",
"new_code": "..."
}`);
populateDiffs(codes);
Why Scientific Code Needs Visible Iteration
Scientific code is not judged only by whether it runs. It has to respect constraints, handle corner cases, and still produce useful results under a budget. SCORE is built for that world, where a cleaner revision is more valuable than a flashy one-shot answer.
This is not an officially supported Google product.
What SCORE Actually Generates
Inside the diffs, the generated Python looks like empirical software, not toy output. You see model version bumps, constrained configuration lists, and forecasting logic that gets rewritten when a previous choice is too brittle. The repo's point is not that AI can write code. It is that AI can be made to revise code under explicit pressure.
MAX_CONFIGS = 8
config_list = [
{"name": "last_value_naive"},
{"name": "additive_seasonal"},
{"name": "hybrid_decomposition"},
]
MODEL_VERSION = 11
# revised after critique
MODEL_VERSION = 12
How the Browser Diff Viewer Makes the System Legible
The browser layer matters because it changes how you read the system. A plain diff would show edits. SCORE's viewer turns those edits into a navigable sequence, which makes the model's behavior inspectable by someone who is not inside the training loop.
function populateDiffs(codes) {
const oldCode = codes.old_code;
const newCode = codes.new_code;
renderDiff(oldCode, newCode);
}
SCORE vs. General-Purpose Code Assistants
General coding copilots optimize for speed and breadth. SCORE optimizes for traceability and domain fit. That makes it narrower, but also more honest about the job in front of it.
| General code assistant | SCORE |
|---|---|
| Optimizes for quick output and broad usefulness | Optimizes for traceable revision in scientific software |
| Usually shows the final answer, not the path to it | Shows the edit path as the main artifact |
| Fits many coding tasks | Fits constrained empirical and forecasting work |
| Measures success by usefulness at the prompt | Measures success by the quality of the improvement trail |
That is the real lesson in this repo. AI coding tools do not all have to chase the same kind of helpfulness. SCORE suggests a more specialized bargain: if the work is scientific, show the steps, expose the constraints, and make the revision history part of the result.