`tips-for-running-modal`: Training code built to survive being killed

A compact Modal scaffold for PyTorch that treats checkpoints, preemption, and `torch.compile` edge cases as first-class problems instead of afterthoughts.

8 min read • View on GitHub • More from chenglou

A cloud GPU workshop where a training loop runs beside a locked checkpoint vault and a visible kill switch. The scene explains the article's core idea: long-running ML jobs are designed to resume after interruption instead of depending on one uninterrupted process.
The repo's real subject is not cloud compute. It is what happens after the process dies.
Key Takeaways

Most ML templates optimize for the happy path. This one optimizes for the moment the run gets preempted, the container disappears, and the only useful thing left is a checkpoint on a volume. That shift changes everything: where state lives, when a job is allowed to resume, and what counts as correctness.

A training job that expects to die

The repository reads like a small operating manual for serverless GPU work. It assumes that cloud training is interruptible, that detached runs are normal, and that the filesystem attached to compute is not where truth lives.

That is why the scaffold is so compact. The repo is not trying to become a framework. It is trying to encode a reliable habit: train somewhere ephemeral, write progress somewhere durable, and make resume a first-class path instead of a lucky accident.

The documentation folders matter here too. Files like AGENTS.md and CLAUDE.md signal that the repository expects to be reused by humans and assistants without losing the operating rules. This is less a library than a copyable contract.

The checkpoint contract

The sharpest idea in the repo is not saving state. It is refusing to resume under the wrong state. The checkpoint loader scans for the latest saved file, compares the saved config with the current config, and fails fast if they do not match.

That is the right trade-off. A silent resume with different hyperparameters is worse than a hard error, because it gives you a story about continuity that the math does not deserve.

A close-up of a file drawer holding checkpoints, with a strict gauge comparing two config cards and a red stamp blocking the wrong one. The image turns resumption into a validation problem, which is the repo's main technical point.
Resumption is only trustworthy when the saved config and the live config still agree.
def load_checkpoint(path, current_config):
    checkpoint = torch.load(path, map_location='cpu')
    if checkpoint['config'] != current_config:
        raise ValueError('Refusing to resume with a changed config')
    return checkpoint


def save_state_dict(model):
    return {
        k.replace('_orig_mod.', ''): v
        for k, v in model.state_dict().items()
    }

A checkpoint is not just a file. It is a state machine with a failure path.

Modal is the transport layer, not the story

modal_run.py does the unglamorous infrastructure work. It defines the app, builds a slim image, mounts a Modal volume at /outputs, and uses dynamic imports so one entry point can launch different experiments without turning the repo into a framework.

That separation is the whole point. Compute is disposable. The volume is not. Detached execution keeps the job running in the cloud while the saved artifacts stay available for the next launch, which is why the workflow feels more like choreography than babysitting.

A laptop feeds code across a narrow bridge into a sealed cloud container while the state storage sits below it in a locked vault. The image makes the boundary obvious: code can move with the job, but durable state has to live somewhere else.
The scaffold separates transient compute from persistent training state.

The repo even encodes small operational annoyances instead of pretending they do not exist. Remote image builds need hardcoded ignore lists because the container build cannot lean on a local shell environment, and the TensorBoard helper is defensive enough to degrade gracefully when the dependency is missing. Those are tiny details, but they are the kind that keep a system from feeling fragile.

Why `torch.compile` changes the rules

`torch.compile` is where the template stops being generic and starts looking battle-tested. Compilation wraps the model, which changes state-dict keys and creates a portability problem if you save the wrong representation. The repo handles that by normalizing the wrapper-specific prefix at save time and by loading checkpoints before compilation gets involved.

That order matters. It keeps the artifact portable across compiled and non-compiled runs, and it prevents a very annoying class of bugs where the model looks loaded but is actually carrying the wrong names under the hood.

def restore(model, checkpoint):
    model.load_state_dict(checkpoint['model'])
    return model


def clean_state_dict(state_dict):
    cleaned = {}
    for key, value in state_dict.items():
        cleaned[key.replace('_orig_mod.', '')] = value
    return cleaned

What this template replaces

The easiest way to understand this repo is as a better alternative to three familiar workflows. The first is the plain local PyTorch script that dies when your laptop sleeps. The second is the persistent VM that works but demands manual SSH patience. The third is the generic cloud tutorial that gets you to a demo, not to a trustworthy rerun.

This scaffold sits in the middle. It keeps the deployment surface small, but it adds the missing operational rules that long jobs need in real life.

QuestionPlain local scriptPersistent VM + SSH`tips-for-running-modal`
Does state survive interruption?Only if you add it yourselfUsually, but tied to one machineYes, through a mounted Modal volume
Is resume correctness enforced?Usually noManual and easy to driftYes, config mismatch fails fast
How much deployment friction is there?Low at first, high laterMedium to highLow once the scaffold is in place
Does `torch.compile` portability need care?Easy to breakEasy to overlookHandled by state-dict cleanup
What is the operational burden?You babysit the terminalYou babysit the boxYou poll a detached run and resume cleanly
How likely is silent corruption?HighModerateLower, because the run refuses bad resumes
What is it best for?Short experimentsOne-off long jobsInterruptible training loops

That is why the repo feels more like a survival kit than a tutorial. It does not promise that cloud training will be pleasant. It promises that when the job disappears, the restart will be boring, and boring is exactly what you want from training infrastructure.