`tips-for-running-modal`: Training code built to survive being killed
A compact Modal scaffold for PyTorch that treats checkpoints, preemption, and `torch.compile` edge cases as first-class problems instead of afterthoughts.
- The repo treats interruption as a normal condition, so restartability matters more than keeping one process alive.
- Checkpointing is framed as a contract, which means a saved run is only valid if the current config still matches it.
- Modal is the transport layer, not the point, because the real design work sits in durable state and failure handling.
- `torch.compile` changes checkpoint portability, and the scaffold handles that by normalizing wrapper-specific state dict keys.
Most ML templates optimize for the happy path. This one optimizes for the moment the run gets preempted, the container disappears, and the only useful thing left is a checkpoint on a volume. That shift changes everything: where state lives, when a job is allowed to resume, and what counts as correctness.
A training job that expects to die
The repository reads like a small operating manual for serverless GPU work. It assumes that cloud training is interruptible, that detached runs are normal, and that the filesystem attached to compute is not where truth lives.
That is why the scaffold is so compact. The repo is not trying to become a framework. It is trying to encode a reliable habit: train somewhere ephemeral, write progress somewhere durable, and make resume a first-class path instead of a lucky accident.
The documentation folders matter here too. Files like AGENTS.md and CLAUDE.md signal that the repository expects to be reused by humans and assistants without losing the operating rules. This is less a library than a copyable contract.
The checkpoint contract
The sharpest idea in the repo is not saving state. It is refusing to resume under the wrong state. The checkpoint loader scans for the latest saved file, compares the saved config with the current config, and fails fast if they do not match.
That is the right trade-off. A silent resume with different hyperparameters is worse than a hard error, because it gives you a story about continuity that the math does not deserve.
def load_checkpoint(path, current_config):
checkpoint = torch.load(path, map_location='cpu')
if checkpoint['config'] != current_config:
raise ValueError('Refusing to resume with a changed config')
return checkpoint
def save_state_dict(model):
return {
k.replace('_orig_mod.', ''): v
for k, v in model.state_dict().items()
}
Modal is the transport layer, not the story
modal_run.py does the unglamorous infrastructure work. It defines the app, builds a slim image, mounts a Modal volume at /outputs, and uses dynamic imports so one entry point can launch different experiments without turning the repo into a framework.
That separation is the whole point. Compute is disposable. The volume is not. Detached execution keeps the job running in the cloud while the saved artifacts stay available for the next launch, which is why the workflow feels more like choreography than babysitting.
The repo even encodes small operational annoyances instead of pretending they do not exist. Remote image builds need hardcoded ignore lists because the container build cannot lean on a local shell environment, and the TensorBoard helper is defensive enough to degrade gracefully when the dependency is missing. Those are tiny details, but they are the kind that keep a system from feeling fragile.
Why `torch.compile` changes the rules
`torch.compile` is where the template stops being generic and starts looking battle-tested. Compilation wraps the model, which changes state-dict keys and creates a portability problem if you save the wrong representation. The repo handles that by normalizing the wrapper-specific prefix at save time and by loading checkpoints before compilation gets involved.
That order matters. It keeps the artifact portable across compiled and non-compiled runs, and it prevents a very annoying class of bugs where the model looks loaded but is actually carrying the wrong names under the hood.
def restore(model, checkpoint):
model.load_state_dict(checkpoint['model'])
return model
def clean_state_dict(state_dict):
cleaned = {}
for key, value in state_dict.items():
cleaned[key.replace('_orig_mod.', '')] = value
return cleaned
What this template replaces
The easiest way to understand this repo is as a better alternative to three familiar workflows. The first is the plain local PyTorch script that dies when your laptop sleeps. The second is the persistent VM that works but demands manual SSH patience. The third is the generic cloud tutorial that gets you to a demo, not to a trustworthy rerun.
This scaffold sits in the middle. It keeps the deployment surface small, but it adds the missing operational rules that long jobs need in real life.
| Question | Plain local script | Persistent VM + SSH | `tips-for-running-modal` |
|---|---|---|---|
| Does state survive interruption? | Only if you add it yourself | Usually, but tied to one machine | Yes, through a mounted Modal volume |
| Is resume correctness enforced? | Usually no | Manual and easy to drift | Yes, config mismatch fails fast |
| How much deployment friction is there? | Low at first, high later | Medium to high | Low once the scaffold is in place |
| Does `torch.compile` portability need care? | Easy to break | Easy to overlook | Handled by state-dict cleanup |
| What is the operational burden? | You babysit the terminal | You babysit the box | You poll a detached run and resume cleanly |
| How likely is silent corruption? | High | Moderate | Lower, because the run refuses bad resumes |
| What is it best for? | Short experiments | One-off long jobs | Interruptible training loops |
That is why the repo feels more like a survival kit than a tutorial. It does not promise that cloud training will be pleasant. It promises that when the job disappears, the restart will be boring, and boring is exactly what you want from training infrastructure.