Artemis2: The Moon Mission Benchmark That Grades Reasoning, Not Answers
Inside a Flight Director simulator where telemetry checks, anomaly handling, and burn decisions are scored by deterministic physics.
- Artemis2 grades whether an agent can run a mission procedure end to end, not whether it can sound smart.
- Its edge is deterministic scoring, which turns telemetry checks, anomaly handling, and burn decisions into measurable behavior.
- The repo keeps physics, mission data, and agent tools separate so the simulation stays reproducible and inspectable.
- Compared with LLM-as-judge benchmarks, Artemis2 trades open-endedness for a much sharper signal on long-horizon state and recovery.
The benchmark is a Flight Director job, not a chat task
Artemis2 is not trying to make an agent play pilot. It asks the model to act like Flight Director, which means reading telemetry, checking subsystems, handling anomalies, and making go or no-go calls over a long mission.
That changes the whole evaluation game. The agent is not rewarded for a fluent answer, it is rewarded for the sequence of actions that proves it can keep state, follow procedure, and recover when the mission gets messy.
Artemis II Description Artemis2 is a multi-step decision environment simulating NASA's Artemis II crewed lunar flyby mission.
All spacecraft parameters are grounded in publicly available NASA mission data for the real Artemis II mission (SLS Block 1 / Orion / European Service Module).
Why the scoring model is the real invention
The clever part is not the space theme. It is that Artemis2 can grade process with a much tighter signal than a rubric-based judge ever could. The environment cares whether the agent checked the right system, asked the right console, and responded in the right order.
That makes the benchmark feel more like operations than conversation. A good run is one where the model keeps the mission moving without losing track of fuel, power, oxygen, attitude, and anomaly status.
| Artemis2 | LLM-as-judge |
|---|---|
| Scores the procedure, not just the final answer. | Scores a response after the fact. |
| Uses deterministic mission state and seeded anomalies. | Depends on rubric interpretation and judge consistency. |
| Exposes state drift, missed checks, and bad recovery steps. | Exposes style issues and shallow reasoning. |
| Rewards the right sequence of actions under pressure. | Rewards plausible-sounding completion. |
| Best for long-horizon operational reasoning. | Best for open-ended text quality or ranking. |
How a mission phase actually unfolds
A phase begins with state, not drama. The agent sees the current mission step, checks the relevant telemetry, and then decides whether to hold, call a subsystem, or execute a maneuver.
The key detail is that the simulator can inject anomalies on a seeded schedule. That means the benchmark can force the model to diagnose a partial failure instead of coasting through a happy path.
The repo's console model makes the job feel real. GNC is not just a label, PROP is not just a prop word, and EECOM is not decorative. Each console maps to a slice of mission truth, so the agent has to ask the right place before it can claim the mission is safe.
- Read the current telemetry and mission phase.
- Query the console that owns the subsystem in question.
- Handle anomalies before committing to a burn or a go call.
- Advance the mission only after the simulator validates the decision.
Under the hood, the repo is built for determinism
The codebase is split the way a serious benchmark should be split. physics.py runs the discrete-event simulation, mission_data.py holds the mission truth, and artemis2.py wraps that state in tools an agent can actually use.
The physics layer avoids heavy orbital theatrics and goes for reproducible behavior. Trajectories are interpolated between waypoints, burns use standard rocket equation logic, and the whole system stays deterministic enough for golden tests to catch regressions.
That choice matters because the benchmark has to be inspectable. If a run fails, you want to know whether the agent missed a check, whether an anomaly was mishandled, or whether the simulation itself drifted, and this repo is set up to tell you.
Why Artemis2 beats LLM-as-judge for this problem
Artemis2 is narrower than a free-form benchmark, and that is the point. It gives up open-ended judgment so it can give back a cleaner signal on long-horizon reasoning, state retention, and procedural recovery.
If you are building agents that must operate over time, that trade is hard to beat. The project is saying that the best test is not whether the model can write a nice explanation, but whether it can keep the room under control when the mission changes shape.
Who built it, and what the repo says about the ecosystem
This reads like a benchmark built by people who care about evaluation infrastructure more than demo polish. The footprint is small, but the separation of concerns, validation layer, and test coverage all point in the same direction: this is research-grade tooling with a strong bias toward reproducibility.
It also fits the broader OpenReward idea. Artemis2 is not a product trying to capture attention, it is a proving ground for agents that need dense feedback, hard state, and a clear mission objective.