Artemis2: The Moon Mission Benchmark That Grades Reasoning, Not Answers

Inside a Flight Director simulator where telemetry checks, anomaly handling, and burn decisions are scored by deterministic physics.

11 min read • View on GitHub • More from EnvCommons

A wide black-ink editorial illustration of a mission control room on a pure white background, with one operator at a console, curved telemetry ribbons, and a lunar trajectory arc overhead. It explains that Artemis2 turns evaluation into procedural work, not chat.
Artemis2 frames the agent as mission control, where every action is a procedure with consequences.
Key Takeaways

The benchmark is a Flight Director job, not a chat task

Artemis2 is not trying to make an agent play pilot. It asks the model to act like Flight Director, which means reading telemetry, checking subsystems, handling anomalies, and making go or no-go calls over a long mission.

That changes the whole evaluation game. The agent is not rewarded for a fluent answer, it is rewarded for the sequence of actions that proves it can keep state, follow procedure, and recover when the mission gets messy.

Artemis II Description Artemis2 is a multi-step decision environment simulating NASA's Artemis II crewed lunar flyby mission.

Project README, Repository documentation · EnvCommons/Artemis2 README

All spacecraft parameters are grounded in publicly available NASA mission data for the real Artemis II mission (SLS Block 1 / Orion / European Service Module).

Project README, Repository documentation · EnvCommons/Artemis2 README

Why the scoring model is the real invention

The clever part is not the space theme. It is that Artemis2 can grade process with a much tighter signal than a rubric-based judge ever could. The environment cares whether the agent checked the right system, asked the right console, and responded in the right order.

That makes the benchmark feel more like operations than conversation. A good run is one where the model keeps the mission moving without losing track of fuel, power, oxygen, attitude, and anomaly status.

Artemis2LLM-as-judge
Scores the procedure, not just the final answer.Scores a response after the fact.
Uses deterministic mission state and seeded anomalies.Depends on rubric interpretation and judge consistency.
Exposes state drift, missed checks, and bad recovery steps.Exposes style issues and shallow reasoning.
Rewards the right sequence of actions under pressure.Rewards plausible-sounding completion.
Best for long-horizon operational reasoning.Best for open-ended text quality or ranking.

The diagram shows Artemis2 as a chain of verifiable decisions, where each phase changes the available telemetry and the score.

How a mission phase actually unfolds

A phase begins with state, not drama. The agent sees the current mission step, checks the relevant telemetry, and then decides whether to hold, call a subsystem, or execute a maneuver.

The key detail is that the simulator can inject anomalies on a seeded schedule. That means the benchmark can force the model to diagnose a partial failure instead of coasting through a happy path.

The repo's console model makes the job feel real. GNC is not just a label, PROP is not just a prop word, and EECOM is not decorative. Each console maps to a slice of mission truth, so the agent has to ask the right place before it can claim the mission is safe.

  1. Read the current telemetry and mission phase.
  2. Query the console that owns the subsystem in question.
  3. Handle anomalies before committing to a burn or a go call.
  4. Advance the mission only after the simulator validates the decision.
A close-up black-ink illustration of a go-no-go console on a pure white background, with subsystem cards, checkmarks, and one anomaly card half-lit. It explains how Artemis2 scores the order and quality of checks, not just the final success.
The benchmark rewards the right checks in the right sequence.

Under the hood, the repo is built for determinism

The codebase is split the way a serious benchmark should be split. physics.py runs the discrete-event simulation, mission_data.py holds the mission truth, and artemis2.py wraps that state in tools an agent can actually use.

The physics layer avoids heavy orbital theatrics and goes for reproducible behavior. Trajectories are interpolated between waypoints, burns use standard rocket equation logic, and the whole system stays deterministic enough for golden tests to catch regressions.

That choice matters because the benchmark has to be inspectable. If a run fails, you want to know whether the agent missed a check, whether an anomaly was mishandled, or whether the simulation itself drifted, and this repo is set up to tell you.

A black-ink illustration of a spacecraft state card floating above a slim mission timeline on a pure white background. Arrows connect the card to fuel, power, oxygen, attitude, and anomaly status, explaining how the repository keeps mission state deterministic and inspectable.
The simulation keeps state legible so every reward can be traced back to the mission model.

Why Artemis2 beats LLM-as-judge for this problem

Artemis2 is narrower than a free-form benchmark, and that is the point. It gives up open-ended judgment so it can give back a cleaner signal on long-horizon reasoning, state retention, and procedural recovery.

If you are building agents that must operate over time, that trade is hard to beat. The project is saying that the best test is not whether the model can write a nice explanation, but whether it can keep the room under control when the mission changes shape.

Who built it, and what the repo says about the ecosystem

This reads like a benchmark built by people who care about evaluation infrastructure more than demo polish. The footprint is small, but the separation of concerns, validation layer, and test coverage all point in the same direction: this is research-grade tooling with a strong bias toward reproducibility.

It also fits the broader OpenReward idea. Artemis2 is not a product trying to capture attention, it is a proving ground for agents that need dense feedback, hard state, and a clear mission objective.