NVIDIA/x86-power-control: The daemon that turns server power into a state machine
Inside the BMC logic that sequences power rails, watches for failure states, and adapts to NVIDIA hardware through configuration instead of hardcoded assumptions.
- This repo turns power control into a guarded state machine, because server power is only safe when every transition waits for the right hardware signal or timeout.
- NVIDIA's real portability layer is configuration, not code, since JSON signal maps let one daemon adapt to different boards without rewriting the sequencing logic.
- Persistence matters because the BMC can reboot while the host state still matters, so the daemon has to remember intent across restarts.
- The project is a strong OpenBMC story because it keeps the shared control model while pushing vendor-specific quirks into small, testable overrides.
The invisible machinery behind “power on”
The repo’s public description says it plainly: NVIDIA calls this “a BMC based power control for Nvidia CPUs.” That sounds modest until you unpack what “power control” means on real hardware. It is the logic that decides when a server can safely wake up, when it should shut down, and what to do when the host, the BMC, and the wiring do not agree.
x86-power-control is a BMC application that manages the power state of an x86 host. It supports power on, power off, reset, and graceful shutdown.
That is the right frame for this repository too. The interesting thing is not that a daemon can assert a pin. It is that the daemon knows which pin matters on this platform, when it matters, and how long it can wait before deciding the machine is stuck.
Why this repo matters
Server power is brittle in the boring way that hurts production. The command is simple. The reality is not. NVIDIA’s implementation sits inside OpenBMC’s control plane and translates a generic request into board-specific hardware actions, with the right polarity, timing, and fallback behavior for each platform.
That is why the project reads less like a utility and more like infrastructure. It is the difference between a script that toggles a line and a service that can survive imperfect hardware, missed interrupts, and BMC restarts.
A state machine built for hardware that can disagree with you
The heart of the code lives in a state machine. The daemon does not assume one clean power-on path. It reacts to events, checks feedback, and moves through guarded states until the host is stable or the attempt is abandoned.
enum class PowerState {
waitForPowerOK,
transitionToOff,
checkForWarmReset,
softOff,
hardOff,
powerOn,
reset,
};
// The real work is in what happens when inputs, timeouts,
// and hardware feedback disagree with each other.
That pattern matters because the daemon has to deal with more than one kind of signal source. Some inputs arrive as direct GPIO interrupts. Others come through polled paths. Some are D-Bus properties. The code is built to handle all of that without collapsing the whole machine into one fragile assumption.
Configuration is the real portability layer
The most important portability trick in the repo is not inheritance. It is JSON. NVIDIA maps logical signals to physical lines in configuration, which means a new board can change wiring, polarity, and timing without asking the daemon to relearn the universe.
| Layer | Hardcoded approach | NVIDIA approach |
|---|---|---|
| Signal mapping | Wires and offsets baked into logic | Logical names mapped through JSON |
| Polarity | Assumed or scattered across code | Declared per signal in config |
| Timing | Fixed in implementation | Parameterized per platform |
| Portability | New board means new code path | New board means new config |
| Failure handling | Often ad hoc | State machine plus watchdog-style limits |
This is why the project feels more durable than a pile of board scripts. The daemon becomes reusable because the board-specific facts are moved out of the code path and into data.
When the BMC remembers what happened before it rebooted
Power state is not ephemeral. If the BMC restarts, the host may still be on, or it may have been in the middle of a shutdown sequence. The repo’s persistent state support exists so the daemon can remember what mattered before the reboot and continue from there instead of guessing.
That matters for last-state policies. If the system was powered on before the BMC lost power, the control plane needs to restore the right behavior when it comes back. This is one of those details that only looks boring until it fails.
This application provides a D-Bus interface for power control of x86-based hosts.
The NVIDIA-specific layer
The `src/platform/nvl144/` code is where the generic model bends to fit NVIDIA hardware. That is the point of a platform layer here: not to reinvent power control, but to encode the quirks that make one board different from another.
In practice, that means the same daemon can keep the high-level sequencing logic while a smaller platform subclass handles the parts that should not be universal. It is a clean split between control philosophy and hardware reality.
| Question | Generic OpenBMC layer | NVIDIA platform layer | |
|---|---|---|---|
| What stays shared? | State machine, event handling, D-Bus interface | Still shared | |
| What changes? | Platform assumptions | NVIDIA sequencing rules and quirks | |
| Where do board facts live? | Common code and defaults | Platform overrides and config | |
| Why split it? | To support many x86 systems | To support NVIDIA hardware without forking the core logic |
Why the details matter: polling, expanders, and watchdogs
The uncomfortable part of power control is that the obvious path is not always available. The research for this repo points to polled input paths such as `gpio_keys_polled`, which means some signals are not handled as neat interrupts. They have to be watched the hard way because the hardware does not hand the BMC a cleaner interface.
The same is true for I2C-attached GPIO expanders. When a signal lives behind an expander instead of a direct SoC pin, the daemon has to treat the board as a network of dependencies, not a flat list of wires. Timing limits and watchdog-like behavior keep the host from hanging forever in a half-finished shutdown.
// Typical control problems this daemon has to solve:
// 1. Wait for the right signal.
// 2. Stop waiting when the signal never comes.
// 3. Fall back to a safer transition path.
// 4. Keep the host from getting stuck in limbo.
That is the real engineering burden here. Power control is reliability work masquerading as plumbing.
How it compares to upstream and alternatives
NVIDIA is not inventing a new category. It is narrowing a known one. The project sits close to upstream OpenBMC power control, but shifts the burden of variation toward configuration and platform overrides, which is exactly what you want when the same control model has to survive different server families.
| Project | Scope | Strength | Trade-off |
|---|---|---|---|
| NVIDIA/x86-power-control | NVIDIA x86 platforms | Vendor-tuned behavior with shared OpenBMC logic | Narrower scope by design |
| openbmc/x86-power-control | General OpenBMC x86 hosts | Broad upstream baseline | Less platform-specific |
| Intel OpenBMC power-control | Intel-centered deployments | Tight fit for Intel ecosystems | Optimized for a different hardware family |
| Commercial BMC stacks like MegaRAC | Proprietary server firmware | Integrated product bundle | Less inspectable and less adaptable |
The better comparison is not feature checklists. It is control philosophy. NVIDIA keeps the OpenBMC model, but pushes the messy hardware facts into the smallest possible surface area.
Why this is a good OpenBMC story
This repository is a good example of why OpenBMC matters. Shared infrastructure only works if it can tolerate hardware diversity without turning into a pile of exceptions. Here, the daemon stays small in concept and broad in consequence.
That is the hidden win. The code makes power boring enough to trust, even when the wiring, platform, and timing details are anything but boring.