NVIDIA/nvcf-otelconfig: Compiling the YAML Wall

How NVIDIA built a multi-stage Python and Go pipeline to generate valid OpenTelemetry configurations for GPU serverless workloads.

6 min read • View on GitHub • More from NVIDIA

A complex vintage telephone switchboard with tangled cords feeding into a mechanical loom that outputs a single pristine cable. This represents the compilation of complex telemetry requirements into a single YAML configuration file.
The 4D configuration matrix requires a compiler to enforce order on chaotic configuration permutations.
Key Takeaways

The YAML Compiler

OpenTelemetry configuration has become so complex that YAML is no longer a human-writable source code. It is a compilation target. The most surprising thing about NVIDIA/nvcf-otelconfig is that NVIDIA had to build a multi-stage compiler just to reliably output a valid config.yaml for GPU serverless workloads.

The repository does not just parse YAML; it compiles it. A suite of Python scripts generates Go templates, which are then embedded into a Go binary via //go:embed. At runtime, this binary ingests a simple JSON request and renders the final OpenTelemetry Collector configuration. Configuration has officially become software engineering.

The compilation pipeline transforms user intent into a valid OpenTelemetry YAML file.

Escaping the 4D Matrix

In the NVIDIA Cloud Functions (NVCF) ecosystem, users Bring Your Own Observability (BYOO). This mandate creates a mathematical nightmare for platform engineers.

NVIDIA Cloud Functions (NVCF) is a platform for deploying, managing, and invoking GPU-accelerated functions at scale.

NVIDIA, Organization · NVIDIA/nvidia-cloud-functions

The system must perfectly map four dimensions: Environment (VM versus Kubernetes), Deployment (Container versus Helm), Protocol (Logs, Metrics, or Traces), and Vendor (Splunk, Datadog, Grafana, or Azure). Hand-writing permutations for this matrix is impossible. The compiler abstraction ensures that the OTel collector knows exactly whether to look for Kubernetes metadata or standard VM metrics based on the deployment target.

Sieving the GPU Noise

Unlike generic OpenTelemetry collectors, this compiled output is highly opinionated for hardware. The embedded Go templates contain expert knowledge specifically tuned for NVIDIA GPUs.

A massive industrial metal sieve suspended in mid-air. Jagged rocks representing noisy container metrics bounce off the mesh, while perfectly smooth glowing spheres representing isolated GPU metrics slip through into a pristine collection basin. This illustrates aggressive metric relabeling.
Aggressive metric_relabel_configs drop irrelevant data to isolate NVIDIA DCGM metrics.

The generated templates aggressively isolate GPU signals. For example, the configurations are hardcoded to scrape nvidia-dcgm-exporter on port 9400. They utilize extensive metric_relabel_configs to drop noisy Kubernetes metrics, strictly filtering for signals like DCGM_FI_DEV_GPU_UTIL. Furthermore, sensitive API keys are never rendered into the YAML string; they are dynamically read from secure mount points using OTel's ${file:...} syntax.

The Declarative Divide

This bespoke, compiled approach stands in stark contrast to the rest of the industry. While the broader OpenTelemetry community is pushing for standardized declarative YAML configurations, NVIDIA's platform requires a dynamic, programmatic generation engine.

ApproachConfiguration MethodFlexibilityUse Case
Standard OTel SDKsProgrammatic (Language-specific)High, but fragmentedApplication-level instrumentation
OTel Declarative ConfigStatic YAMLLanguage-agnostic, rigidStandardized collector deployments
nvcf-otelconfigCompiled (Go/Python engine)Dynamic, highly opinionatedMulti-tenant GPU serverless platforms

Centralized configuration management works for static environments. But when deploying multi-tenant, GPU-accelerated serverless functions at scale, static YAML breaks down. nvcf-otelconfig proves that at a certain scale of complexity, configuration files must graduate into compiled software.