dgx-serve-vllm-gemma4: A private Gemma 4 server that behaves like infrastructure, not a demo

anthonywu's deployment stack turns vLLM, Tailscale, Docker, and systemd into a bootable AI appliance for DGX hardware.

10 min read • View on GitHub • More from anthonywu

A sealed server cabinet sits behind a private network boundary, with one cable arriving from a trusted mesh and another looping back to a local terminal. The image shows that the model server is treated like a private appliance, reachable only through controlled paths.
The repo’s core move is simple: make a large model feel like a machine you can place behind a private boundary, not a public endpoint.
Key Takeaways

Most local LLM stacks start as a command and end as a liability. This repo takes a different path: it wraps Gemma 4 31B in a private, bootable service that only answers on Tailscale or localhost, while the model cache stays offline and the runtime speaks OpenAI-compatible API semantics.

With the debut of Gemma 4, vLLM introduces immediate support for Google's most sophisticated open model lineup, spanning multiple hardware backends, with first-ever Day 0 support on Google TPUs, AMD GPUs, Intel XPUs.

vLLM Team, Project Team · vLLM Gemma 4 blog

The thesis is smaller than it looks

The codebase is compact by design. Compose, just, scripts, and a systemd unit do almost all the work. That makes the repository feel less like an app and more like a deployment pattern: clone it, wire your hardware, and get a reproducible service instead of an ad hoc shell history.

That discipline matters because the target is not generic cloud inference. The repo assumes NVIDIA DGX Spark hardware, a large unified-memory footprint, and a Gemma 4 model that benefits from specific vLLM flags for reasoning, tool calling, and long context.

vLLM is to LLM serving what TensorRT was to DL inference: high performance, low frills, and production first.

David Gao, Blogger / Developer · vLLM deep dive

What makes the stack different

The most important line in the whole project is the port binding. The container is never exposed on 0.0.0.0. Instead, one listener binds to the host's Tailscale IP and another to 127.0.0.1. That is a small detail with a large consequence: the model is reachable from the tailnet and the machine itself, but not from a public ethernet interface.

Typical vLLM rolloutThis repo
Network exposureOften bound to a broad host interface or whatever the container runtime exposesBound to Tailscale and localhost only
Model accessDepends on the surrounding network and manual firewall habitsOffline Hub mode with a mounted Hugging Face cache
StartupA one-off docker command or a manually managed compose filesystemd waits for tailscaled and docker before launching the service
OperationsHumans remember the edge casesScripts and just recipes check privacy and health
Gemma 4 tuningGeneric defaultsReasoning and tool-call parsers, fp8 KV cache, and multimodal limits

The service is not just a container. It is a network boundary, a model runtime, and a boot sequence tied together on purpose.

How the repo keeps the server honest

ports:
  - "${TAILSCALE_IP}:${VLLM_PORT}:8000"
  - "127.0.0.1:${VLLM_LOCAL_PORT}:8000"
environment:
  HF_HUB_OFFLINE: "1"
  VLLM_NO_USAGE_STATS: "1"
  DO_NOT_TRACK: "1"

Those flags are the real contract. They keep the runtime offline, suppress usage telemetry, and make the model server act like a sealed appliance instead of a chat app with opinions. The gemma4-specific parser flags, plus the fp8 cache setting and multimodal limits, push the stack toward a very particular kind of performance: high-capacity serving without pretending the hardware is generic.

The scripts are not garnish. healthcheck.sh and chat-test.sh verify the service end to end, and the fallback path can inspect the container to reach its internal IP if the local port is unavailable. That turns a brittle deployment into something an operator can actually recover.

Purpose-built for advanced reasoning and agentic workflows, Gemma 4 delivers an unprecedented level of intelligence-per-parameter, now accessible to the vLLM community under a commercially permissive Apache 2.0 license.

vLLM Team, Project Team · vLLM Gemma 4 blog

That is why the systemd unit matters. It moves the project from a developer experiment to infrastructure, with startup dependencies that make sense for a machine that should come online only after the secure network and the container engine are ready.

The comparison that matters

A lot of LLM deployment repos optimize for the first successful prompt. This one optimizes for the second month of operation. It cares about private reachability, offline behavior, reproducible startup, and whether the server still feels contained after the novelty wears off.

That makes the audience unusually specific and unusually real: teams handling sensitive data, operators who want a private inference endpoint, and technically capable builders who do not want to rediscover their network model every time the box reboots. For them, the repo is not a toy. It is a template for making a frontier-ish open model feel domesticated.