dgx-serve-vllm-gemma4: A private Gemma 4 server that behaves like infrastructure, not a demo
anthonywu's deployment stack turns vLLM, Tailscale, Docker, and systemd into a bootable AI appliance for DGX hardware.
- This repo turns Gemma 4 into a private appliance by binding the server to Tailscale and localhost instead of exposing a generic container port.
- Its real product is the deployment envelope, where offline caches, privacy flags, and systemd startup turn a model server into something you can trust to stay put.
- Gemma 4 specific vLLM flags, including reasoning, tool calling, fp8 cache settings, and multimodal limits, make the stack feel tuned rather than templated.
- The repository is a reference architecture for operators who want an OpenAI-compatible endpoint that behaves like infrastructure, not a demo.
Most local LLM stacks start as a command and end as a liability. This repo takes a different path: it wraps Gemma 4 31B in a private, bootable service that only answers on Tailscale or localhost, while the model cache stays offline and the runtime speaks OpenAI-compatible API semantics.
With the debut of Gemma 4, vLLM introduces immediate support for Google's most sophisticated open model lineup, spanning multiple hardware backends, with first-ever Day 0 support on Google TPUs, AMD GPUs, Intel XPUs.
The thesis is smaller than it looks
The codebase is compact by design. Compose, just, scripts, and a systemd unit do almost all the work. That makes the repository feel less like an app and more like a deployment pattern: clone it, wire your hardware, and get a reproducible service instead of an ad hoc shell history.
That discipline matters because the target is not generic cloud inference. The repo assumes NVIDIA DGX Spark hardware, a large unified-memory footprint, and a Gemma 4 model that benefits from specific vLLM flags for reasoning, tool calling, and long context.
vLLM is to LLM serving what TensorRT was to DL inference: high performance, low frills, and production first.
What makes the stack different
The most important line in the whole project is the port binding. The container is never exposed on 0.0.0.0. Instead, one listener binds to the host's Tailscale IP and another to 127.0.0.1. That is a small detail with a large consequence: the model is reachable from the tailnet and the machine itself, but not from a public ethernet interface.
| Typical vLLM rollout | This repo | |
|---|---|---|
| Network exposure | Often bound to a broad host interface or whatever the container runtime exposes | Bound to Tailscale and localhost only |
| Model access | Depends on the surrounding network and manual firewall habits | Offline Hub mode with a mounted Hugging Face cache |
| Startup | A one-off docker command or a manually managed compose file | systemd waits for tailscaled and docker before launching the service |
| Operations | Humans remember the edge cases | Scripts and just recipes check privacy and health |
| Gemma 4 tuning | Generic defaults | Reasoning and tool-call parsers, fp8 KV cache, and multimodal limits |
How the repo keeps the server honest
ports:
- "${TAILSCALE_IP}:${VLLM_PORT}:8000"
- "127.0.0.1:${VLLM_LOCAL_PORT}:8000"
environment:
HF_HUB_OFFLINE: "1"
VLLM_NO_USAGE_STATS: "1"
DO_NOT_TRACK: "1"
Those flags are the real contract. They keep the runtime offline, suppress usage telemetry, and make the model server act like a sealed appliance instead of a chat app with opinions. The gemma4-specific parser flags, plus the fp8 cache setting and multimodal limits, push the stack toward a very particular kind of performance: high-capacity serving without pretending the hardware is generic.
The scripts are not garnish. healthcheck.sh and chat-test.sh verify the service end to end, and the fallback path can inspect the container to reach its internal IP if the local port is unavailable. That turns a brittle deployment into something an operator can actually recover.
Purpose-built for advanced reasoning and agentic workflows, Gemma 4 delivers an unprecedented level of intelligence-per-parameter, now accessible to the vLLM community under a commercially permissive Apache 2.0 license.
That is why the systemd unit matters. It moves the project from a developer experiment to infrastructure, with startup dependencies that make sense for a machine that should come online only after the secure network and the container engine are ready.
The comparison that matters
A lot of LLM deployment repos optimize for the first successful prompt. This one optimizes for the second month of operation. It cares about private reachability, offline behavior, reproducible startup, and whether the server still feels contained after the novelty wears off.
That makes the audience unusually specific and unusually real: teams handling sensitive data, operators who want a private inference endpoint, and technically capable builders who do not want to rediscover their network model every time the box reboots. For them, the repo is not a toy. It is a template for making a frontier-ish open model feel domesticated.