All posts

A NUT Server in an LXC, and a Grafana Dashboard to Watch It

Running Network UPS Tools inside a Proxmox container, passing the USB cable through, and turning battery runtime into a graph I actually look at.

The UPS in my rack sat there for two years doing exactly one job: keeping things alive through the 30-second brownouts my neighborhood specializes in. It did that job silently and invisibly, which sounds great until you realize you have no idea how old the battery is, how much runtime you actually have, or whether anything will shut down cleanly when the runtime runs out.

So: NUT. Network UPS Tools has been the answer to this for about twenty years, it’s in every distro’s repos, and its documentation reads like it was written for people who already know how it works. Here’s the version I wish I’d read first.

The mental model

NUT splits into three pieces, and understanding the split makes the config files stop looking arbitrary:

  • The driver (usbhid-ups for most USB units) talks to the physical UPS and translates its HID reports into NUT variables.
  • upsd is a small network daemon that serves those variables to anything that asks, over TCP 3493.
  • upsmon is a client. It connects to upsd, watches for OB LB (on battery, low battery), and runs a shutdown command when things go bad.

The important consequence: upsmon runs on every machine you want to shut down, not just the one with the USB cable in it. One primary, many secondaries. That’s the whole design.

Why an LXC, and the one catch

I run everything on Proxmox, so a 256MB unprivileged container was the obvious home for this. The catch is that the UPS is a USB HID device, and unprivileged containers map root to an unprivileged host UID — so the /dev/hidraw* node shows up inside the container owned by nobody:nogroup and NUT can’t open it.

You have two options. The clean one is a udev rule on the host that chowns the device to the mapped UID; the fast one is running this single container privileged. I went privileged. It does one thing, it has no network exposure beyond 3493 on the lab VLAN, and I’d rather spend the complexity budget elsewhere. Know that you’re making that trade.

Find the device on the host first:

lsusb                    # note the Bus/Device numbers
ls -l /dev/hidraw*       # note the major number (243 on my host, it varies)

Then add to /etc/pve/lxc/<CTID>.conf:

lxc.cgroup2.devices.allow: c 189:* rwm
lxc.cgroup2.devices.allow: c 243:* rwm
lxc.mount.entry: /dev/bus/usb/003 dev/bus/usb/003 none bind,optional,create=dir
lxc.mount.entry: /dev/hidraw0 dev/hidraw0 none bind,optional,create=file

189 is USB device nodes, 243 is hidraw on my kernel — check yours. The bind mount of the whole bus directory matters: if the UPS renumbers after a power event, a bind of the single device node goes stale and the driver comes back with “no appropriate HID device found” after the next outage, which is a delightful thing to discover during an outage.

The four config files

Inside the container, apt install nut-server nut-client, then:

/etc/nut/ups.conf — describe the hardware.

[rack]
    driver = usbhid-ups
    port = auto
    desc = "Rack UPS"
    pollinterval = 5

/etc/nut/upsd.conf — listen on the lab network, not just loopback, or the other nodes can’t reach it.

LISTEN 0.0.0.0 3493

/etc/nut/upsd.users — one account for local monitoring, one read-only account for metrics.

[upsmon-primary]
    password = <redacted>
    upsmon primary

[metrics]
    password = <redacted>
    actions = SET
    instcmds = none

/etc/nut/upsmon.conf — the part that actually does something when the power goes out.

MONITOR rack@localhost 1 upsmon-primary <redacted> primary
MINSUPPLIES 1
SHUTDOWNCMD "/sbin/shutdown -h +0"
POWERDOWNFLAG /etc/killpower

Finally set MODE=netserver in /etc/nut/nut.conf, restart both services, and confirm:

upsc rack@localhost

You should get a wall of variables — battery.charge, ups.load, input.voltage, battery.runtime. If you get “Driver not connected,” it’s the device passthrough, ninety percent of the time.

Making it a graph

upsc in a terminal is fine for a one-off. For history you want the numbers in a TSDB. I used nut_exporter, which speaks NUT on one side and Prometheus on the other, running in the same container:

--nut.server=127.0.0.1
--nut.username=metrics
--nut.vars_enable=battery.charge,battery.runtime,ups.load,input.voltage,ups.status

Scrape it on 9199, then point Grafana at it. There are community NUT dashboards on grafana.com worth starting from, but the four panels I actually look at are:

PanelMetricWhy
Battery chargebattery.chargeThe obvious one
Runtime remainingbattery.runtime / 60In minutes, because seconds are useless to a human
Loadups.loadTells you if you’ve quietly outgrown the unit
Input voltageinput.voltageThe interesting one

That last one is the reason to do any of this. Charge and load are boring — they sit at 100% and 34% forever. Input voltage is a live readout of your utility feed, and mine is bad: regular dips into the low 100s, a handful of sub-90V sags a week, each one a few hundred milliseconds. None of them long enough to notice. All of them being absorbed by a battery that’s slowly earning its retirement.

I set one alert — runtime below 5 minutes — and one recording rule counting transfer events per week. The trend line on that second one is what’ll tell me the battery is dying, months before it fails a self-test.

Wiring up the rest of the cluster

Every other Proxmox node gets nut-client only, with a single line in its upsmon.conf:

MONITOR rack@10.0.20.15 1 metrics <redacted> secondary

Secondaries shut themselves down and then tell the primary they’re done. The primary waits for FINALDELAY, then goes.

There’s a wrinkle here I’m saving for next week’s post: in a Proxmox HA cluster, nodes disappearing one at a time looks exactly like nodes failing one at a time. Shut down two of three nodes and the survivor loses quorum, decides it’s the broken one, and gets rebooted by its own watchdog — while on battery. Getting the ordering right is its own problem, and it’s the reason I ended up clustering properly instead of the half-cluster I’d been running.

Was it worth it

A weekend, mostly spent on device passthrough. What I got: a number for how much runtime I actually have (11 minutes at current load, not the 25 the box claims), automatic clean shutdown, and a graph that has quietly convinced me my utility power is worse than I thought.

The graph is the part I didn’t expect to care about.