Your Software Says 45 Watts. Your Wall Says Otherwise.
Every homelab has a spreadsheet somewhere with a “power draw” column filled in by vibes. Someone read a forum post, someone else squinted at a UPS display, and now the whole rack is “about 200 watts, I think.”
Here is the thesis. Software power readings lie, and they lie in predictable directions. They almost always read low, because they measure one slice of the machine and skip the rest. The meter at the wall is the truth. Software is what you use for trends once the wall meter has told you how far off the software is.
This article is about measuring, not about what the numbers cost. For watts to kWh to dollars, see the power math post. For UPS setup, see the NUT post. Here we walk the rack box by box, pick the right tool for each, and note which tool is fibbing.
Full example: Clone the working files at github.com/KingPin/sumguy-examples/observability/kill-a-watt-homelab-power-tour
The Toolbox, and How Each One Lies
Five ways to get a number. Each one is wrong in its own way.
| Tool | What it actually measures | How it lies |
|---|---|---|
| Plug-in meter (Kill A Watt) | AC power at the wall | Barely. Limited to one outlet’s rating |
| Smart plug (Shelly, Tasmota) | AC power at the wall | Accuracy depends on the metering chip; low loads read worst |
| UPS via NUT | Load on the UPS output | Reports percent of rated capacity; not every model reports watts |
| BMC via IPMI DCMI | Whole-server power as the BMC sees it | Not every board supports it; sampling windows vary |
RAPL (turbostat, powerstat) | CPU package, optionally DRAM | Skips PSU losses, fans, disks, NICs, GPU |
The Kill A Watt is the reference. The P3 International page for the P4400 model lists 0.2% accuracy and a maximum of 125 VAC, 15 A, and 1875 VA. That is the fine print: it fits a normal 120 V outlet and nothing bigger. Per the manual, it displays watts, kWh, volts, amps, hertz, VA, and power factor.
Use the display’s instantaneous watts to see what a box does right now. Use the cumulative kWh and elapsed-time readout for anything that matters. A spinning-disk NAS, a server with variable fan speed, and a mini PC doing a backup at 3 AM all change by the minute. One glance at the screen tells you almost nothing. Leave the meter in for 24 hours or longer, then divide kWh by hours. That average is the number you want.
Box 1: Modem, Router, and Firewall
Start with the boring stuff, because it never turns off. Anything that runs 8,760 hours a year deserves a real measurement, even if it is small.
Put the Kill A Watt on the router’s power brick and leave it for a day. These devices are flat consumers, so the average and the instantaneous reading land close together. That makes them the easiest calibration check. If your smart plug and your Kill A Watt disagree on a steady 10 W load, you now know the plug’s error at low power, and you can apply it to every other low-power reading you take.
If the firewall is a mini PC running OPNsense or pfSense, treat it as a mini PC (Box 3). Wall power includes the power brick’s own conversion loss, which software cannot see.
Box 2: The PoE Switch and What “PoE Budget” Means
A switch’s PoE budget is a ceiling. It is the most power the switch can hand out to powered devices, and it says nothing about what is flowing right now. A switch with a 150 W budget and two cameras plugged in does not draw 150 W.
So measure the switch twice:
- With no PoE devices connected. This is the switch’s own base draw.
- With all PoE devices connected and running.
The difference is the PoE load plus the conversion loss to deliver it. It will exceed what the cameras and access points consume, because the switch’s PSU and PoE stage are not 100% efficient. Many managed switches also report per-port PoE draw in their web UI or over SNMP. Treat that as a software number: useful for trends, to be checked against the wall once.
Do the measurement with the devices at their busiest. A PTZ camera with infrared LEDs on at night pulls more than the same camera at noon. A one-time reading on a sunny afternoon misses that.
Box 3: The Mini PC
Mini PCs idle low, and that is why people buy them. It is also where software readings tempt you most, because RAPL is right there on any modern Intel box.
RAPL exposes cumulative energy counters, not instantaneous power. The kernel docs say as much: there is no power_uw file. You read energy_uj (microjoules), wait, read again, and divide. The zones live under /sys/class/powercap/, with names like package-0 in each zone’s name file. Since the PLATYPUS fixes (CVE-2020-8694), the counters are readable by root only on current kernels.
# Package power over 10 seconds, no tools neededcd /sys/class/powercap/intel-rapl:0cat namee1=$(sudo cat energy_uj); sleep 10; e2=$(sudo cat energy_uj)echo "scale=2; ($e2 - $e1) / 10 / 1000000" | bcThe counter wraps at max_energy_range_uj. Over 10 seconds on a small box you will not notice. In a long-running logger you must handle it.
For something friendlier, turbostat prints package watts per interval. It must run as root, --Summary gives one line per interval, and --interval sets the period (5 seconds by default):
sudo turbostat --quiet --Summary --interval 10 --show PkgWatt,RAMWattpowerstat does the same job with a nicer summary. The -R flag selects RAPL, and the two positional arguments are the delay between samples and the sample count:
sudo powerstat -R 10 60Here is where RAPL lies. PkgWatt is the CPU package. It does not include the SSD, the NIC, the fan, the Wi-Fi card, the board’s voltage regulators, or the loss in the power brick. The RAMWatt column exists only on server processors, per the turbostat man page. A mini PC that reports a small package number at idle can still show a much bigger number at the wall. The next section measures that gap.
The Procedure: Idle, Load, and the Delta
This is the method for any box. Do it once per machine and write down the results.
- Plug the box into a Kill A Watt (or a smart plug you have checked against one).
- Let it idle for 24 hours or longer. Do not touch it. Note kWh and elapsed hours, then divide.
- Run a sustained load for at least 30 minutes (a CPU stress test, a scrub, a transcode, whatever the box does for a living). Note the instantaneous watts and the kWh over that window.
- During both phases, log the software number (RAPL, DCMI, or UPS watts) at the same time.
- Subtract: wall minus software.
That delta holds more than PSU loss. It includes everything the software number skips: drives, fans, NICs, the motherboard, and any add-in card. You cannot split it into parts without more meters. What you can do is learn how the delta behaves. If it stays roughly constant across idle and load, you have a fixed overhead. If it grows with load, conversion loss and fans are the main contributors. Either way, you now have an offset for that machine, and its software readings become useful trend data.
Box 4: The NAS, Spun Up and Spun Down
A NAS with spinning disks has two personalities. Spun down, the drives draw very little. During spin-up, each drive pulls a surge for several seconds. A one-second sample on the wrong second misses it. A 60-second average hides it.
Measure three states:
- All disks spun down (set the idle timeout, wait, confirm with
hdparm -C /dev/sdX). - All disks spun up and idle.
- Spin-up transient, if you want the peak (a smart plug polled every second does better here than glancing at a meter).
For the averages, only the 24-hour-plus kWh readout is honest. If your disks spin down overnight and spin up for a morning backup, the daily average sits between your two idle numbers, and no single reading will show that.
For disk choices and power per drive, see the DIY NAS build post. The point here is the method: never trust one spot reading on a machine with moving parts.
Box 5: The Used Enterprise Server
Old rack servers are where measuring pays for itself. They idle high, have redundant PSUs that each carry part of the load, and have fans that scale with temperature.
Most of these have a BMC (iDRAC, iLO, or generic IPMI). If the board supports DCMI, you can ask it for power:
ipmitool -I lanplus -H 192.168.1.60 -U admin -P 'changeme' dcmi power readingThe output looks like this (field names from published Supermicro and Dell examples):
Instantaneous power reading: 211 WattsMinimum during sampling period: 1 WattsMaximum during sampling period: 334 WattsAverage power reading over sample period: 214 WattsThe numbers above are just a sample of the format from a public forum post, not a recommendation for what your box should read.
How this lies:
- Not every board supports DCMI. Some return an error or “not supported”. If so, check the vendor’s own tools or the web UI instead.
- The sampling period differs by vendor. Some report a short window, some a long one. Read the “Sampling period” line before you trust the average.
- It is the BMC’s idea of power. Treat it as a software number and compare it to the wall.
With dual PSUs, the wall reading depends on how you plug in. Measure with both PSUs connected, since that is how the server runs. Meter each feed if you can. If you wonder whether one PSU would use less than two, unplug one and measure again. The wall meter answers that in a day.
For BMC access details, see the IPMI, iDRAC, and iLO post.
Box 6: The GPU Box
GPU boxes make fools of idle numbers. The card idles at one level and jumps when a model loads, then drops again between prompts. RAPL sees the CPU only. It will tell you nothing about a graphics card.
Use nvidia-smi for the card’s own reading, and the wall for everything else:
nvidia-smi --query-gpu=power.draw --format=csv,noheader -l 1Log the wall meter alongside. The wall number is the sum of CPU, GPU, board, fans, and PSU loss. The GPU’s self-reported number covers the card only, so the delta between wall and GPU plus RAPL is your overhead for this box. Run a real workload (your actual inference job, not a synthetic benchmark) and keep the meter on for the whole run.
Box 7: The UPS Itself
The UPS is a box in the rack too, and it burns power to keep your other boxes alive. Measure it two ways.
First, plug the UPS into a Kill A Watt with nothing connected to it. That is the UPS’s own overhead. Then connect your gear and measure again. The difference between the second reading and what your gear draws on its own is the UPS’s loss while carrying load.
Second, ask it. With NUT running, upsc reads the UPS variables. It takes one variable per call, so loop:
for v in ups.load ups.realpower ups.realpower.nominal ups.power; do echo "$v: $(upsc myups@localhost $v)"doneIn the NUT variable names list, ups.load is “Load on UPS (percent)”, ups.realpower is “Current value of real power (Watts)”, ups.realpower.nominal is “Nominal value of real power (Watts)”, and ups.power is “Current value of apparent power (Volt-Amps)”.
Here is how this lies. Many consumer UPS models do not report ups.realpower at all, and give you only ups.load. That percentage is a fraction of the rated capacity, and your rating might be in VA while the real capacity is in watts. If you only have a percentage, you can estimate watts as ups.load times the rated watts divided by 100. Check the box’s data sheet for the watt rating, since VA and W differ by the unit’s power factor. Then verify against the Kill A Watt once. After that, your NUT-to-Grafana graph becomes a useful trend line.
Box 8: Vampires, PDUs, and Chargers
Power strips with glowing switches, wall warts with no load attached, a PDU with metering you forgot was on, the phone charger that lives in the rack. Each one is tiny. Together they are a tax you pay forever.
The Kill A Watt is the tool here. Put it on the strip with nothing connected. Many strips with indicator lights show a small but nonzero number. Smart plugs are worse for this job, because their own electronics draw power too, and their accuracy at very low loads is the weakest part of their spec. If a metered PDU reports a number, check it against the Kill A Watt once. The metered PDU’s own electronics draw power, so the reading you want is wall power with the PDU on and nothing plugged in.
Smart Plugs: The Poor Person’s Logger
A Kill A Watt shows you a number only when you stand in front of it. For a month of data, you want a plug that talks to your network. Two common routes are a Shelly plug running its own firmware, or a plug flashed with Tasmota.
The accuracy caveat is real. These plugs use inexpensive metering chips, and they read worst at low loads. Tasmota has built-in calibration commands (PowerCal, VoltageCal, CurrentCal), and the Kill A Watt gives you a reference to calibrate against. Plug the Kill A Watt into the wall, the smart plug into the Kill A Watt, and the load into the smart plug. Compare the readings and adjust. The Kill A Watt also counts the smart plug’s own draw, so calibrate with a load of at least 50 W, where that small offset barely matters.
A Shelly Gen2 or newer plug answers over HTTP RPC. Switch.GetStatus returns apower (instantaneous active power in watts), voltage, current, and an aenergy block whose total is in watt-hours:
curl -s 'http://192.168.1.50/rpc/Switch.GetStatus?id=0'On Tasmota, Status 10 returns sensor data (“replaces Status 8”, per the command reference, which is retained for backwards compatibility). Commands go over HTTP as /cm?cmnd=... with spaces written as %20. The numbers are under StatusSNS.ENERGY:
curl -s 'http://192.168.1.51/cm?cmnd=Status%2010' | jq '.StatusSNS.ENERGY | {Power, Today, Voltage}'A Logger You Can Run Today
This script polls a Shelly Gen2+ plug every N seconds and appends a CSV line of epoch time and watts. It uses only the Python standard library:
#!/usr/bin/env python3"""Poll a Shelly Gen2+ plug and append 'epoch,watts' lines to a CSV."""import jsonimport sysimport timeimport urllib.request
host = sys.argv[1] # e.g. 192.168.1.50interval = float(sys.argv[2]) # seconds between samplesout = sys.argv[3] # e.g. nas.csv
url = f"http://{host}/rpc/Switch.GetStatus?id=0"
with open(out, "a", buffering=1) as f: while True: try: with urllib.request.urlopen(url, timeout=5) as r: watts = json.load(r)["apower"] f.write(f"{int(time.time())},{watts}\n") except Exception as e: print(f"sample failed: {e}", file=sys.stderr) time.sleep(interval)Run it in a tmux session or a systemd unit:
python3 poll_plug.py 192.168.1.50 5 nas.csvNow compute the average watts and kWh per day from the CSV. This awk integrates between samples using the timestamps, so a missed sample does not skew the result:
awk -F, 'NR > 1 { dt = $1 - pt if (dt > 0 && dt < 300) { wh += pw * dt / 3600; secs += dt }}{ pt = $1; pw = $2 }END { if (secs > 0) printf "avg %.1f W over %.1f h, %.3f kWh/day\n", wh / (secs / 3600), secs / 3600, wh / (secs / 3600) * 24 / 1000}' nas.csvThe dt < 300 guard skips gaps longer than five minutes, such as when the logger was down. Without it, a long gap would credit the last known wattage to the whole outage.
Pick a sampling interval that fits the question. For daily averages, 10 to 30 seconds is plenty. For spin-up peaks, go to 1 second and run it for a short window only.
Where Each Tool Belongs
Here is the summary I would pin on the rack:
- Kill A Watt: the truth, for one box at a time. Use it for calibration, idle baselines, and vampire hunting.
- Smart plug: the always-on logger after you have checked it against the Kill A Watt. Good for daily averages, bad for sub-watt loads.
- UPS via NUT: a whole-rack trend line. Verify the watt rating once.
- IPMI DCMI: useful on boards that support it, and always compared to the wall.
- RAPL: a CPU-only trend. It always reads low versus the wall, by a delta you now know.
Your 2 AM self will appreciate a CSV that says exactly what the rack drew last Tuesday.
Common Questions
Is a Kill A Watt accurate enough for homelab measurements?
Yes. P3 International lists 0.2% accuracy for the P4400 model. That is tighter than a typical smart plug or any software estimate, so the Kill A Watt serves as your reference. Its limit is capacity: it tops out at 125 VAC, 15 A, and 1875 VA, so it cannot meter a 240 V server.
Does RAPL show total system power?
No. RAPL reports energy for the CPU package, plus DRAM on server processors, and nothing else. It skips the power supply’s conversion loss, drives, fans, NICs, and add-in cards. Expect wall power to read higher than RAPL on every machine. Measure the gap once with a wall meter.
How long should I measure idle power?
Measure for at least 24 hours. A single reading catches one moment, while a full day catches backups, disk spin-downs, fan changes, and scheduled jobs. Divide the cumulative kWh by elapsed hours for the average. A weekly run is better if your workload has a weekend pattern.
Why does my UPS show load percent but no watts?
Many UPS models report only ups.load and leave out ups.realpower. Percent load is a fraction of the unit’s rated capacity, so multiply it by the rated watts, not VA, to estimate real power. Run upsc on your UPS to see which variables it exposes, then check the result once with a wall meter.