AI workstation · Ep. 3

First Boot: The Sleep Trap, the Watchdog, and the Health Monitor

The workstation's VRAM bar crossing its 95 percent alert line, the watchdog's green pulse ring, and the phone showing the health-monitor alert that finds the problem first.

The machine arrived. I plugged it in. First boot was clean — Windows 11, the HP Z2 firmware handled the Ultra 7 and the RTX PRO 5000 without drama, nvidia-smi showed driver 595.79 and CUDA 13.2, and the 48 GB of VRAM was visible. I loaded Ollama, pulled Qwen 3.8 27B, ran a test prompt. Tokens came out. All good.

Then the machine went to sleep.

Not “standby” sleep. Not “hibernate” sleep. Something worse. I came back an hour later and the box was cold to the touch, the fans were off, and nvidia-smi timed out. I checked the power settings: sleep was set to “Never.” I checked the Wi-Fi power management: “allow the computer to turn off this device” was unchecked. I rebooted. It worked again. Then it slept again.

What was actually happening

powercfg /a told the truth:

The following sleep states are available on this system:
    Standby (S0 Low Power Idle) Network Connected
    Hibernate
    Fast Startup

This machine uses Modern Standby (S0 Low Power Idle), not the classic S3 “sleep” that the Control Panel sleep slider actually controls. With Modern Standby the CPU enters a low-power state, the Wi-Fi radio (Intel BE200) can power down even when the OS says “stay connected,” and the box looks off but isn’t technically in any state Windows is tracking. The power button is on the front of the tower — I couldn’t reach it without getting up, and the box was behind a desk.

The fix was three layers:

1. Kill the sleep path. I disabled the Wi-Fi power-management “allow the computer to turn off this device” setting and added a powercfg registry override so the S0 state never engages when the machine is plugged into power. The command:

powercfg /setdcvalueindex SCHEME_CURRENT SUB_SLEEP STANDBYIDLE 0
powercfg /setacvalueindex SCHEME_CURRENT SUB_SLEEP STANDBYIDLE 0
powercfg /setactive SCHEME_CURRENT

2. Keep the network alive. Even with S0 disabled, the Wi-Fi driver was still dropping the link under load. I set the adapter’s power management to “Never” and disabled the Intel BE200’s “Power Save Mode.” The Tailscale tunnel now stays up.

3. Wake on LAN — the part that took the longest. The machine had no Ethernet cable; only Wi-Fi was up. HP Z2 firmware has a “Wake on Ring” option but no standard WoL for Wi-Fi-only setups. The working fix: a scheduled task with “Wake the computer to run this task” enabled, triggered every 5 minutes during the hours I expect the machine to be on. Combined with the Tailscale tunnel, I can now wake it from my phone or laptop without touching the box.

The watchdog

Even after the sleep was fixed, services died. Ollama would crash after a bad model load. ComfyUI would OOM on a large image. The Gateway would hang after a long-running inference. Task Scheduler’s “restart on failure” only fires for processes it saw fail — a process that segfaults silently or gets killed by OOM doesn’t trigger it.

So I wrote C:\Users\admin\watchdog\watchdog.ps1. It runs every minute as a scheduled task. It checks 8 services:

Task Port Kill list
OllamaServe 11434 ollama.exe, llama-server.exe
Gateway 5050 (path match)
PrivateAIApp 3939 (path match)
PublicAIApp 4040 (path match)
PublicGujAIApp 4041 (path match)
ComfyUI 8188 (path match)
PdfWerkApi 5272 —
CloudflaredTunnel — cloudflared.exe

Each check is: is the port listening (or the process running)? If not, increment a counter. Second miss → kill any leftover processes and Start-ScheduledTask. Seventh consecutive miss → give up and log “Check its log.” This prevents a crash loop from hammering the box.

The log caps at 1 MB and keeps the last 2000 lines.

The health monitor

The watchdog keeps services alive. It doesn’t tell me anything.

C:\Users\admin\health-monitor\monitor.js is the one that emails me. It runs as a separate scheduled task (HealthMonitor), and does:

  1. Port check on the same 5 ports (11434, 8188, 5050, 3939, 4040)
  2. URL check on https://chat.jambuvan.app and https://private-ai.jambuvan.app (catches the cloudflared tunnel being down)
  3. Process check for cloudflared.exe
  4. GPU VRAM check via nvidia-smi — alerts if >95 % used
  5. Disk space check — alerts if <50 GB free on C:

If anything fails, it sends an email via Brevo (sender: alerts@jambuvan.app, to: joshihrn.it@gmail.com) with the list of failures and a timestamp. There’s a 1-hour cooldown so a flapping service doesn’t spam me, and a recovery email when things come back.

The Brevo API key is read from C:\Users\admin\private-app\data\brevo_key.txt at runtime. No secrets in the script.

What I got out of it

  • The machine doesn’t go to sleep anymore.
  • If a service dies, it’s back within 2 minutes without me knowing.
  • If a service can’t come back, I get an email on my phone.
  • If the GPU is about to OOM, I know before the inference fails.
  • If the disk is getting full, I know before the models can’t write.

The setup took about a day of trial and error. Most of it was the sleep trap — the other stuff was straightforward. The sleep trap is the part I’d tell anyone setting up a Modern Standby Windows box in a room they don’t sit near: check powercfg /a first, and if it says S0 Low Power Idle, the sleep settings in the UI are a lie.

What’s next

The box is awake, watched, and loud enough to email me when it’s not. Next is putting the actual work on it: the first local inference runs, and wiring the services behind episode 2’s “12 models in the local library” into one endpoint Jambu can call. That is episode 4: first local inference.