How it fits together

XIAO ESP32-S3  --ESP-NOW-->  ESP32 master  --UART-->  ESP32 bridge  --HTTP-->  Controller  --JSON API-->  Blue Iris
(battery, asleep)            (no WiFi)                (WiFi + OTA)             (Python)                   (9 cameras)

A press travels four hops. Two acknowledgements come back the other way to the button's LED: one when the master hears the press (a few milliseconds), and one when Blue Iris has been told to save (about a second).

Most of the interesting problems were in that chain, so that's what this write-up covers first.

The button

The button is a Seeed XIAO ESP32-S3 on a 1S LiPo, asleep almost all the time. It has one GPIO for the switch, one for an LED, no WiFi password and no IP stack.

The board choice decides the battery life. Both a generic ESP32 devkit and the XIAO have chips that sleep at 7 to 10 µA, but the devkit's USB-serial chip and LDO keep drawing 4 to 20 mA in deep sleep. The XIAO has native USB and a buck regulator, and Seeed rates it around 14 µA asleep (fed from the BAT pads; through the 5 V pin it's closer to 300 µA). At twenty presses a day that works out to roughly a year and a half on an 1150 mAh cell, and sleep current is most of the budget until you press more than about five times a day.

The wake path is short on purpose:

  1. Wake on the button pin going low (or a 12-hour timer, which sends a telemetry frame and goes back to sleep).
  2. Debounce before spending any radio time: 5 samples, 5 ms apart, must agree. A wake is a claim until the debouncer confirms it.
  3. Send the PRESS frame to the master, up to 3 tries with the same sequence number.
  4. Wait up to 400 ms for the first ack, then up to 9 s for the Controller's verdict.
  5. Wait for release, lock out for 1.5 s (so an enthusiastic double-tap saves one moment, not two), and sleep.

Three things bit me on the S3 and are easy to mistake for firmware bugs:

Also: re-arm the pull-up in the RTC domain before sleeping (the normal GPIO pull-ups power down), and never arm the wake pin while the button is still held, or a stuck button wakes the chip forever on a battery.

ESP-NOW: why three boards instead of one

ESP-NOW and WiFi station mode share one radio, and a station follows its access point's channel. Put the button's receiver on the same chip as the WiFi connection and every ESP-NOW node has to follow the router's channel, including after a reboot picks a new one. That's the failure where the button works for a month and then stops, which is the worst behaviour for a device whose only job is catching something that already happened.

So the work is split across three boards:

The extra board costs about $4. In exchange the ESP-NOW channel is fixed forever, the buttons hold no WiFi credentials (lose one and nothing on the network is exposed), a press involves no WiFi association (about 10 ms of radio instead of 1 to 3 s of joining), and the master is about 300 lines that never need updating.

Pick an ESP-NOW channel at least 5 away from your access point (AP on 6, ESP-NOW on 1 or 11). An adjacent channel is worse than sharing one.

The frame

Every frame is a 28-byte packed header plus up to 160 bytes of ASCII payload. The header, byte by byte:

Frames are authenticated, not encrypted. ESP-NOW's built-in encryption tops out at 6 encrypted peers on an ESP32 and can't cover broadcast, and I want to add lights and sensors to the same network later. So the check is done in the application with a shared 32-byte key: someone sniffing can see that a button was pressed, and can't fake or replay a press.

Dedupe is keyed on (node, session, seq). A retry reuses its sequence number, so the master acks a duplicate but forwards only the first. A lost ack costs one more radio frame, never a second set of clips. New buttons need no pairing step; the master registers a node on its first authenticated frame.

The LED is the interface

The first blink is why it feels like a button and not a switch you hope did something.

The bridge: UART in, HTTP out

The master and bridge are joined by three wires: TX and RX crossed, and a shared ground. (Leave out the ground and it works on the bench, where both boards share a laptop's ground, and fails the moment they're on separate supplies.)

The UART carries newline-delimited JSON at 115200 baud in both directions. At a few frames a day there's nothing to gain from packing bytes, and a lot to gain from being able to clip on a USB-serial adapter and read the traffic, or type a fake press by hand:

{"t":"press","node":"a3f1","dev":"button","seq":12,"rssi":-61,"data":{"boot":42,"press":17}}
{"t":"ack","node":"a3f1","seq":12,"ok":true,"code":200}

A press becomes POST /api/twab on the Controller. Because a Blue Iris trigger can't safely be repeated, the HTTP status codes are a contract:

The bridge is the only board that updates over the air (signed images, automatic rollback if the new firmware can't reach the Controller). It also serves GET /status: firmware, RSSI, whether the master is talking, and how long ago the last press came in. When a press goes missing, a healthy master link plus a stale last-press time puts the fault at the button (on mine, the battery connector had pulled out). That's one HTTP request instead of a serial cable.

Talking to Blue Iris

The Controller is Python (FastAPI) with a phone-sized web UI. Everything it asks of Blue Iris is POST http://<host>:81/json with a JSON body.

Login is a two-step challenge. The first {"cmd":"login"} returns "fail" plus a nonce (that's expected), and the second sends md5("user:nonce:password"). The nonce becomes the session for every later command, and any non-success result is treated as an expired session: log in again, retry once.

The calls that matter:

Two traps. manrec has to be a top-level key; nested inside data it's accepted and does nothing. And the Blue Iris user needs the clip-creation permission, or manual recording silently does nothing too.

What a press does

  1. One at a time. Presses are serialised behind a lock (three simultaneous presses once produced three receipts all claiming a full minute for one merged clip).
  2. Check each camera before triggering it. Blue Iris returns success for a trigger on a dead camera, and it reports a dead RTSP source as online for about 20 seconds. The Controller trusts its own health poller instead.
  3. Trigger each armed camera, publish a receipt to the UI, and answer the bridge (which sends the second ack to the button).
  4. Verify in the background. Wait 16 s for the post-roll to finish, find the newest clip at or before the trigger with cliplist, and run ffprobe on it. This step exists because one file in 99 came back 5.56 MB with no H.264 start code, inside a press that had returned 200.
  5. File it. Copy (never move, so Blue Iris's database keeps working) all the clips into one folder named for the moment, tag them full-range, and transcribe the microphone track with Whisper so I can search the day's clips by what I said.

Measured press-to-verdict: 765 ms with one camera armed, 1,423 ms with three, 1,164 ms with ten. It doesn't scale with the camera count; the variance is Blue Iris's response time. The radio hop is about 10 ms.

The pre-roll setting nobody mentions

Blue Iris can do all of this out of the box, but almost none of it is the default. Each camera is set to record "when triggered" with motion detection off (so only the button can trigger it), a 60-second pre-trigger time, and 10 seconds of break time after.

The setting that trips everyone up is movieroll. The name suggests file rollover. It's the stream buffer, it defaults to 5 seconds, and it silently caps the pre-trigger time. With the pre-trigger set to 60 s and movieroll left alone, my button saved 4.1 s of lead-in. With movieroll at 60 s: 58.7 s. The button looks like it works either way.

(the registry stores these in deciseconds, per camera, across seven profiles. the repo has the exact keys and a script that sets them.)

The buffer also behaves like a bucket. It refills in real time and a press empties it, so a second press 12 seconds after the first gets about 2 seconds of lead-in, and it takes 71 seconds to fill back up. Blue Iris doesn't expose the buffer depth, so the Controller predicts it and shows it on every receipt ("saved, but with only 1.2 s of lead-in").

Turning USB cameras into H.264 cameras

Three of my cameras are USB (two OBSBOT webcams and a varifocal bench camera), and a fourth feed is the shop PC's own screen. Blue Iris will record them, but it only keeps an encoded buffer for network cameras recording direct to disk. A USB camera's buffer is raw frames, capped at one second. I measured it: the "60-second" pre-roll on a USB camera held about 0.1 s.

So every USB device goes through a bridge: one ffmpeg process per device encodes it to H.264 and publishes it to a local RTSP server (MediaMTX), and Blue Iris adds that as an ordinary IP camera. The same camera went from 0.1 s of pre-roll to 58.2 s.

The encode I settled on for a 1080p USB camera:

ffmpeg -f dshow -vcodec mjpeg -video_size 1920x1080 -framerate 30 \
       -use_wallclock_as_timestamps 1 -i video="<camera>":audio="<camera mic>" \
       -pix_fmt yuvj420p \
       -c:v h264_nvenc -preset p4 -tune ll -rc cbr -b:v 8M -maxrate 8M -bufsize 8M \
       -g 30 -bf 0 -delay 0 \
       -c:a aac -b:a 128k -ar 48000 -ac 1 \
       -f rtsp -rtsp_transport tcp rtsp://127.0.0.1:8554/cam1

Every flag in there is the answer to something that went wrong:

Blue Iris has no API for creating a camera, so a script in the repo stops it, clones a working camera's registry settings, and points the copy at the new RTSP path. A small supervisor checks every bridge every 30 seconds by pulling one frame, and shuts idle bridges off after two hours.

The wireless lav mic goes through the same trick. Its receiver is its own Blue Iris "camera" (the audio, plus a waveform drawn as video so the health check has a picture), so one press saves the voice track along with all nine angles.

The cameras

The other five are cheap IP cameras from different brands, and matching nine mismatched cameras so they cut together was a project of its own: five different control APIs (Dahua CGI, ONVIF, a XiongMai binary protocol, Foscam CGI, and UVC for the USB ones), each with its own ways of reporting a setting it didn't apply. The calibration toolkit and every per-camera setting are in the repo, and the full technical notes cover it. The short version: a setting that reads back correctly is a claim, and the recorded file is the evidence.

Numbers

Build your own

Everything is in the repo: the Controller, all three firmware trees (with instructions for where your MACs and key go), the USB bridge scripts, and the Blue Iris settings with the measurement behind each one. The parts list, with links, is on the build page. You'll need Blue Iris (paid, Windows only), a PC for it, two plain ESP32 boards and one XIAO ESP32-S3 per button.