How it fits together
XIAO ESP32-S3 --ESP-NOW--> ESP32 master --UART--> ESP32 bridge --HTTP--> Controller --JSON API--> Blue Iris
(battery, asleep) (no WiFi) (WiFi + OTA) (Python) (9 cameras)
A press travels four hops. Two acknowledgements come back the other way to the button's LED: one when the master hears the press (a few milliseconds), and one when Blue Iris has been told to save (about a second).
Most of the interesting problems were in that chain, so that's what this write-up covers first.
The button
The button is a Seeed XIAO ESP32-S3 on a 1S LiPo, asleep almost all the time. It has one GPIO for the switch, one for an LED, no WiFi password and no IP stack.
The board choice decides the battery life. Both a generic ESP32 devkit and the XIAO have chips that sleep at 7 to 10 µA, but the devkit's USB-serial chip and LDO keep drawing 4 to 20 mA in deep sleep. The XIAO has native USB and a buck regulator, and Seeed rates it around 14 µA asleep (fed from the BAT pads; through the 5 V pin it's closer to 300 µA). At twenty presses a day that works out to roughly a year and a half on an 1150 mAh cell, and sleep current is most of the budget until you press more than about five times a day.
The wake path is short on purpose:
- Wake on the button pin going low (or a 12-hour timer, which sends a telemetry frame and goes back to sleep).
- Debounce before spending any radio time: 5 samples, 5 ms apart, must agree. A wake is a claim until the debouncer confirms it.
- Send the PRESS frame to the master, up to 3 tries with the same sequence number.
- Wait up to 400 ms for the first ack, then up to 9 s for the Controller's verdict.
- Wait for release, lock out for 1.5 s (so an enthusiastic double-tap saves one moment, not two), and sleep.
Three things bit me on the S3 and are easy to mistake for firmware bugs:
- The XIAO has no onboard antenna. It ships with a loose U.FL antenna. Unplugged, the range collapses and it looks like the master is down.
- The USB console is the chip. In deep sleep the COM port vanishes and comes back on every wake. I build a bench variant with deep sleep disabled instead of editing code I'd forget to undo.
CONFIG_FREERTOS_HZmust be 1000. At the IDF default of 100,pdMS_TO_TICKS(5)rounds to zero and the debounce does nothing.
Also: re-arm the pull-up in the RTC domain before sleeping (the normal GPIO pull-ups power down), and never arm the wake pin while the button is still held, or a stuck button wakes the chip forever on a battery.
ESP-NOW: why three boards instead of one
ESP-NOW and WiFi station mode share one radio, and a station follows its access point's channel. Put the button's receiver on the same chip as the WiFi connection and every ESP-NOW node has to follow the router's channel, including after a reboot picks a new one. That's the failure where the button works for a month and then stops, which is the worst behaviour for a device whose only job is catching something that already happened.
So the work is split across three boards:
- The button (XIAO ESP32-S3, on the LiPo) wakes, sends one frame, shows the result and goes back to sleep.
- The master (a plain ESP32 on mains power) is the ESP-NOW receiver for the whole shop. It verifies, dedupes and acks, and never joins WiFi.
- The bridge (another plain ESP32 on mains) joins WiFi, takes frames from the master over UART, and POSTs them to the Controller.
The extra board costs about $4. In exchange the ESP-NOW channel is fixed forever, the buttons hold no WiFi credentials (lose one and nothing on the network is exposed), a press involves no WiFi association (about 10 ms of radio instead of 1 to 3 s of joining), and the master is about 300 lines that never need updating.
Pick an ESP-NOW channel at least 5 away from your access point (AP on 6, ESP-NOW on 1 or 11). An adjacent channel is worse than sharing one.
The frame
Every frame is a 28-byte packed header plus up to 160 bytes of ASCII payload. The header, byte by byte:
- Bytes 0-1,
magic:0x4157 - Byte 2,
version: 1 - Byte 3,
msg: HELLO, PRESS, TELEMETRY, ACK, CMD or EVENT - Byte 4,
dev: BUTTON, LIGHT, SENSOR or MASTER - Byte 5,
payload_len: up to 160 - Bytes 6-7,
node_id: from the low MAC bytes - Bytes 8-11,
session: random per boot - Bytes 12-15,
seq: increments per frame - Bytes 16-19,
ref_seq: on an ACK, the seq being acknowledged - Bytes 20-27,
tag: the first 8 bytes of HMAC-SHA256 over the frame, with the tag zeroed
Frames are authenticated, not encrypted. ESP-NOW's built-in encryption tops out at 6 encrypted peers on an ESP32 and can't cover broadcast, and I want to add lights and sensors to the same network later. So the check is done in the application with a shared 32-byte key: someone sniffing can see that a button was pressed, and can't fake or replay a press.
Dedupe is keyed on (node, session, seq). A retry reuses its sequence number, so the master acks a duplicate but forwards only the first. A lost ack costs one more radio frame, never a second set of clips. New buttons need no pairing step; the master registers a node on its first authenticated frame.
The LED is the interface
- One short blink: the master heard the press.
- One long solid light (about 0.6 s): Blue Iris was triggered. Saved.
- Two medium blinks: the master heard it, but the save failed.
- Four fast blinks: nobody answered (master down, wrong key, or the antenna isn't plugged in).
The first blink is why it feels like a button and not a switch you hope did something.
The bridge: UART in, HTTP out
The master and bridge are joined by three wires: TX and RX crossed, and a shared ground. (Leave out the ground and it works on the bench, where both boards share a laptop's ground, and fails the moment they're on separate supplies.)
The UART carries newline-delimited JSON at 115200 baud in both directions. At a few frames a day there's nothing to gain from packing bytes, and a lot to gain from being able to clip on a USB-serial adapter and read the traffic, or type a fake press by hand:
{"t":"press","node":"a3f1","dev":"button","seq":12,"rssi":-61,"data":{"boot":42,"press":17}}
{"t":"ack","node":"a3f1","seq":12,"ok":true,"code":200}
A press becomes POST /api/twab on the Controller. Because a Blue Iris trigger can't safely be repeated, the HTTP status codes are a contract:
- 200 means at least one camera triggered. The bridge stops; a retry would trigger the cameras again.
- 409 means no cameras are armed. The bridge stops.
- 502 means nothing triggered. The bridge retries, up to 3 times with an 8 s timeout.
The bridge is the only board that updates over the air (signed images, automatic rollback if the new firmware can't reach the Controller). It also serves GET /status: firmware, RSSI, whether the master is talking, and how long ago the last press came in. When a press goes missing, a healthy master link plus a stale last-press time puts the fault at the button (on mine, the battery connector had pulled out). That's one HTTP request instead of a serial cable.
Talking to Blue Iris
The Controller is Python (FastAPI) with a phone-sized web UI. Everything it asks of Blue Iris is POST http://<host>:81/json with a JSON body.
Login is a two-step challenge. The first {"cmd":"login"} returns "fail" plus a nonce (that's expected), and the second sends md5("user:nonce:password"). The nonce becomes the session for every later command, and any non-success result is treated as an expired session: log in again, retry once.
The calls that matter:
- The button:
{"cmd":"trigger","camera":"CAM1"} - List cameras (online, no signal, paused, fps, audio):
{"cmd":"camlist"} - Find the saved clip:
{"cmd":"cliplist","camera":"CAM1","startdate":<epoch>} - Manual take:
{"cmd":"camconfig","camera":"CAM1","manrec":true}
Two traps. manrec has to be a top-level key; nested inside data it's accepted and does nothing. And the Blue Iris user needs the clip-creation permission, or manual recording silently does nothing too.
What a press does
- One at a time. Presses are serialised behind a lock (three simultaneous presses once produced three receipts all claiming a full minute for one merged clip).
- Check each camera before triggering it. Blue Iris returns success for a trigger on a dead camera, and it reports a dead RTSP source as online for about 20 seconds. The Controller trusts its own health poller instead.
- Trigger each armed camera, publish a receipt to the UI, and answer the bridge (which sends the second ack to the button).
- Verify in the background. Wait 16 s for the post-roll to finish, find the newest clip at or before the trigger with
cliplist, and runffprobeon it. This step exists because one file in 99 came back 5.56 MB with no H.264 start code, inside a press that had returned 200. - File it. Copy (never move, so Blue Iris's database keeps working) all the clips into one folder named for the moment, tag them full-range, and transcribe the microphone track with Whisper so I can search the day's clips by what I said.
Measured press-to-verdict: 765 ms with one camera armed, 1,423 ms with three, 1,164 ms with ten. It doesn't scale with the camera count; the variance is Blue Iris's response time. The radio hop is about 10 ms.
The pre-roll setting nobody mentions
Blue Iris can do all of this out of the box, but almost none of it is the default. Each camera is set to record "when triggered" with motion detection off (so only the button can trigger it), a 60-second pre-trigger time, and 10 seconds of break time after.
The setting that trips everyone up is movieroll. The name suggests file rollover. It's the stream buffer, it defaults to 5 seconds, and it silently caps the pre-trigger time. With the pre-trigger set to 60 s and movieroll left alone, my button saved 4.1 s of lead-in. With movieroll at 60 s: 58.7 s. The button looks like it works either way.
(the registry stores these in deciseconds, per camera, across seven profiles. the repo has the exact keys and a script that sets them.)
The buffer also behaves like a bucket. It refills in real time and a press empties it, so a second press 12 seconds after the first gets about 2 seconds of lead-in, and it takes 71 seconds to fill back up. Blue Iris doesn't expose the buffer depth, so the Controller predicts it and shows it on every receipt ("saved, but with only 1.2 s of lead-in").
Turning USB cameras into H.264 cameras
Three of my cameras are USB (two OBSBOT webcams and a varifocal bench camera), and a fourth feed is the shop PC's own screen. Blue Iris will record them, but it only keeps an encoded buffer for network cameras recording direct to disk. A USB camera's buffer is raw frames, capped at one second. I measured it: the "60-second" pre-roll on a USB camera held about 0.1 s.
So every USB device goes through a bridge: one ffmpeg process per device encodes it to H.264 and publishes it to a local RTSP server (MediaMTX), and Blue Iris adds that as an ordinary IP camera. The same camera went from 0.1 s of pre-roll to 58.2 s.
The encode I settled on for a 1080p USB camera:
ffmpeg -f dshow -vcodec mjpeg -video_size 1920x1080 -framerate 30 \
-use_wallclock_as_timestamps 1 -i video="<camera>":audio="<camera mic>" \
-pix_fmt yuvj420p \
-c:v h264_nvenc -preset p4 -tune ll -rc cbr -b:v 8M -maxrate 8M -bufsize 8M \
-g 30 -bf 0 -delay 0 \
-c:a aac -b:a 128k -ar 48000 -ac 1 \
-f rtsp -rtsp_transport tcp rtsp://127.0.0.1:8554/cam1
Every flag in there is the answer to something that went wrong:
- MJPEG input. The camera's own H.264 output delivered zero frames over DirectShow.
-bf 0. Blue Iris trims B-frames from the oldest part of the pre-roll, which left 107 ms gaps once a second for the first 8 seconds of every clip.-g 30. One keyframe per second on every source. It costs about 11% in file size and makes scrubbing the timeline far better, and it lines keyframes up across cameras.-pix_fmtpinned. Unpinned, one camera came through a few points off in red and green inside Blue Iris. An ffmpeg-to-ffmpeg test can't see that.- Wall-clock timestamps. One camera claims 30 fps and delivers 32.3.
-use_wallclock_as_timestamps 1(plus-vf fps=30on that one) fixes it;-r 30doesn't. - No sliced threads with x264. If you use software x264 with
-tune zerolatency, it splits each frame into slices, Blue Iris counts every slice as a frame, decides the camera runs at 75 fps, and throws away audio to "fix" the sync (a 265 ms hole every 2.3 s). Use-threads 1 -x264-params sliced-threads=0. - No substream. Letting Blue Iris pull a second, low-res stream from a bridged camera cost 2 to 10 video gaps per two-minute take on the main recording. Without it: none.
Blue Iris has no API for creating a camera, so a script in the repo stops it, clones a working camera's registry settings, and points the copy at the new RTSP path. A small supervisor checks every bridge every 30 seconds by pulling one frame, and shuts idle bridges off after two hours.
The wireless lav mic goes through the same trick. Its receiver is its own Blue Iris "camera" (the audio, plus a waveform drawn as video so the health check has a picture), so one press saves the voice track along with all nine angles.
The cameras
The other five are cheap IP cameras from different brands, and matching nine mismatched cameras so they cut together was a project of its own: five different control APIs (Dahua CGI, ONVIF, a XiongMai binary protocol, Foscam CGI, and UVC for the USB ones), each with its own ways of reporting a setting it didn't apply. The calibration toolkit and every per-camera setting are in the repo, and the full technical notes cover it. The short version: a setting that reads back correctly is a claim, and the recorded file is the evidence.
Numbers
- Recording all nine continuously: 21.7 GB an hour. One nine-camera press: about 460 MB.
- Blue Iris CPU with seven cameras stream-copying to disk: 31 to 35%.
- Press to verdict: 0.77 to 1.42 s. Radio hop: about 10 ms.
- Battery (estimate, sleep current not yet measured on my board): about 19 months at 20 presses a day on 1150 mAh.
Build your own
Everything is in the repo: the Controller, all three firmware trees (with instructions for where your MACs and key go), the USB bridge scripts, and the Blue Iris settings with the measurement behind each one. The parts list, with links, is on the build page. You'll need Blue Iris (paid, Windows only), a PC for it, two plain ESP32 boards and one XIAO ESP32-S3 per button.
gearscodeandfire