Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ble-loadgen

BLE advertisement load generator for the ESPresense HIL bench. It advertises from a static random address ~40 times a second, rotating through a pool of 4096 of them, so every rotation costs a listening node a new fingerprint slot — a room full of phones, compressed. It exists to make the slow heap decline in ESPresense#2309 show up in a HIL window instead of over days on a shelf.

Split out of firmware-tester because it shares nothing with it: no PlatformIO, no serial, no toolchain — just Python stdlib and a raw HCI socket. Its own image (ghcr.io/espresense/ble-loadgen) stays tiny and releases on its own cadence.

How it runs

As a detached step in the ESPresense HIL pipeline (.woodpecker/hil.yml): it floods for the run and the runner kills it at the end — no host service, no gating.

- name: ble-flood
  image: ghcr.io/espresense/ble-loadgen:1
  detach: true
  privileged: true
  network_mode: host          # raw HCI only works in the host netns (see below)
  # no commands: the image ENTRYPOINT is `python3 ble_flood.py`, CMD supplies --index/--rate

Requirements

  • network_mode: host. HCI_CHANNEL_USER only works in the host network namespace — a bridge-net container can't even open an AF_BLUETOOTH socket (EAFNOSUPPORT). privileged supplies CAP_NET_ADMIN.
  • A USB Bluetooth adapter on the host. ls /sys/class/bluetooth should show hci0.
  • BlueZ absent. HCI_CHANNEL_USER takes exclusive control of a down adapter; bluetoothd would fight for it. Don't install it, or mask it.

On EBUSY the flood retries for 30s before giving up — busy is usually transient. If it still fails, something is holding the adapter: bluetoothd (systemctl mask --now bluetooth) or a leftover detached ble-flood from a previous HIL run (pkill -f ble_flood.py). Missing CAP_NET_ADMIN or a missing adapter fails immediately instead, and the message says which.

Usage

The script is pure stdlib — run it directly:

python3 ble_flood.py --selftest              # framing + address rules, no hardware
sudo python3 ble_flood.py --index 0 --rate 40        # flood until killed
sudo python3 ble_flood.py --index 0 --seconds 30     # one 30s burst
sudo python3 ble_flood.py --index 0 --pool 512       # smaller address pool

Address pool

Addresses come from a pool (default 4096, or $BLE_FLOOD_POOL) and repeat once it wraps. The churn a node sees is unchanged — every rotation is still a different MAC until the pool wraps — but anything downstream that keys off the address stops growing at --pool rows. That matters because the flood's addresses escape the bench: ESPresense Companion turns each one into an MQTT discovery config and Home Assistant into a device_tracker, and BlueZ caches each under /var/lib/bluetooth/*/cache. A 58-hour unbounded soak minted 8.3M addresses, left 53k orphaned HA entities, and exhausted the inodes on the HA host.

--pool 0 restores the old unbounded behaviour. Only use it against a bench whose subscribers you are willing to rebuild.

Docker

The image entrypoint is python3 ble_flood.py, so arguments go straight after the image. Flooding needs the host network namespace (raw HCI) and CAP_NET_ADMIN:

# flood until stopped — default args are --index 0 --rate 40
docker run --rm --network host --privileged ghcr.io/espresense/ble-loadgen:1

# one 30s burst on hci0
docker run --rm --network host --privileged ghcr.io/espresense/ble-loadgen:1 --seconds 30

# selftest needs neither host net nor privileged (it touches no socket)
docker run --rm ghcr.io/espresense/ble-loadgen:1 --selftest

--cap-add NET_ADMIN in place of --privileged also works; --privileged is what the HIL pipeline already grants, so the docs use it for parity.

Troubleshooting

HCI_Reset unanswered, or no completion event — the controller is claimed (the bind succeeded, so nothing else can be driving it) but it is not answering. The reset is retried 3 × 15s for exactly this reason: a command sent into the window right after a user-channel bind can be lost, or answered only once a USB part finishes its firmware setup. The bench hit this while crow did not. If it still gives up, the message prints the rfkill state and the causes a successful bind cannot explain — a soft rfkill block, an autosuspended USB port, a dongle that needs a re-plug, or a flood leaked from an earlier run (pkill -f ble_flood.py).

The bench is running old code. :1 and :1.0 are semver tags: they only move when a v* git tag is pushed, while latest/sha-* track main. So a fix merged to main does not reach a pipeline that pins :1 until a release is cut — which is how the bench spent six weeks on an image that predated the SIGTERM cleanup, leaving hci0 wedged between runs.

Why static random addresses with the MAC in the name

ESPresense keys ID_TYPE_RAND_STATIC_MAC off the top two bits of the address MSB, so each rotation is a distinct identity. Each advert also carries a name containing its own MAC — without that, a live node collapsed two addresses to one id (ID_TYPE_NAME outranks ID_TYPE_RAND_STATIC_MAC), and the id space this exists to exercise never churned. That detail came from the bench, not theory.

About

BLE advertisement load generator for the ESPresense HIL bench

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages