A camera watching your backyard only sees what’s in its cone of view, and every camera you bolt on is another privacy trade-off. IronEar takes the opposite bet: three tiny microphones and a neural network that can tell a barking dog from breaking glass, entirely on-device.
What IronEar actually does
IronEar is an open-source sensor node built around an ESP32-S3-WROOM-1 with 16 MB of flash and 8 MB of PSRAM. It classifies audio locally about once per second, matching against more than 500 sound categories: speech, barking, running water, breaking glass, laughter, vehicles, using TinyEar, a lightweight port of Google’s YAMNet model. Nothing gets uploaded to a server, and there’s no subscription keeping the classifier alive.
A scene-learning mode lets you record a few minutes of a specific environment and teach the device to recognize that acoustic signature later. That’s useful for sounds that don’t map cleanly onto YAMNet’s built-in classes, like one particular door latch or a machine that’s starting to fail.
The hardware stack
Sound capture comes from three Infineon IM73D122 MEMS microphones arranged in a triangular array. For connectivity, IronEar talks to Home Assistant over MQTT when Wi-Fi is available, and falls back to a Semtech SX1262 LoRa radio running a private Meshtastic channel when it isn’t, which is useful for a farm shed or an off-grid cabin. A Quectel LC76G GNSS receiver plus a magnetometer let each node report both position and heading, so a scattered sensor mesh can tell you which unit heard the sound and roughly where it was standing.
- ESP32-S3-WROOM-1, 16 MB flash, 8 MB PSRAM
- 3x Infineon IM73D122 MEMS mics, triangular array
- TinyEar (YAMNet port), 500+ sound classes, ~1 Hz inference
- Semtech SX1262 LoRa + Meshtastic fallback, Quectel LC76G GNSS
Build it yourself
Most of the design, hardware files, firmware, mesh protocol docs, and TinyEar weights are slated to go open source (the inference engine and its training data stay closed). Pricing and release date aren’t out yet, so the fastest way in right now is the ESP32-S3 side of it: wire an ESP32-S3-WROOM-1 dev board to an I2S MEMS mic on a breadboard, drop in a TensorFlow Lite Micro library running a YAMNet-class model, and publish the classifications to Home Assistant over MQTT. That’s a realistic thesis-scope build for an ECE elective, and it’s the same pipeline IronEar itself runs on. Full writeup and update signups: Hackster.io.
Frequently Asked Questions
What hardware powers IronEar’s sound recognition?
An ESP32-S3-WROOM-1 (16 MB flash, 8 MB PSRAM) paired with three Infineon IM73D122 MEMS microphones in a triangular array. Classification runs locally about once per second using TinyEar, a lightweight port of Google’s YAMNet audio model, recognizing over 500 sound classes.
How does IronEar stay connected without Wi-Fi?
It falls back to a Semtech SX1262 LoRa radio on a private Meshtastic mesh, so it keeps working off-grid at a farm, cabin, or campsite. A Quectel LC76G GNSS receiver and a magnetometer let each node report its position and heading so you know which sensor triggered.
What will I learn if I build something like this?
You’d practice I2S audio capture on a microcontroller, running a TensorFlow Lite Micro model on-device instead of in the cloud, and wiring a sensor into Home Assistant over MQTT, a solid combo of embedded ML and IoT skills for an ECE thesis or capstone.
