---
title: "natKit-IMU"
url: "https://maker.wiznet.io/Lihan__/projects/natkit-imu/"
markdown_url: "https://maker.wiznet.io/Lihan__/projects/natkit-imu/md"
type: "UCC: User Created Content"
author: "neuralbertatech"
author_url: "https://github.com/neuralbertatech/natKit-IMU"
editor: "WIZnet"
editor_url: "https://maker.wiznet.io/"
original_author: "neuralbertatech"
original_url: "https://github.com/neuralbertatech/natKit-IMU"
published: "2026-08-30"
language: "en"
likes: 0
views: 185
comments: 0
source: "WIZnet Makers (https://maker.wiznet.io/)"
---

# natKit-IMU

> ESP32-based IMU sensor device for the natKit BCI toolkit. This device collects inertial measurement data and streams it to the natKit backend via MQTT.

Original author: neuralbertatech (source: https://github.com/neuralbertatech/natKit-IMU)

## Article

## atKit-IMU — A Sensor Network That Chose a Cable to Save the Airwaves

`#ESP32S3` `#W5500` `#MACRAW` `#esp_eth` `#ESP-NOW` `#MQTT` `#IMU` `#BCI` `#ESP-IDF`

> 📚 A BCI toolkit project by NeurAlbertaTech (NAT), a non-profit student research group at the University of Alberta A rare, genuinely verifiable repository — the source comments record measured numbers **and the wrong calls the team made**

---

### 01 — What is this project?

#### Background: BCI and motion artifacts

**BCI (Brain-Computer Interface)** is the technology of reading brain signals to control machines. The most common approach is **EEG**, which measures brainwaves through electrodes on the scalp — and it comes with one chronic problem.

The brain signals EEG picks up are **extraordinarily faint, on the order of microvolts (µV)**. But when a person turns their head, blinks, or clenches their jaw, the electrical signal that movement produces bleeds in at **tens of times the amplitude** of the brainwave itself. This is called a **motion artifact**. You are trying to look at a brain signal, and body movement paints over the whole screen.

The fix is **to record separately when the movement happened**. If you know the timing and magnitude of the motion, you can subtract that contamination afterward. The component that keeps that record is the **IMU (Inertial Measurement Unit)** — an inertial sensor package containing an accelerometer, gyroscope, and magnetometer.

#### So this project is

![](https://maker.wiznet.io/upload/ckeditor5/629382507%5F1788174297%2Epng)

The hub (primary) runs on an **Espressif ESP Thread Border Router board**, with the **W5500 sitting on the separately-sold Sub-Ethernet daughter board** rather than the main board itself. Every other board is a plain off-the-shelf ESP32 or ESP32-C3 devkit.

**natKit-IMU** is the **motion sensor arm** of `natKit`, the BCI toolkit NeurAlbertaTech is building. IMU nodes are attached at multiple points on the body, and 100 samples per second stream out to an MQTT broker. The end goal is a **labeled dataset for training activity-recognition models**.

From that goal comes the one constraint that governs every design decision in this project.

> **Data from multiple sensors must align precisely on a single timeline.**

To remove an artifact, "this instant in the EEG" and "this movement in the IMU" have to be the same instant. The same holds when labeling training data. **If clocks between nodes drift apart by milliseconds, labels land on the wrong window and the entire dataset is contaminated.**

#### Why multiple boards is hard

One sensor board is easy. Put WiFi on an ESP32, spin up a single MQTT client, done.

**The original Arduino-based version worked exactly that way.** Every sensor board associated to the access point on its own, ran its own MQTT client, and pulled its own time over NTP. There was no hub and no hierarchy.

As boards were added, that approach collapsed. The airwaves got crowded, batteries drained, and — worst of all — **every board's clock ran on its own.** Each pulled time from the internet independently, so errors accumulated differently on each one, which runs directly against the single-timeline requirement above.

So the team redesigned the firmware from the ground up. **WiFi, MQTT, and NTP were stripped out of the sensor boards entirely**, leaving them to transmit over ESP-NOW to a hub — a **role-separated hierarchy** (pure ESP-IDF).

> **ESP-NOW** — Espressif's connectionless transport: it uses the WiFi radio but skips association, IP, and TCP entirely, sending frames straight to a peer's MAC address.

Now exactly one board reaches the internet and distributes time. **And this is where the W5500 enters** — the old design had no place for Ethernet, but once a hub role existed, something had to carry its uplink.

Here is the head-to-head, same boards, same conditions, three hours.

|  | Old (WiFi → MQTT) | New (ESP-NOW → hub → MQTT) |
| --- | --- | --- |
| Usable samples per board | 56–80 /s | **98 /s** |
| Inter-arrival jitter | 17 ms | **1.4 ms** |
| Board-to-board clock agreement | 2.7–5.5 ms | **0.1–0.2 ms** |
| Bandwidth | baseline | **half** |

**Board-to-board clock error went from 5.5 ms to 0.2 ms.** A human arm takes several hundred milliseconds to complete a motion. At 5 ms of skew you cannot tell whether the shoulder or the wrist moved first; at 0.2 ms that ordering becomes something you can actually reason about. This is not an "improvement" — it is **the line between usable and unusable**.

---

### 02 — Data flow

There are three roles. **leaf** only reads sensors and transmits wirelessly, **primary** is the hub that collects them and acts as timing master, and **gateway** appears only when needed to hand data off to the broker.

![](https://maker.wiznet.io/upload/ckeditor5/629382507%5F1788167232%2Epng)

**Three uplink branches.** The one that misleads people is #2. **UART is not a destination — it is a short bridge to the chip next door**, and what actually reaches the broker is that chip's WiFi. So #1 and #2 both ultimately go out over the air, and **only #3 is wired end to end**.

| Path | Boards required | Result |
| --- | --- | --- |
| 1. WiFi | 1× ESP32 devkit | **85% frame loss** (→ 03) |
| 2. UART → WiFi | ESP32 devkit + C3 devkit, wired with 3 UART lines | no loss, but one more chip |
| 3. Ethernet | Thread BR board + W5500 daughter board | no loss, and one chip |

This choice is also fixed at build time and does not change once flashed. On the ESP32-S3, #3 is the default.

**Behind the broker.** The broker (mosquitto) only delivers; a bridge forwards to Kafka, which stores the stream as a time series. From there it forks four ways — live signal monitoring, experiment session records such as calibration and task runs, **labeling and model training**, and export for external analysis.

---

### 03 — Why the architecture looks like this: 85% frame loss

Looking at the flow alone, you might wonder why the roles were split three ways and the uplink three ways. The answer sits in a single measurement, recorded in the `ethernet_net.hpp` comments.

> With the primary associated to WiFi, **roughly 85% of ESP-NOW frames were lost.** The MAC layer was even ACKing them — the WiFi task above it, busy servicing association, threw them away. The leaves themselves were perfectly fine.

The MAC sending an ACK means **the radio waves arrived correctly**. This is not something you fix by adjusting an antenna or transmit power; it is **something you fix by changing the structure**.

Which is where the role separation seen in 02 came from.

![](https://maker.wiznet.io/upload/ckeditor5/629382507%5F1788180557%2Epng)

| Role | What it does | Networking |
| --- | --- | --- |
| **leaf** | reads sensors and transmits | ESP-NOW **only** |
| **primary** | collection hub + timing master | ESP-NOW + uplink |
| **gateway** | receives and forwards to broker | WiFi (only when needed) |

Roles are not a runtime switch but a **Kconfig build option**. The leaf image contains **no WiFi and no MQTT code at all.** Since ESP-NOW does not go through LwIP, a leaf never even calls `esp_netif_init()`. As a bonus, **only one board in the entire system holds WiFi credentials** — the one doing the uplink.

---

### 04 — Why W5500? ⭐

#### 🔷 In the author's own words: "sidesteps the argument entirely"

The title of the comment block in `main/ethernet_net.hpp` is the answer — **"Why this exists, and why it is the better answer."**

> This epic assumed primary and gateway had to be two chips, because one radio cannot do ESP-NOW and WiFi association at the same time. Ethernet sidesteps that whole argument. **The uplink is OFF THE RADIO.** There is no association to inherit a channel from. **One chip, no serial bridge, no contention.**

The "off-radio uplink" that the UART approach barely achieves with two boards and three wires, **the W5500 delivers with a single chip.**

#### 🔷 Two ways it fails silently — and first-timers hit both

**1. The W5500 has no MAC address of its own.**

> No eFuse, no OTP. Leave it unset and it comes up all zeros or at a vendor default, and **two boards on the same network collide.**

This project derives one from the S3's base MAC (`ESP_MAC_ETH`).

**2. The GPIO ISR service must be installed before the driver.**

> Without it the driver logs "GPIO isr service is not installed" and **the link never comes up. It does not fail loudly; it simply never reports a connection.**

And losing `default y if IDF_TARGET_ESP32S3` produces this — *"it boots with no uplink, and the uplink counters cheerfully report frames 'sent' over a UART nobody is reading. This cost two debugging cycles."*

#### 🔷 Verified evidence ✅

The console prints `WIRED UPLINK: link up, ip …, mqtt up, published …`, with the primary publishing MQTT directly over Ethernet. The pin map is taken verbatim from Espressif's official example.

---

### 05 — Key components

**🌐 W5500 (MACRAW)** — The S3 has **no built-in Ethernet MAC**, so both MAC and PHY must come from outside. The W5500 is an integrated MAC+PHY solution behind SPI, which fits that requirement exactly.

**📡 ESP-NOW v2** — A 4-byte envelope with a magic byte and a version byte. The reasoning is practical: **on a shared channel, broadcasts from unrelated ESP-NOW devices also land in your receive callback.🎯 BNO08x + CEVA sh2** — Accel, gyro, mag, and rotation arrive as separate reports at different rates. So against a 100 Hz frame, acceleration is fresh nearly every time, orientation about 95% of the time, and magnetometer about 91%; each frame carries a `has_data` bit to mark which.

**🧭 Time synchronization** — A leaf has no NTP and no wall clock. The primary broadcasts two packets per second, and the leaf fits a rolling least-squares line over the most recent 32 pairs. The key detail is that **the transmit timestamp is read inside the send callback, not before calling **`**esp_now_send**` — read it beforehand and you are measuring transmit-queue latency, which **swings between 1.4 and 13.8 ms**, an error three orders of magnitude larger than the quantity being estimated. And a leaf **never rewrites its own timestamps**; correction is left to the consumer so it can be undone or improved later.

---

### 06 — Application scenarios

- **Motion capture · gait analysis** — With clock error down at the microsecond scale, the **ordering** of movement between joints becomes something you can discuss.

- **BCI motion artifact removal** — The project's original purpose (see 01).

- **Rehabilitation · sports form measurement** — Leaves carry no credentials, so field deployment stays simple.

- **Industrial vibration · multi-point measurement** — Factory WiFi is crowded and racks usually already have an Ethernet port. **A W5500 uplink fits this environment particularly well.**

---

### 07 — Field notes: wireless links fail quietly

#### 🔻 One channel separated everything from nothing

Measured with a leaf a few centimetres from the hub, changing only the channel.

| Channel | RSSI | Delivery |
| --- | --- | --- |
| 3 | −22 dBm | **all 10 frames/s** |
| 6 | −21 to −57 dBm | erratic |
| 9 / 11 | −61 / −83 dBm | **nothing delivered** |
| 1 | −82 dBm | **jammed by the board itself** |

Same board, same distance, same firmware — and the result is all or nothing. Worse, this is **a property of the location, not the firmware**, so a hardcoded value that works in one room is wrong in the next, and the failure mode looks like "a bad radio," which drags diagnosis out.

**The fix is clever.** All 13 channels are surveyed at boot, and it costs nothing — the primary cannot publish anything for roughly 33 seconds after reset anyway while NTP is unsynced, and that dead window is exactly enough to sweep them.

#### 🔻 A healthy board was convicted twice

One board was **written off as "physically damaged" and removed.** Its neighbour was clean while this one showed 274 sequence gaps, which looked like reasonable grounds. Four days later it was plugged back in, worked normally, and posted **the strongest RSSI of all four nodes at −36 dBm.**

The author's conclusion is the striking part — *"the original diagnosis was not so much wrong as **unsafe**."* The pattern that convicted it (one of two adjacent leaves degrading) went on appearing **between two known-good boards, with nothing changed, alternating on a 10–30 minute cycle.**

The other case was **simply unplugged.** The hub had just kept publishing a registry entry for a leaf that was no longer there. One principle came out of it — **"the absence of a data topic is not evidence about a board until you have confirmed it has power."**

#### 🔻 The plausible-looking metric is the dangerous one

The only metrics trusted are `**beacons_missed**`** and **`**leaf_send_failures**`**.** A healthy node missed 22 of 913 beacons with 4 send failures; the problem node missed **616 with 1800 failures.**

What got discarded:

- **Clock-fit residual** — the healthy node and the dying node were **equally saturated**. Zero discriminating power.

- **Hardware receive timestamp** — a genuine hardware stamp, yet it scattered **20–57 ms against **`**esp_timer**`**.** Just reading `esp_timer` in the callback gives 25 µs.

#### 🔻 Still open

Powering the Ethernet rig's S3 and the WiFi rig at the same time makes them **adopt each other as neighbouring hubs** on the shared channel. The current response is avoidance — the S3 stays powered down while the WiFi rig runs. But opening a USB console on the S3 resets and re-enumerates it, so this board **must be diagnosed from its published status topics rather than the console.**

Also outstanding on that board: **the S3's transmissions are heard fine by the leaves, but leaf→S3 unicasts mostly go unacknowledged.** Transmitter healthy, receiver deaf — the signature of an in-band interferer sitting right next to it. The suspect is the **ESP32-H2 co-processor** a few millimetres away, whose stock firmware transmits 802.15.4 in the same 2.4 GHz band. A Kconfig option that holds the H2 in reset was added to **turn the hypothesis into an experiment.** The Ethernet uplink was disabled during this investigation and the symptom persisted, which cleared the W5500 as a cause.

---

### Conclusion

> **"What you have not measured, you do not know" — this project proves that principle in its source comments rather than its README.**

- ✅ Role separation took **clock error from 5.5 ms to 0.2 ms**, jitter from **17 ms to 1.4 ms**, and halved bandwidth

- ✅ Measured **85% frame loss** when one chip does both, and pinned the cause on the upper task rather than the MAC

- ✅ **Removed the contention structurally with a wired W5500 uplink** — one chip, no bridge

- ✅ Documented the W5500's two traps (**no MAC address**, **ISR service first**) in code and comments

- ✅ Surveys 13 channels at boot — **free, by using the 33-second NTP wait**

- ✅ Publicly corrected two hardware misdiagnoses, and separated trustworthy metrics from plausible ones by measurement

- ✅ A complete pipeline running from sensor all the way to a **labeled training dataset**

---

### Q&A

**Q. Why is a W5500 needed to add Ethernet to an ESP32-S3?** A. **The S3 has no built-in Ethernet MAC (EMAC).** On a classic ESP32 you attach a PHY to the internal EMAC, but the S3 needs both MAC and PHY from outside. The W5500 is an integrated MAC+PHY behind SPI, which fits exactly.

**Q. What are the most common traps in W5500 initialization?** A. ① **It has no MAC address of its own** — leave it unset and two boards on the same network collide. ② **The GPIO ISR service must be installed before the driver.** Both fail silently, with no error.

**Q. Can the three uplinks be switched at runtime?** A. No. They are **Kconfig build options**, decided at compile time, with the code excluded entirely via `#ifdef`. On the S3, Ethernet is the default.

**Q. Why not just hardcode the channel?** A. Channel 3 delivered all 10 frames/s while channels 9 and 11 delivered **none at all.** And because this is **a property of the location**, a value that is right in one room is wrong in the next.

**Q. Where does the collected data ultimately go?** A. The broker only delivers; a bridge forwards to Kafka for storage. From there it splits into live monitoring, experiment session records, **labeling and model training**, and export. The final product is a labeled training dataset.

---

## [한글 버전] natKit-IMU — 전파를 아끼려고 랜선을 고른 센서 네트워크

`#ESP32S3` `#W5500` `#MACRAW` `#esp_eth` `#ESP-NOW` `#MQTT` `#IMU` `#BCI` `#ESP-IDF`

> 📚 캐나다 앨버타 대학 기반 비영리 연구단체 NeurAlbertaTech(NAT)의 BCI 툴킷 프로젝트 소스 주석에 실측 수치와 **틀린 판단까지** 남겨둔, 보기 드물게 검증 가능한 저장소

---

### 01 — 무엇을 만들었나

#### 배경: BCI와 모션 아티팩트

**BCI(Brain-Computer Interface)** 는 뇌 신호를 읽어 기계를 제어하는 기술입니다. 두피에 전극을 붙여 뇌파를 측정하는 **EEG**가 가장 흔한 방식인데, 여기엔 고질적인 문제가 하나 있습니다.

EEG가 잡아내는 뇌 신호는 **마이크로볼트(µV) 단위로 극도로 미약**합니다. 그런데 사람이 고개를 돌리거나 눈을 깜빡이거나 턱에 힘을 주면, 그 움직임이 만드는 전기 신호가 뇌파보다 **수십 배 크게** 섞여 들어옵니다. 이걸 **모션 아티팩트(motion artifact)** 라고 합니다. 뇌 신호를 보려는데 몸 움직임이 화면을 덮어버리는 셈입니다.

해결책은 **"언제 움직였는지를 따로 기록해두는 것"** 입니다. 움직임의 시각과 크기를 알면, 사후에 그 구간의 오염을 분리해낼 수 있습니다. 그 기록을 담당하는 게 **IMU(Inertial Measurement Unit)** — 가속도계·자이로스코프·지자기계가 든 관성측정 센서입니다.

#### 그래서 이 프로젝트는

**natKit-IMU**는 NeurAlbertaTech가 만드는 BCI 툴킷 `natKit`의 **모션 센서 파트**입니다. 몸 여러 곳에 IMU 노드를 붙이고 초당 100샘플을 MQTT 브로커로 흘려보냅니다. 최종 목적은 **동작 인식 모델 학습용 라벨링 데이터셋**입니다.

여기서 이 프로젝트의 모든 설계 판단을 지배하는 제약이 하나 나옵니다.

> **여러 센서의 데이터가 하나의 타임라인 위에 정확히 정렬돼야 한다.**

아티팩트를 제거하려면 "EEG의 이 시점"과 "IMU의 이 움직임"이 같은 순간이어야 합니다. 학습 데이터에 라벨을 붙일 때도 마찬가지입니다. **노드 간 시각이 밀리초 단위로 어긋나면 라벨이 엉뚱한 구간에 붙고, 데이터셋 전체가 오염됩니다.**

#### 왜 여러 노드가 어려운가

센서 보드가 하나면 쉽습니다. ESP32에 WiFi 붙이고 MQTT 클라이언트 하나 띄우면 끝입니다.

**초기 버전(Arduino 기반)이 실제로 그 방식이었습니다.** 센서 보드가 저마다 공유기에 접속해 각자 MQTT 클라이언트를 띄우고, 각자 NTP로 시간을 받아왔습니다. 허브도, 계층도 없었습니다.

보드가 늘어나자 이 방식이 무너졌습니다. 전파는 붐비고, 배터리는 녹고, 무엇보다 **보드마다 시계가 따로 놉니다.** 각자 인터넷에서 시간을 받아오니 오차가 제각각 쌓이는데, 이건 앞서 말한 "하나의 타임라인" 요구와 정면으로 충돌합니다.

그래서 팀은 펌웨어를 통째로 다시 설계했습니다. **센서 보드에서 WiFi와 MQTT와 NTP를 전부 걷어내고**, ESP-NOW로 허브에만 쏘도록 **역할을 나눈 계층 구조**(순정 ESP-IDF)로 갈아엎었습니다.

> **ESP-NOW** — WiFi 라디오를 쓰되 접속·IP·TCP 절차를 모두 건너뛰고, 상대 MAC 주소만으로 프레임을 바로 쏘는 Espressif의 연결 없는 통신 방식.

이제 인터넷에 접속하고 시간을 배포하는 건 허브 하나뿐입니다. 같은 보드·같은 조건 3시간 비교 결과입니다.

|  | 기존 (WiFi → MQTT브로커) | 신규 (ESP-NOW → 허브 → MQTT브로커) |
| --- | --- | --- |
| 보드당 유효 샘플 | 56–80 /초 | **98 /초** |
| 도착 간격 지터 | 17 ms | **1.4 ms** |
| 보드 간 시각 일치도 | 2.7–5.5 ms | **0.1–0.2 ms** |
| 대역폭 | 기준 | **절반** |

**보드 간 시각 오차 5.5ms → 0.2ms.** 사람의 팔 하나가 움직이는 데 걸리는 시간이 수백 밀리초입니다. 5ms가 어긋나면 어깨와 손목 중 뭐가 먼저 움직였는지 알 수 없지만, 0.2ms면 그 순서를 논할 수 있습니다. 이건 "개선"이 아니라 **쓸 수 있느냐 없느냐의 경계**입니다.

---

### 02 — 데이터 흐름

역할이 셋으로 나뉩니다. **leaf**는 센서만 읽어 무선으로 쏘고, **primary**는 그걸 모으는 허브이자 타이밍 마스터이며, **gateway**는 필요할 때만 등장해 브로커로 넘깁니다.![](https://maker.wiznet.io/upload/ckeditor5/629382507%5F1788171737%2Epng)

**업링크 세 갈래.** 여기서 오해하기 쉬운 게 2번입니다. **UART는 종착지가 아니라 옆 칩까지 가는 짧은 다리**이고, 브로커로 나가는 건 결국 그 칩의 WiFi입니다. 즉 1번과 2번은 결국 무선으로 나가고, **3번만 처음부터 끝까지 유선**입니다.

![](https://maker.wiznet.io/upload/ckeditor5/629382507%5F1788175808%2Epng)

| 경로 | 필요한 보드 | 결과 |
| --- | --- | --- |
| 1. WiFi | ESP32 devkit 1장 | 프레임 **85% 손실** (→ 03) |
| 2. UART → WiFi | ESP32 devkit + C3 devkit, UART 3선 배선 | 손실 없음, 칩이 하나 더 |
| 3. Ethernet | Thread BR 보드 + W5500 도터보드 | 손실 없음, 칩도 하나 |

이 선택도 빌드 타임에 결정되고 보드에 올린 뒤엔 안 바뀝니다. ESP32-S3에서는 3번이 기본값입니다.

**브로커 뒤.** 브로커(mosquitto)는 배달만 하고, 브리지가 Kafka로 넘겨 시계열로 쌓습니다. 거기서 네 갈래로 갈립니다 — 실시간 신호 관측, 캘리브레이션·과제 실행 같은 실험 세션 기록, **라벨링과 모델 학습**, 외부 분석용 내보내기.

---

### 03 — 왜 이런 구조가 됐나: 프레임 85% 손실

흐름만 보면 "왜 굳이 역할을 셋으로 쪼개고 업링크를 세 갈래로 뒀나" 싶습니다. 답은 이 측정치 하나에 있습니다. `ethernet_net.hpp` 주석입니다.

> primary가 WiFi에 associate된 상태에서 **ESP-NOW 프레임의 약 85%가 유실됐다.** MAC 계층에서는 ACK까지 받았는데, association을 처리하느라 바쁜 WiFi 태스크가 그 위에서 버렸다. 정작 leaf들은 완벽히 정상이었다.

MAC이 ACK를 보냈다는 건 **전파는 제대로 도착했다**는 뜻입니다. 안테나나 출력을 만져서 될 문제가 아니라 **구조를 바꿔야 하는 문제**입니다.

그래서 02에서 본 역할 분리가 나왔습니다.

| 역할 | 하는 일 | 네트워크 |
| --- | --- | --- |
| **leaf** | 센서만 읽고 쏜다 | ESP-NOW **only** |
| **primary** | 수집 허브 + 타이밍 마스터 | ESP-NOW + 업링크 |
| **gateway** | 받아서 브로커로 | WiFi (필요할 때만) |

역할은 런타임 스위치가 아니라 **Kconfig 빌드 옵션**입니다. leaf 이미지에는 WiFi도 MQTT도 **코드 자체가 안 들어갑니다.** ESP-NOW는 LwIP를 안 거치므로 leaf는 `esp_netif_init()`조차 호출하지 않습니다. 덤으로 **WiFi 자격증명을 가진 보드가 업링크 담당 하나뿐**이 됩니다.

---

### 04 — 왜 W5500인가? ⭐

#### 🔷 저자 본인의 표현: "논쟁 자체를 우회한다"

`main/ethernet_net.hpp`의 주석 제목이 그대로 답입니다 — **"Why this exists, and why it is the better answer."**

> 이 에픽은 primary와 gateway가 두 칩이어야 한다고 가정했다. 한 라디오가 ESP-NOW와 WiFi association을 동시에 못 하기 때문이다. 이더넷은 이 논쟁 전체를 우회한다. **업링크가 라디오 밖에 있다(OFF THE RADIO).** 채널을 물려받을 association도 없다. **칩 하나, 시리얼 브릿지 없음, 경합도 없음.**

UART 방식이 보드 두 장과 전선 세 가닥으로 겨우 만들어낸 "라디오 밖 업링크"를, **W5500은 칩 하나로 만듭니다.**

#### 🔷 조용히 실패하는 두 가지 — 처음 쓰는 사람이 그대로 밟는다

**1. W5500은 자기 MAC 주소가 없습니다.**

> eFuse도 OTP도 없다. 설정하지 않으면 전부 0이거나 벤더 기본값으로 올라오고, **같은 네트워크의 두 보드가 충돌한다.**

이 프로젝트는 S3의 base MAC(`ESP_MAC_ETH`)에서 유도해 넣습니다.

**2. GPIO ISR 서비스를 드라이버보다 먼저 설치해야 합니다.**

> 없으면 "GPIO isr service is not installed"를 찍고 **링크가 영원히 올라오지 않는다. 요란하게 실패하지 않고, 그냥 연결을 영영 보고하지 않는다.**

그리고 `default y if IDF_TARGET_ESP32S3`를 잃으면 이렇게 됩니다 — *"업링크 없이 부팅되고, 업링크 카운터는 아무도 안 읽는 UART로 프레임을 '보냈다'고 태연히 보고한다. 이걸로 디버깅 사이클을 두 번 날렸다."*

#### 🔷 검증된 증거 ✅

로그에 `WIRED UPLINK: link up, ip …, mqtt up, published …`가 찍히며 primary가 이더넷으로 직접 MQTT를 발행합니다. 핀맵은 Espressif 공식 예제 값 그대로입니다.

---

### 05 — 핵심 구성 요소

**🌐 W5500 (MACRAW)** — S3에는 **내장 이더넷 MAC이 없어서** MAC과 PHY를 모두 외부에서 가져와야 합니다. W5500은 SPI 너머의 MAC+PHY 통합 솔루션이라 이 요구에 정확히 맞습니다.

**📡 ESP-NOW v2** — 4바이트 봉투에 매직 + 버전 바이트. 이유가 실용적입니다 — **공유 채널에서는 무관한 ESP-NOW 기기의 브로드캐스트도 수신 콜백에 들어오기** 때문입니다.

**🎯 BNO08x + CEVA sh2** — accel/gyro/mag/rotation을 각각 다른 주기로 냅니다. 그래서 100Hz 프레임 기준 가속도는 거의 매번, 자세는 95%, 지자기는 91%만 새 값이고, 프레임마다 `has_data` 비트로 표시합니다.

**🧭 타이밍 동기화** — leaf에는 NTP도 벽시계도 없습니다. primary가 초당 두 패킷을 뿌리고 leaf가 최근 32쌍에 최소자승 직선을 적합합니다. 핵심은 **송신 시각을 **`**esp_now_send**`** 호출 전이 아니라 송신 콜백 안에서 읽는 것** — 호출 전에 읽으면 **1.4~13.8ms로 요동치는 송신 큐 대기시간**을 재게 되는데, 추정하려는 값보다 세 자릿수 큰 오차입니다. 그리고 leaf는 **자기 타임스탬프를 고쳐 쓰지 않습니다.** 보정을 되돌리거나 나중에 개선할 수 있게 소비자에게 맡깁니다.

---

### 06 — 응용 시나리오

- **모션 캡처 · 보행 분석** — 시각 오차가 마이크로초 단위라 관절 간 움직임의 **선후 관계**를 논할 수 있습니다.

- **BCI 모션 아티팩트 제거** — 이 프로젝트의 본래 목적(01 참조).

- **재활 · 스포츠 자세 측정** — leaf에 자격증명이 없어 현장 배포가 단순합니다.

- **산업 진동 · 다지점 계측** — 공장 WiFi는 붐비고 랙엔 랜 포트가 이미 있습니다. **W5500 업링크가 특히 잘 맞는 환경.**

---

### 07 — 필드 기록: 무선 링크는 조용히 끊어진다

#### 🔻 채널 하나가 전부와 전무를 갈랐다

허브에서 몇 센티 떨어진 leaf를 채널만 바꿔 측정한 결과입니다.

| 채널 | RSSI | 전달 |
| --- | --- | --- |
| 3 | −22 dBm | **10 frames/s 전부** |
| 6 | −21 ~ −57 dBm | 들쭉날쭉 |
| 9 / 11 | −61 / −83 dBm | **전달 없음** |
| 1 | −82 dBm | **보드 자신이 방해** |

같은 보드, 같은 거리, 같은 펌웨어인데 전부 아니면 전무입니다. 게다가 이건 펌웨어가 아니라 **장소의 속성**이라 한 방에서 맞는 하드코딩 값이 옆방에서 틀리고, 실패 양상이 "나쁜 라디오"처럼 보여 진단이 오래 걸립니다.

**해법이 영리합니다.** 부팅 시 13채널을 전부 훑는데, 이게 공짜입니다 — primary는 리셋 후 약 33초간 NTP 미동기라 어차피 아무것도 발행하지 못하고, 그 죽은 시간이 정확히 스캔에 충분합니다.

#### 🔻 멀쩡한 보드를 두 번 유죄로 만들었다

한 보드는 **"물리적 파손"으로 판정돼 제거됐습니다.** 옆 노드가 깨끗한데 이 보드만 시퀀스 갭 274개였으니 합리적으로 보였습니다. 나흘 뒤 다시 꽂으니 정상 동작했고 **RSSI −36dBm으로 네 노드 중 가장 강했습니다.**

저자의 결론이 인상적입니다 — *"원래 진단이 틀렸다기보다, **안전하지 않은** 진단이었다."* 유죄의 근거였던 패턴(나란한 두 leaf 중 하나만 나빠짐)이, 이후 **멀쩡한 두 보드 사이에서 아무것도 안 건드린 채 10~30분 간격으로 계속 번갈아** 나타났기 때문입니다.

또 하나는 **그냥 전원이 뽑혀 있었습니다.** 허브가 사라진 leaf의 레지스트리 항목을 계속 발행하고 있었을 뿐입니다. 원칙 하나가 남았습니다 — **"데이터 토픽이 없다는 사실은, 전원을 확인하기 전까지 증거가 아니다."**

#### 🔻 그럴듯한 지표가 가장 위험하다

신뢰하는 지표는 `**beacons_missed**`**와 **`**leaf_send_failures**`** 둘뿐**입니다. 정상 노드는 비콘 913개 중 22개 누락·실패 4회, 문제 노드는 **616개 누락·실패 1800회.**

반면 버려진 것들:

- **클럭 동기화 잔차** — 정상 노드와 죽어가는 노드가 **똑같이 포화**. 판별력 0.

- **하드웨어 수신 타임스탬프** — 진짜 하드웨어 스탬프인데 `esp_timer` 대비 **20~57ms 흩어짐.** 콜백에서 그냥 `esp_timer`를 읽으면 25µs입니다.

#### 🔻 아직 열려 있는 것

이더넷 릭의 S3와 WiFi 릭을 동시에 켜면 같은 채널에서 **서로를 인접 허브로 인식(adopt)** 해버립니다. 현재 대응은 회피 — WiFi 릭이 도는 동안 S3는 꺼둡니다. 그런데 S3는 USB 콘솔을 열면 리셋되며 재열거되므로, **콘솔이 아니라 발행된 상태 토픽으로만 진단**해야 합니다.

그리고 같은 보드에서 **S3의 송신은 잘 들리는데 leaf→S3 유니캐스트만 ACK를 못 받는** 현상이 남아 있습니다. 송신기는 멀쩡하고 수신기만 먹은 모습 — 용의자는 몇 밀리미터 옆에 얹힌 **ESP32-H2 코프로세서**입니다(같은 2.4GHz에서 802.15.4 송신). H2를 리셋에 붙잡아두는 Kconfig 옵션을 만들어 **가설을 실험으로 전환**해뒀습니다. 이 과정에서 이더넷도 꺼봤지만 증상이 그대로여서 W5500은 원인에서 배제됐습니다.

---

### Conclusion

> **"측정하지 않은 것은 알지 못한 것이다" — 이 프로젝트는 그 원칙을 README가 아니라 소스 주석으로 증명한다.**

- ✅ 역할 분리로 **시각 오차 5.5ms → 0.2ms**, 지터 **17ms → 1.4ms**, 대역폭 절반

- ✅ 한 칩 겸용 시 **프레임 85% 손실**을 측정하고 원인을 MAC이 아닌 상위 태스크로 특정

- ✅ **W5500 유선 업링크로 경합을 구조적으로 제거** — 칩 하나, 브릿지 없음

- ✅ W5500의 두 함정(**MAC 주소 부재**, **ISR 선행 설치**)을 코드와 주석으로 문서화

- ✅ 부팅 시 13채널 측정 — **NTP 대기 33초를 활용해 비용 0**

- ✅ 하드웨어 오진 두 건을 **공개 정정**하고, 믿을 지표와 아닌 지표를 측정으로 구분

- ✅ 센서에서 **라벨링·학습 데이터셋까지** 이어지는 완결된 파이프라인

---

### Q&A

**Q. ESP32-S3에 이더넷을 붙일 때 W5500이 필요한 이유는?** A. **S3에는 내장 이더넷 MAC(EMAC)이 없습니다.** 클래식 ESP32라면 내장 EMAC에 PHY만 붙이면 되지만, S3는 MAC과 PHY를 모두 외부에서 가져와야 합니다. W5500은 SPI 너머의 MAC+PHY 통합 솔루션이라 정확히 맞습니다.

**Q. W5500 초기화에서 가장 자주 걸리는 함정은?** A. ① **자기 MAC 주소가 없습니다** — 안 정해주면 같은 네트워크의 두 보드가 충돌합니다. ② **GPIO ISR 서비스를 드라이버보다 먼저** 설치해야 합니다. 둘 다 에러 없이 조용히 실패합니다.

**Q. 업링크 세 가지는 실행 중에 바꿀 수 있나요?** A. 아니요. **Kconfig 빌드 옵션**이라 컴파일 시점에 결정되고 `#ifdef`로 코드가 아예 빠집니다. S3에서는 이더넷이 기본값입니다.

**Q. 채널을 왜 하드코딩하지 않나요?** A. 채널 3은 10 frames/s를 전부 전달했고 9와 11은 **하나도 전달하지 못했습니다.** 그리고 이건 **장소의 속성**이라 한 방에서 맞는 값이 옆방에서 틀립니다.

**Q. 수집한 데이터는 최종적으로 어디에 쓰이나요?** A. 브로커는 배달만 하고, 브리지가 Kafka로 넘겨 저장합니다. 거기서 실시간 관측, 실험 세션 기록, **라벨링과 모델 학습**, 내보내기로 갈라집니다. 최종 산출물은 라벨이 붙은 학습 데이터셋입니다.

---

Source: https://maker.wiznet.io/Lihan__/projects/natkit-imu/
