---
title: "How Much Does W5500 Hardware Offload Save Over MACRAW on the Same Board?"
url: "https://maker.wiznet.io/Grace_Koo/projects/nucleo-h723zg-udp-echo/"
markdown_url: "https://maker.wiznet.io/Grace_Koo/projects/nucleo-h723zg-udp-echo/md"
type: "UCC: User Created Content"
author: "mixxen"
author_url: "https://github.com/mixxen/nucleo-h723zg-udp-echo"
editor: "WIZnet"
editor_url: "https://maker.wiznet.io/"
original_author: "mixxen"
original_url: "https://github.com/mixxen/nucleo-h723zg-udp-echo"
published: "2026-09-03"
language: "en"
hardware: ["WIZnet W5500"]
likes: 0
views: 107
comments: 0
source: "WIZnet Makers (https://maker.wiznet.io/)"
---

# How Much Does W5500 Hardware Offload Save Over MACRAW on the Same Board?

> Four UDP echo builds on one NUCLEO-H723ZG, measured for an hour each: W5500 hardware offload against MACRAW and the STM32's own MAC.

Original author: mixxen (source: https://github.com/mixxen/nucleo-h723zg-udp-echo)

## Components

- **WIZnet W5500** x 1 ([docs](https://docs.wiznet.io/Product/Chip/Ethernet/W5500))

WIZnet parts: W5500 ([Datasheet](https://docs.wiznet.io/Product/Chip/Ethernet/W5500/datasheet?utm_source=maker&utm_medium=project&utm_campaign=w5500), [product hub](https://maker.wiznet.io/products/w5500/))

## Article

## nucleo-h723zg-udp-echo — Four Ethernet Implementations on One Board, Each Run for an Hour

`#Rust` `#embassy` `#W5500` `#TOE` `#MACRAW` `#STM32H723` `#UDP` `#Benchmark` `#SocketOffload`

> 📚 **Context**: [mixxen/nucleo-h723zg-udp-echo](https://github.com/mixxen/nucleo-h723zg-udp-echo) — one UDP echo server on a NUCLEO-H723ZG, implemented four ways and compared. Created 2026-07-27, last pushed 2026-08-13. ✅ **Verification status**: Read from `main`: `Rust/TRADE_STUDY.md`, `Rust/STREAM_BENCHMARK_REPORT.md`, `Rust/PROFILED_STREAM_BENCHMARK_REPORT.md`, `Rust/W5500.md`, and the sources `src/bringup/w5500_offload.rs`, `w5500_macraw.rs`, `w5500_spi.rs` and `src/servers/w5500_offload_udp_echo.rs`. The figures are the repository's own and were not reproduced here.

---

### 01 — What the repository does

One board, one application (UDP echo on port 7), four implementations — and a measurement of all four under the same conditions.

The measurement tooling is purpose-built. `Rust/tools/udp-benchmark/` is a host-side Rust program whose `runner.rs` alone is 38 KB, counting loss, lateness, duplication, reordering, corruption and application RTT separately.

The sources are split for the sake of the measurement. From a bring-up file's header:

> This module ends at a configured hardware UDP socket. The echo algorithm is intentionally kept in `servers/w5500_offload_udp_echo.rs` so its SLOC can be measured separately from network bring-up.

---

### 02 — The four implementations

- **C/LwIP native RMII** — stack in the MCU (LwIP raw API), over the STM32 MAC and a LAN8742A PHY.

- **Rust/Embassy native RMII** — stack in the MCU (Embassy), over the same MAC and PHY.

- **W5500 MACRAW** — stack in the MCU (Embassy); `embassy-net-wiznet` carries raw frames over SPI1.

- **W5500 hardware offload** — stack inside the W5500, driven as hardware sockets through the `w5500-dhcp` / `w5500-hl` crates.

The two W5500 variants use **the same shield on the same SPI wiring**, plugged into the Arduino headers:

| Shield signal | Arduino pin | STM32 pin |
| --- | --- | --- |
| SCLK | D13 | PA5 / SPI1_SCK |
| MISO | D12 | PA6 / SPI1_MISO |
| MOSI | D11 | PB5 / SPI1_MOSI |
| CS | D10 | PD14 / GPIO |
| RESET | RESET | Board NRST |

Checksums land in different places too: the two native RMII variants use the STM32 MAC's hardware offload, MACRAW does them in **MCU software**, and offload does them in **W5500 hardware**.

---

### 03 — The results

#### Code size and static memory

Reproducible through `measure-variants.ps1`. NCLOC excludes blank and full-line-comment lines. Release build, 2026-08-13.

| Variant | Bring-up | Server | Flash | RAM |
| --- | --- | --- | --- | --- |
| C/LwIP RMII | 620 | 29 | 126,824 B | 53,051 B |
| Embassy RMII | 144 | 48 | 59,420 B | 29,488 B |
| W5500 MACRAW | 269 | 48 | 66,160 B | 30,232 B |
| W5500 offload | 262 | 60 | 25,444 B | 3,188 B |

The first two columns are NCLOC.

**Same chip, same shield: offload uses a ninth of MACRAW's RAM and two-fifths of its flash.** MACRAW lifts frames into the MCU for the Embassy stack, so the buffers and stack have to live in MCU memory; with offload they live in the chip.

#### One hour of continuous streaming

100-byte datagrams at 1,000 per second for an hour — 3.6 million in total, measured on profiling builds.

| Variant | Missing | p50 | p99 |
| --- | --- | --- | --- |
| Native RMII | 19 (0.000528%) | 0.303 ms | 0.543 ms |
| W5500 MACRAW | 21 (0.000583%) | 0.674 ms | 1.500 ms |
| W5500 offload | 4 (0.000111%) | 0.448 ms | 0.565 ms |

| Variant | CPU | Stack high-water | Static RAM |
| --- | --- | --- | --- |
| Native RMII | 25.68% | 25,492 B | 20,560 B |
| W5500 MACRAW | 54.06% | 31,404 B | 21,304 B |
| W5500 offload | 34.20% | 920 B | 3,168 B |

This is where the three separate. **Offload loses the fewest packets, has a third of MACRAW's p99, and a thirty-fourth of its stack high-water.** MACRAW spends the most CPU of the three — every frame crosses SPI and its checksums are computed in software.

#### Where raising the rate separates them

The 30-second sweep walks every integer rate from 1 to 20 kHz, where kHz means **datagrams per second**, not a frequency: 3 kHz is 3,000 hundred-byte UDP datagrams sent and echoed back every second.

The unit matters before the numbers do. The report's definition:

> An error event is one missing, late, duplicate, reordered, corrupt, foreign, or host send-error observation. **One packet can contribute to more than one event.**

So the table below is not a count of lost packets. Where the author does break the components out, lateness dominates: MACRAW's 57,630 events at 2 kHz are **1,648 missing and 55,982 late**.

| Send rate | Native RMII | W5500 MACRAW | W5500 offload |
| --- | --- | --- | --- |
| 1 KHz | 0 | 0 | 0 |
| 2 KHz | 0 | 57,630 | 0 |
| 3 KHz | 0 | 89,744 | 0 |
| 4 KHz | 1 | 119,826 | 21,364 |
| 7 KHz | 13,326 | 209,876 | 111,360 |
| 10 KHz | 133,650 | 299,888 | 201,349 |
| 15 KHz | 341,813 | 449,896 | 351,347 |
| 20 KHz | 554,345 | 599,900 | 501,351 |

**Within the same Embassy framework, offload and the MCU's internal MAC are both error-free to 3 kHz.** They separate through the 4–10 kHz band and converge again above it: 341,813 against 351,347 at 15 kHz, and at 20 kHz offload's 501,351 is below the internal MAC's 554,345. All three are saturated there, so the ordering carries little meaning.

What breaks down first is MACRAW. **Lifting whole frames across SPI for the MCU to process in software is the bottleneck** — and driving the same chip through its hardware stack removes it.

The conditions belong with the numbers. This sweep ran on **profiling builds** with the **W5500 SPI clock at 20 MHz** — not the release setting quoted in section 04, where offload runs at 50 MHz. The author also warns against using the reported offload CPU percentage at 4 kHz and above: a continuously-ready task can stay in one executor poll past the Cortex-M7 DWT counter's roughly 10.7-second wrap, undercounting it.

The table also carries C/LwIP native RMII at 1–15 kHz, but that row is not directly comparable: the report marks it **3-second validation**, where the other three are 30-second measurements.

#### What the comparison presupposes

**The STM32H723 has an Ethernet MAC on chip.** A LAN8742A PHY was added to use it, which is why a native RMII row exists at all. On an MCU without a MAC — STM32F103, RP2040, RP2350, nRF52840, SAMD51 — that row does not exist, and the choice is offload against MACRAW, where offload leads on every measure taken.

So what the table really asks is not which is faster but **whether an MCU that has a MAC needs to go past 3 kHz**. If it does, the internal MAC is the better answer. Below that, offload does the same job in 3,168 B of static RAM against the internal MAC's 20,560 B.

---

### 04 — What the W5500 code shows

#### 🔷 Empty UDP payloads are dropped

```
// The W5500 never completes SEND for an empty payload. Drop// it so this hardware limitation cannot wedge the socket.if length == 0 {    warn!("W5500 offload cannot echo a zero-byte UDP payload");    continue;}
```

`W5500.md` records this as a functional difference between the two variants: the MACRAW version can echo an empty datagram, the offload version cannot.

#### 🔷 Destination registers are cached

```
if self.last_source != Some(source) {    unwrap!(device.set_sn_dest(ECHO_SOCKET, &source));    self.last_source = Some(source);}
```

The reasoning in the comment is that a command stream normally comes from one host, so as long as the peer does not change this saves one SPI register transaction per echo.

#### 🔷 The interrupt is cleared before draining

`Sn_IR` is cleared before reading, so a packet arriving between the last receive and the clear is not lost — the later arrival asserts INTn again and wakes the executor.

#### 🔷 Socket allocation and interrupt masks

Sn0 is DHCP, Sn1 the echo, Sn2 profiling (feature-gated). SIMR enables the socket interrupts, and each socket's `Sn_IMR` **unmasks RECV only**. DHCP is serviced only when its own socket interrupts or its maintenance timer is due.

#### 🔷 SPI clocks were settled by measurement

MACRAW runs at 20 MHz in release and a validated 40 MHz in the performance build. Offload runs at 50 MHz in release; the performance build derives a dedicated 80 MHz SPI1 kernel clock from PLL2-P and divides it to 40 MHz.

> Independent 50 MHz SPI and the earlier approximately 65 MHz configuration obtained no DHCP lease or network response with the performance CPU clock.

#### 🔷 The shield reset is never driven

The MACRAW side's `BoardReset` is an empty implementation whose `set_low` and `set_high` both return `Ok(())` — the shield's RESET is tied to the board's NRST, so there is nothing for the driver to hold.

---

### 05 — What was not measured

The C/LwIP one-hour profile is **still blank**. It sits in the table as "Pending one-hour run", and CPU and runtime-stack instrumentation are not implemented on the C side; only its static RAM, taken from the release ELF, is filled in.

There is no TCP. All four variants are UDP echo, and the comparison is only valid within UDP.

Nor is component cost in scope. Using the internal MAC requires a PHY, magnetics and RMII routing including a 50 MHz reference clock, where the W5500 is one SPI device — but BOM, board area and layout effort are not what this repository measures.

The author also bounds the NCLOC metric: it is not a claim that every line carries equal complexity, only a simple auditable project metric, and external crates are excluded from first-party SLOC — their effect is captured separately by binary size and the dependency inventory.

---

### 06 — Similar Projects on WIZnet Makers

- [**How Does a Zephyr Driver Use W5500 Hardware TCP/IP When the In-Tree One Is MACRAW Only?**](https://maker.wiznet.io/Grace_Koo/projects/iot-zephyr-app/) is the companion piece. That one is somebody building an offload path Zephyr lacks; this one is somebody measuring what that choice costs. Where that post notes there is no side-by-side measurement of MACRAW against offload on the same hardware, this repository supplies it.

- [**zweidraehte: A Rust KNX Device Stack with W5500-Based KNX/IP Firmware**](https://maker.wiznet.io/josephsr/projects/zweidraehte-a-rust-knx-device-stack-with-w5500-based-knx-ip-firmware/) shares the foundation — `no_std` Rust on embassy driving a W5500. The difference is purpose: it implements KNX/IP, while this repository implements the same echo four times in order to compare them.

- [**SSH Stamp: Secure Remote UART Access with W6300-EVB-Pico2**](https://maker.wiznet.io/Grace_Koo/projects/ssh-stamp-pull-125/) uses the same crate as this repository's MACRAW variant, `embassy-net-wiznet`. There the shared driver code was modified; here it is used as-is and placed next to the alternative.

---

### Q&A

**Q. Offload or MACRAW?** On this board, below a few thousand datagrams per second, offload wins on every measure taken. Over the one-hour stream: 4 missing against 21, p99 0.565 ms against 1.500 ms, stack high-water 920 B against 31,404 B, static RAM 3,168 B against 21,304 B.

**Q. Then is there a reason to use MACRAW?** The stack is in the MCU, so anything the chip does not know — IPv6, TLS, an arbitrary protocol — can sit on top, and the socket count is not fixed at eight. This comparison covers one UDP echo.

**Q. What if a higher packet rate is needed?** In the same Embassy build, offload and the internal MAC were both error-free to 3 kHz, with about twice the headroom above that on the internal MAC (knee at 6–7 kHz). The 15 kHz figure in the table belongs to the C/LwIP build, and that row alone was a 3-second validation rather than the same measurement.

**Q. What SPI clocks are used?** MACRAW: 20 MHz release, 40 MHz in the performance build. Offload: 50 MHz release, and in the performance build an 80 MHz PLL2-P kernel clock divided to 40 MHz. The report states that 50 MHz and roughly 65 MHz obtained no DHCP lease at the performance CPU clock.

**Q. Why are the C figures blank?** The one-hour run has not been done, and CPU and runtime-stack instrumentation are not implemented on the C side. Only static RAM, from the release ELF, is present.

---

**Original Link**: <https://github.com/mixxen/nucleo-h723zg-udp-echo>

**Trade study**: `Rust/TRADE_STUDY.md` · **One-hour measurement**: `Rust/PROFILED_STREAM_BENCHMARK_REPORT.md` · **W5500 setup**: `Rust/W5500.md`

---

Source: https://maker.wiznet.io/Grace_Koo/projects/nucleo-h723zg-udp-echo/
