Wiznet makers

scott

Published September 07, 2026 ©

146 UCC

20 WCC

52 VAR

0 Contests

0 Followers

0 Following

Original Link

ivan-bohun

A MAC No Factory Ever Issued — How Six ESP32-S3 Boards Share One W5500 Address

COMPONENTS Hardware components

WIZnet - W5500

x 1


PROJECT DESCRIPTION

요약

런던의 한 아파트 선반 위에서 ESP32-S3 여섯 대가 실제 웹사이트를 서비스하고 있습니다. 클라우드도, CDN도, 라즈베리파이도, 리버스 프록시 장비도 없습니다. 이 시스템의 페일오버는 W5500의 MAC 주소가 소프트웨어 레지스터라는 사실에 통째로 기대고 있습니다. 리더로 뽑힌 보드가 공장이 발급한 적 없는 가상 MAC을 자기 W5500에 써넣고, 그 보드가 죽으면 4초 안에 다른 보드가 같은 주소를 이어받습니다.

개요

자기 하드웨어로 자기 웹사이트를 호스팅한다는 말은 요즘 거의 농담처럼 들립니다. IVAN BOHUN 프로젝트는 그 통념을 반증하려고 만들어졌고, 지금 이 순간에도 kyrylonovotarskyi.com이 이 하드웨어에서 응답합니다. 데모가 아니라 운영 중인 사이트입니다.

만든 사람은 런던에 거주하는 Kyrylo Novotarskyi로, 스스로를 Principal Engineer라 밝히고 있습니다. 주목할 것은 그의 GitHub 이력입니다. 2015년에 그는 Netflix의 Eureka를 포크했습니다"탄력적인 미들티어 부하 분산과 페일오버를 위한 AWS 서비스 레지스트리". 그리고 11년 뒤, 같은 문제를 클라우드가 아니라 £12짜리 마이크로컨트롤러 여섯 대 위에서 다시 풀었습니다. 선거·이중 헬스체크·스플릿 브레인 회피 같은 설계가 즉흥이 아닌 이유가 여기 있습니다.

프로젝트 이름은 코사크 대령 Ivan Bohun에서 왔습니다. 페레야슬라우 조약에서 선서를 거부하고 사브르를 부러뜨린 인물로, 저자의 표현을 빌리면 *"누구에게도 호스팅되기를 거부하는 웹사이트"*에 어울리는 수호자입니다.

이 프로젝트를 알게 된 경로

저희는 메이커 소식을 다루는 fabscene의 기사를 통해 이 프로젝트를 처음 접했습니다. 거슬러 올라가면 출발점은 저자가 r/esp32 서브레딧에 올린 *"여섯 대의 ESP32-S3 스웜으로 내 웹사이트를 서비스한다"*는 글이었고, 거기서 여러 매체로 번졌습니다. XDA Developers도 2026년 8월 17일 자로 다루면서 *"ESP32 서브레딧에서 disco-jack이라는 사용자가 ESP32 여섯 대로 서버를 돌리는 방법을 설명했다"*고 출처를 밝히고 있습니다.

같은 시기 Hacker News에도 Show HN으로 올라왔습니다(2026년 8월 15일). 반응은 갈렸습니다. 대시보드를 더 보고 싶다는 호평이 있었던 반면, 웹사이트의 시각 디자인이 AI로 만든 티가 난다는 지적과 접속이 안 된다는 보고도 있었습니다. 후자는 이 프로젝트의 성격을 오히려 잘 보여줍니다 — 마이크로컨트롤러 여섯 대가 감당할 수 있는 동시 접속에는 실제로 천장이 있고, 저자 본인도 그 수치를 공개하며 방문자에게 양해를 구하고 있습니다.

→ fabscene 기사: https://fabscene.com/new/make/esp32-s3-swarm-web-server-vmac-failover/
→ r/esp32 원 게시물: https://www.reddit.com/r/esp32/comments/1vosibt/i_serve_my_website_with_a_swarm_of_six_esp32s3s/
→ XDA Developers: https://www.xda-developers.com/you-too-can-make-this-web-server-that-runs-off-six-esp32s-working-together/
→ Hacker News: https://news.ycombinator.com/item?id=49307171

→ 프로젝트 저장소: https://github.com/Novotarskyi/ivan-bohun
→ 라이브 사이트: https://kyrylonovotarskyi.com

선반 위의 랙 — 스모크 아크릴 뒤의 블레이드 네 대, 켜진 화면 두 개, 빛나는 삼지창

아키텍처

시스템은 역할이 고정되지 않는다는 원칙 위에 서 있습니다. 블레이드 네 대가 모두 같은 펌웨어를 갖고, 그중 하나가 선거로 리더 자리를 맡습니다. 유선은 서비스를 나르고, 무선은 그 유선을 감시합니다.

시스템 토폴로지 — 유선 서빙 경로와 무선 제어 평면

하드웨어 구성

구분내용
블레이드 ×4LILYGO T-ETH-Lite S3 — ESP32-S3 + W5500, 8MB 옥탈 PSRAM
대체 보드Waveshare ESP32-S3-POE-ETH (핀맵 별도 정의, 벤치 검증됨)
디스플레이 ×22.8" 로스터 화면, 4.3" 터치 패널 + 6픽셀 LED 레일
스위치TP-Link SG108 — p1~p4 노드, p5 라우터 업링크
SDKESP-IDF v6.0.2, FreeRTOS
총 소비전력약 22W (팬·디스플레이 포함)
랙 후면 — 배선과 전원 분배

기술 배경

locally administered MAC과 gratuitous ARP

MAC 주소의 첫 바이트에는 두 개의 특별한 비트가 있습니다. 최하위 비트는 유니캐스트/멀티캐스트를 가르고, 그 다음 비트가 1이면 "locally administered" — 제조사가 IEEE에서 할당받은 OUI가 아니라 관리자가 임의로 정한 주소라는 뜻입니다. 이 프로젝트의 가상 MAC은 0x02로 시작합니다. 어떤 제조사와도 충돌하지 않는 안전한 대역입니다.

주소를 옮기는 실제 동작은 gratuitous ARP가 맡습니다. 일반적인 ARP는 "이 IP를 쓰는 사람 누구냐"고 묻지만, gratuitous ARP는 묻지 않고 자기 IP-MAC 짝을 일방적으로 알립니다. 스위치는 그 프레임이 어느 포트로 들어왔는지를 보고 CAM 테이블을 갱신합니다. 같은 MAC이 다른 물리 포트로 이사한 것으로 처리되는 것입니다.

이 두 가지가 이 시스템의 페일오버 전부를 설명합니다. 주소는 그대로 두고 그 주소를 들고 있는 실리콘만 바꾸면, 라우터의 포트포워드도, 방문자의 브라우저도, TLS 인증서도 아무것도 바뀔 필요가 없습니다.

ESP-NOW를 제어 평면으로 쓴다는 것

ESP-NOW는 Espressif의 커넥션리스 무선 프로토콜입니다. AP도, 핸드셰이크도, IP 주소도 없이 MAC 주소만으로 프레임을 주고받습니다. 이 프로젝트는 여기에 암호화를 걸고 1Hz 하트비트를 실어 보냅니다.

핵심은 왜 무선이냐입니다. 선거가 이더넷을 타면, 스위치나 케이블이 죽는 순간 선거 자체가 마비되면서 스플릿 브레인이 생깁니다. 서로를 못 보는 두 진영이 각자 리더를 뽑는 상황입니다. 이 프로젝트는 선거를 유선 밖으로 빼서 그 가능성을 구조적으로 제거했습니다. 저자의 표현으로 "선거는 자신이 보호하는 회선을 타지 않는다" 입니다.

하트비트 프레임 구조

L4 TCP 스플라이싱

일반적인 리버스 프록시는 L7에서 동작합니다. HTTP 요청을 파싱하고, 헤더를 읽고, 다시 조립해 백엔드로 보냅니다. 그러려면 프록시가 TLS를 풀어야 하고, 인증서와 키를 쥐고 있어야 합니다.

L4 스플라이싱은 다릅니다. TCP 연결을 수락한 뒤 바이트를 해석하지 않고 그대로 다른 소켓으로 흘려보냅니다. 프록시는 자기가 나르는 것이 HTTPS인지 무엇인지 알지 못합니다.

이 선택이 MCU 급 자원에서는 사실상 유일한 답입니다. TLS 종단을 백엔드 블레이드가 맡으므로 스플라이서는 인증서를 몰라도 되고, 암호 연산 부하도 지지 않습니다. 리더 한 대에 TLS가 몰리는 병목이 애초에 생기지 않습니다.

스플라이서의 요청 경로

W5500이 이 시스템에서 하는 일

대부분의 프로젝트에서 이더넷 칩의 MAC 주소는 한 번 정하고 잊는 값입니다. 이 프로젝트에서는 런타임에 보드 사이를 옮겨 다니는 자원입니다.

W5500은 SPI로 붙어 있고 espressif/ethernet_init 컴포넌트 v1.3.0을 통해 20MHz로 동작합니다. 핀 배치는 보드마다 다르지만 사용하는 GPIO 집합은 같습니다.

신호LILYGO T-ETH-Lite S3Waveshare ESP32-S3-POE-ETH
SCLKGPIO10GPIO13
MOSIGPIO12GPIO11
MISOGPIO11GPIO12
CSGPIO9GPIO14
INTGPIO13GPIO10
RSTGPIO14GPIO9

마스크가 옮겨가는 순간은 코드 다섯 줄입니다.

esp_eth_stop(s_handle);
esp_eth_ioctl(s_handle, ETH_CMD_S_MAC_ADDR, (void *)mac);   /* W5500 SHAR 재기록 */
esp_netif_set_mac(s_netif, (uint8_t *)mac);
esp_netif_set_hostname(s_netif, hostname);
esp_eth_start(s_handle);

여기서 벌어지는 일을 순서대로 보면 이렇습니다.

  1. 선거에서 이긴 블레이드가 링크를 내립니다. 주소를 바꾸는 동안 프레임이 오가면 안 되기 때문입니다.
  2. W5500의 SHAR 레지스터에 가상 MAC을 덮어씁니다. ETH_CMD_S_MAC_ADDR가 그 일을 합니다. 칩 안의 소스 하드웨어 주소가 이 시점에 바뀝니다.
  3. 링크를 다시 올리고 고정 IP를 설정합니다. lwIP가 gratuitous ARP를 쏘고, 스위치 CAM이 새 포트를 학습합니다.
  4. 주소를 바꾸는 동안에도 서버는 살아 있습니다. HTTP 서버가 0.0.0.0으로 리슨하기 때문에 인터페이스가 갈리는 사이에도 계속 서비스합니다. MCU는 연결을 수락하고 가장 한가한 블레이드로 바이트를 중계하는 일로 돌아갑니다.

짚어둘 것 — 이 프로젝트는 하드웨어 TCP/IP를 쓰지 않습니다

W5500 하면 흔히 하드와이어드 TCP/IP 엔진(TOE) 을 떠올립니다. 칩이 TCP 세션을 직접 물고 MCU는 데이터만 주고받는 방식입니다. 이 프로젝트는 그 방식을 쓰지 않습니다.

여기서 W5500은 MACRAW 모드로 동작합니다. 즉 순수한 NIC(MAC + PHY) 로만 쓰이고, TCP/IP는 ESP32 위의 lwIP가 처리합니다. 앞의 3번에서 gratuitous ARP를 쏜 주체가 칩이 아니라 lwIP인 것도 그 때문입니다.

의도된 선택입니다. 이 시스템이 필요로 하는 것들이 전부 소프트웨어 스택 위에서만 가능하기 때문입니다.

  • TLS 1.3 종단 — 백엔드 블레이드가 mbedTLS로 직접 처리합니다
  • L4 TCP 스플라이싱 — 두 소켓 사이로 바이트를 흘려보내려면 소켓 API가 필요합니다
  • 소켓 수 — lwIP를 최대 48개로 잡아두었습니다. 하드와이어드 소켓은 8개가 상한입니다

대가도 분명합니다. 다음 장에서 다룰 MACRAW 16KB 버퍼 문제가 바로 그것입니다. W5500은 두 방식을 모두 지원하고, 이 프로젝트는 소켓의 자유를 얻는 대신 버퍼 제약을 떠안는 쪽을 골랐습니다.

비유하자면 W5500은 명패를 바꿔 다는 창구입니다. 창구 뒤에 앉은 직원이 바뀌어도 명패의 이름과 창구 번호는 그대로여서, 줄 서 있던 사람들은 무슨 일이 있었는지 모릅니다. 바뀐 것은 누가 앉아 있느냐뿐입니다.

마스크의 생애주기 — 착용, 유지, 반납

안전장치 — 자기 자신을 고립시키지 않는다

이 설계에는 위험이 하나 있습니다. 함대에 접근할 수 있는 유일한 경로가 바로 이 유선이라는 것입니다. 주소를 잘못 바꾸면 보드에 접속할 방법이 사라집니다. 저장소는 그에 대한 규칙을 명시해 두었습니다.

  • 모든 노드는 자기 정체성으로 부팅합니다 — 자기 MAC, 자기 IP. 따라서 언제나 OTA가 가능합니다.
  • 팔로워는 자기 정체성을 유지합니다. 마스크를 쓰는 것은 리더 한 대뿐입니다.
  • 마스크 획득이 제한 시간 안에 주소를 못 받으면 자기 정체성으로 되돌립니다.

이 칩이 아니었다면

ESP32-S3에는 내장 이더넷 MAC이 없습니다. Espressif 공식 문서가 명시합니다 — "ESP32-S3 has no integrated Ethernet MAC ... can only be used with an external ethernet interface such as an SPI-Ethernet device." 초대 ESP32에는 있던 EMAC이 S3에서는 빠졌습니다. 즉 이 보드에서 유선 이더넷의 경로는 외부 SPI 이더넷 칩 하나뿐입니다.

그리고 이 프로젝트가 W5500을 고른 이유는 더 구체적입니다. W5500의 MAC 주소는 SHAR(Source Hardware Address Register)에 소프트웨어로 써넣는 값입니다. WIZnet 문서는 소프트웨어 리셋이 칩 안의 MAC·네트워크 설정을 지우므로 필요하면 저장했다가 다시 복원해야 한다고 안내합니다. 즉 이 주소는 칩에 고정된 것이 아니라 매번 소프트웨어가 정해 주는 값입니다.

보통은 번거로움으로 취급되는 성질입니다. 이 프로젝트는 바로 그 성질을 아키텍처의 축으로 삼았습니다.

→ 참고: https://docs.espressif.com/projects/esp-idf/en/latest/esp32s3/api-reference/network/esp_eth.html
→ W5500 문서: https://docs.wiznet.io/Product/Chip/Ethernet/W5500

MACRAW를 택한 대가

이 저장소에서 가장 값어치 있는 부분은 자랑이 아니라 실패의 기록입니다. 그리고 그 실패는 앞 장에서 본 선택 — 하드웨어 TCP/IP 대신 MACRAW + lwIP — 에서 곧바로 나옵니다.

저자는 TCP 수신 윈도를 키우면 페이지가 한 번에 날아갈 것이라 기대하고 23,040바이트(16×MSS)로 올렸습니다. 결과는 반대였습니다.

단일 요청이 798ms에서 2,263ms로 세 배 느려졌습니다. 그런데 두 코어 모두 80% 유휴였습니다. 이것은 작업량이 아니라 재전송의 징후입니다.

원인을 저장소가 직접 적어두었습니다.

*"W5500은 MACRAW 모드에서 총 16KB를 갖는다. 따라서 23KB 버스트는 NIC 자체를 넘치게 하고, TCP는 RTO 시간 척도로 복구한다. 기본 윈도는 여기서 레거시 기본값이 아니라 이 하드웨어가 실제로 가진 버퍼에 맞춘 값이다."*

그리고 다음에 이 부분을 손댈 사람에게 남긴 경고가 이어집니다 — "단일 요청 지연을 먼저 측정하라. 동시성 상태의 처리량은 이 정체를 가린다."

같은 파일에는 mbedTLS 동적 버퍼를 시도했다가 같은 날 되돌린 기록도 있습니다. 이론은 타당했지만 실측이 받쳐주지 않았습니다. 반면 PSRAM 활성화는 성공했습니다. TLS 레코드 버퍼가 내부 SRAM 밖으로 나가면서 힙 최저치가 약 67KB에서 8.34MB로 올라갔고, HEAPLOW 이벤트가 완전히 사라졌으며, 킵얼라이브 소켓을 4개에서 8개로 늘릴 수 있었습니다.

고가용성 설계

W5500의 주소 이전은 페일오버의 마지막 한 걸음일 뿐입니다. 그 앞에 붙은 규칙들이 시스템을 실제로 살립니다.

  • 선거는 선점형 산술입니다"건강하고 서비스 중인 가장 낮은 id" 가 리더가 됩니다. 협상도, 타임아웃도, 교착될 상태 기계도 없습니다.
  • 부하 분산기는 장비가 아니라 선출되는 역할입니다. 모든 블레이드가 두 프로그램을 다 갖습니다. 백엔드 집합은 매초 라디오 가십으로 다시 만들어지며, 어떤 설정 파일도 백엔드를 지목하지 않습니다.
  • 건강은 양쪽에서 두 번 증명됩니다. 각 블레이드가 7초마다 자기 공개 포트에 스스로 TCP 접속합니다. 실패가 2회 쌓이면 스스로 "서비스 중 아님"을 알려 선거에서 빠지고, 6회에 이르면(약 42초간 서버가 죽어 있는 상태) 코어덤프를 남기고 재부팅합니다. 스플라이서는 이와 별개로 실제로 돌아온 바이트를 기준으로 백엔드를 평가합니다 — "자가 보고 플래그는 거짓말할 수 있기 때문" 입니다.
  • 한 번의 딸꾹질로 두 번 넘기지 않습니다. 물러난 블레이드는 60초 홀드다운 동안 복귀하지 않습니다. 선거에 관성이 없어서, 이것이 없으면 풀이 잠깐 빈 리더가 곧바로 마스크를 되찾아 한 번의 문제로 인계가 두 번 일어납니다.
  • 전부 죽고 하나만 남으면 루프백으로 자기 자신에게 스플라이스하고 계속 서비스합니다.
자가 치유 흐름

배포도 같은 원칙을 따릅니다. OTA는 키로 잠겨 있고 변종을 검사하며, 팔로워를 먼저, 리더를 마지막에 굴립니다. 블레이드의 새 이미지는 90초 연속 서빙 시험을 통과해야 유효한 것으로 표시되고, 5분 안에 그러지 못하면 부트로더가 이전 슬롯으로 되돌립니다.

이더넷이 없는 디스플레이 두 대는 경로가 다릅니다. 터치로 여는 온디맨드 WiFi OTA 창을 쓰고, 시험 조건도 60초 생존 + 5분 마감으로 따로 잡혀 있습니다. 블레이드는 상시 관리 평면을 갖고 있어 이 창이 필요 없습니다. 롤백이 걸린다는 원칙은 같고, 시험 시간은 각자의 역할에 맞춰 다릅니다.

배포 파이프라인

관측도 빠지지 않았습니다. 노드마다 4,096 레코드 블랙박스 링과 코어덤프 파티션이 플래시에 있고, 모든 부팅의 원인이 하트비트에 실려 터치 패널의 연대기에 이름으로 표시됩니다. 저자의 표현으로 "설명되지 않는 재부팅은 없다" 입니다.

유리 화면과 LED 레일 — 멤버당 한 픽셀

실측 성능

저자가 직접 계측해 공개한 수치입니다.

항목
신규 TLS 연결2.9~3.0/s (무손실 기준)
P-256 핸드셰이크약 650ms — 신규 연결 비용의 99%
TTFB약 325ms
웜 킵얼라이브 요청40~60ms, 함대 전체 40~50 RPS
스플라이서 단일 블레이드 한계100 RPS
페일오버 복구4초 이내
총 소비전력약 22W

과부하 상황에서 무너지지 않고 천장에서 서비스하며 초과분을 점진적으로 흘려보내도록 설계돼 있습니다.

OTA 창 — 터치로 열리는 갱신 구간

WIZnet Makers의 유사 사례

How to Register an ESP32-S3 as a Kubernetes Node with W5500 Ethernet? — ESP32-S3와 W5500을 분산 시스템의 정식 멤버로 편입시킨다는 점이 같고, 유선이 그 자격의 전제라는 것도 같습니다. 차이는 조율의 주체입니다. 이쪽은 외부 오케스트레이터에 노드로 등록되지만, 본 프로젝트는 오케스트레이터 자체가 없습니다.
→ 프로젝트 링크: https://maker.wiznet.io/viktor/projects/how-to-register-an-esp32-s3-as-a-kubernetes-node-with-w5500-ethernet/

ESP32 dual-W5500 transparent Ethernet bridge — W5500을 상위 프로토콜이 아니라 L2 프레임 계층에서 다룬다는 점이 같습니다. 다만 이쪽은 칩 두 개로 프레임을 통과시키고, 본 프로젝트는 칩 하나의 주소를 보드 넷이 돌려 씁니다.
→ 프로젝트 링크: https://maker.wiznet.io/gunn/projects/esp32-dual-w5500-transparent-ethernet-bridge/

ESP32-S3 Hybrid Aeroponic Controller with W5500, RS485 and ESP-NOW — W5500 유선과 ESP-NOW 무선을 한 시스템에서 병용합니다. 역할은 반대입니다. 이쪽에서 무선은 센서 노드를 늘리는 수단이고, 본 프로젝트에서 무선은 유선을 감시하는 제어 평면입니다.
→ 프로젝트 링크: https://maker.wiznet.io/josephsr/projects/esp32-s3-hybrid-aeroponic-controller-with-w5500-rs485-and-esp-now/

항목본 프로젝트Kubernetes Nodedual-W5500 bridgeAeroponic
W5500 개수노드당 1 (총 4)121
W5500을 쓰는 층위MAC 주소 자체IP 이상L2 프레임IP 이상
무선 병용제어 평면없음없음센서 확장
조율 주체자족 선거외부 K8s없음중앙 컨트롤러

세 사례 모두 W5500을 연결 수단으로 씁니다. MAC 주소 자체를 옮겨 다니는 자원으로 쓴 사례는 이 프로젝트가 처음입니다.

비즈니스 가치

외부 관점 — 어디에 쓸 수 있는가

  • 엣지 고가용성 — 서버를 둘 수 없는 현장에 22W로 이중화된 서비스 지점을 만듭니다. 산업 현장의 로컬 대시보드나 설비 상태 페이지처럼 끊기면 안 되지만 서버를 둘 수는 없는 자리가 이 구조와 맞습니다.
  • 단일 IP를 유지해야 하는 이중화 — 상위 시스템이 IP나 MAC으로 등록돼 있어 주소를 바꿀 수 없는 환경에서, 하드웨어만 이중화하고 주소는 고정할 수 있습니다. 산업용 프로토콜 다수가 여기 해당합니다.
  • 전력 제약 환경 — 태양광·배터리로 도는 원격 설비에서 리눅스 서버 한 대보다 적은 전력으로 이중화를 구성할 수 있습니다.

내부 관점 — WIZnet에 주는 시사점

이 사례의 값어치는 W5500을 쓴 또 하나의 프로젝트라는 데 있지 않습니다. W5500의 어떤 성질이 설계를 가능하게 했는가를 보여준다는 데 있습니다.

  • 소프트웨어 MAC은 제약이 아니라 설계 여지입니다. 공장 각인 MAC이 없다는 점은 보통 단점으로 읽히지만, 이 프로젝트에서는 주소 이전 기반 페일오버를 가능하게 한 유일한 조건이었습니다. 같은 패턴은 산업 이중화 전반에 적용됩니다.
  • ESP32-S3 생태계에서 SPI 이더넷은 선택이 아니라 전제입니다. 초대 ESP32의 EMAC이 S3에서 사라졌기 때문입니다. ESP32 계열이 S3로 옮겨가는 만큼 이 자리의 수요는 구조적으로 늘어납니다.
  • 3자가 남긴 제약 실측 데이터는 벤더가 직접 만들기 어렵습니다. MACRAW 16KB 버퍼와 TCP 윈도의 관계를 798ms 대 2,263ms라는 숫자로 보여준 문서는 응용 노트로서 가치가 있습니다.

한계 및 개선 방향

훌륭한 구현이지만, 이 구조를 다른 곳에 옮길 때 알아야 할 것들이 있습니다.

  • TLS 핸드셰이크가 천장입니다. 신규 연결 하나에 약 650ms, 3 conn/s를 넘으면 연결이 떨어지기 시작합니다. 저자도 이를 숨기지 않고 수치로 공개하며 방문자에게 양해를 구합니다. 트래픽이 몰리는 사이트에는 맞지 않습니다.
  • 무중단이 아니라 빠른 복구입니다. 4초 동안 신규 연결은 끊깁니다. 진행 중이던 연결은 마스크를 벗기 전에 정리됩니다.
  • vMAC은 L2 범위입니다. 같은 브로드캐스트 도메인 안에서만 유효하고, 라우터 너머로는 옮길 수 없습니다.
  • 스위치의 협조가 필요합니다. gratuitous ARP를 무시하거나 포트 보안으로 MAC 이동을 막는 스위치에서는 동작하지 않을 수 있습니다. 저자는 SG108 한 종류에서만 검증했습니다. [추정 — 다른 스위치 시험 기록 없음]
  • 외부 검증이 적습니다. Fork 3, Issue 0으로 사실상 1인 프로젝트입니다. 저자 스스로 C의 break가 의도한 루프가 아닌 곳에 걸린 결함이 있었다고 밝히고 있으며, 그 뒤 7초 자가 점검과 2회 연속 실패 시 자진 사퇴를 넣었습니다.
  • 동시 접속 천장이 실제로 관찰됩니다. Hacker News 스레드에 접속이 안 된다는 보고가 올라온 적이 있습니다. 트래픽이 몰리면 실제로 도달하는 한계이며, 저자도 이를 예상하고 수치를 공개해 두었습니다.
  • r/esp32 스레드의 반응은 대조하지 못했습니다. 최초 게시물은 확인했으나 댓글 내용은 접근이 차단돼 검증하지 못했습니다. 위에 적은 반응은 Hacker News 쪽에서 확인한 것입니다.

FAQ

Q. 따라 만들려면 난이도가 어느 정도입니까? 하드웨어 자체는 어렵지 않습니다. ESP32-S3 + W5500 보드와 8포트 스위치면 됩니다. 어려운 쪽은 펌웨어입니다. 선거·스플라이서·마스크·OTA·블랙박스가 서로 맞물려 있어 일부만 떼어 쓰기 쉽지 않습니다. 다만 저장소가 MIT 라이선스이고 하드웨어 조립부터 펌웨어까지 문서가 갖춰져 있습니다.

Q. 어떤 스위치에서든 됩니까? 보장할 수 없습니다. 이 기법은 스위치가 gratuitous ARP를 받아 CAM 테이블을 갱신해 준다는 전제 위에 있습니다. 관리형 스위치에서 포트 보안이나 MAC 고정이 켜져 있으면 주소 이동 자체가 차단될 수 있습니다. 도입 전에 해당 스위치에서 MAC 이동을 실제로 시험해 보아야 합니다.

Q. 라즈베리파이 한 대면 될 일 아닙니까? 성능만 보면 그렇습니다. 저자도 "라즈베리파이로 가는 게 아마 신중한 방법이었을 것" 이라고 인정합니다. 다만 이 프로젝트가 보여주는 것은 성능이 아니라 구조입니다. 파이 한 대는 단일 장애점이지만, 이 함대는 어떤 한 대가 죽어도 서비스가 이어집니다. 그리고 그 이중화를 22W로 만들었습니다.

Q. 상용 제품에 이 방식을 쓸 수 있습니까? 페일오버 기법 자체는 표준 기술의 조합입니다 — locally administered MAC과 gratuitous ARP는 모두 규격 안에 있습니다. 다만 스위치 의존성과 L2 범위 제약을 설계에 반영해야 하고, TLS 부하가 큰 용도라면 종단을 어디서 할지 다시 따져야 합니다. 주소를 바꿀 수 없는 이중화가 필요한 자리라면 검토할 값어치가 있습니다.



Summary

On a shelf in a London flat, six ESP32-S3 boards serve a live website. No cloud, no CDN, no Raspberry Pi, no reverse-proxy appliance. The failover in this system rests entirely on one fact: the W5500's MAC address is a software register. The elected leader writes a virtual MAC that no factory ever issued into its own W5500, and when that board dies, another board takes over the same address in under four seconds.

Overview

Hosting your own website on hardware you own sounds close to a joke these days. The IVAN BOHUN project exists to disprove that, and kyrylonovotarskyi.com answers from this hardware right now. It is not a demo — it is a site in production.

The author is Kyrylo Novotarskyi, based in London, who describes himself as a Principal Engineer. His GitHub history is the interesting part. In 2015 he forked Netflix's Eureka"AWS service registry for resilient mid-tier load balancing and failover." Eleven years later, he solved the same problem again, not in the cloud but on six £12 microcontrollers. That is why the elections, the double health checks, and the split-brain avoidance do not read as improvisation.

The project is named after the Cossack colonel Ivan Bohun, who refused the oath at Pereiaslav and broke his sabre instead — in the author's words, the right patron for a website that refuses to be hosted by anyone.

How We Found It

We came across this project through an article on fabscene, which covers maker news. Tracing it back, the starting point was the author's own post on the r/esp32 subreddit"I serve my website with a swarm of six ESP32-S3s" — and it spread from there. XDA Developers picked it up on August 17, 2026, citing the same origin: "Over on the ESP32 subreddit, user disco-jack explained how they managed to get six ESP32s to host a server."

It also reached Hacker News as a Show HN on August 15, 2026. Reception was mixed. One commenter praised it and asked to see more of the dashboard; others said the website's visual design looked AI-generated, and one reported the site was unreachable. That last point illustrates the project rather well — six microcontrollers really do have a ceiling on concurrent connections, and the author publishes that number and asks visitors to be gentle.

→ fabscene article: https://fabscene.com/new/make/esp32-s3-swarm-web-server-vmac-failover/
→ Original r/esp32 post: https://www.reddit.com/r/esp32/comments/1vosibt/i_serve_my_website_with_a_swarm_of_six_esp32s3s/
→ XDA Developers: https://www.xda-developers.com/you-too-can-make-this-web-server-that-runs-off-six-esp32s-working-together/
→ Hacker News: https://news.ycombinator.com/item?id=49307171

→ Project repository: https://github.com/Novotarskyi/ivan-bohun
→ Live site: https://kyrylonovotarskyi.com

The rack on the shelf — four blades behind smoked acrylic, two lit screens, the glowing trident

Architecture

The system rests on one principle: no role is fixed. All four blades carry the same firmware, and one of them wins the election and takes the leader seat. The wire carries the service; the radio watches the wire.

System topology — the wired serving path and the wireless control plane

Hardware

ItemDetail
Blades ×4LILYGO T-ETH-Lite S3 — ESP32-S3 + W5500, 8 MB octal PSRAM
Alternate boardWaveshare ESP32-S3-POE-ETH (separate pin map, bench-proven)
Displays ×22.8" roster screen, 4.3" touch panel + six-pixel LED rail
SwitchTP-Link SG108 — p1–p4 nodes, p5 router uplink
SDKESP-IDF v6.0.2, FreeRTOS
Total power drawabout 22 W (fans and displays included)
Rear of the rack — cabling and power distribution

Technology Background

Locally Administered MAC Addresses and Gratuitous ARP

The first byte of a MAC address carries two special bits. The lowest bit separates unicast from multicast, and the next bit, when set, means "locally administered" — the address was chosen by an administrator rather than assigned from a manufacturer's IEEE OUI. This project's virtual MAC begins with 0x02, a range that cannot collide with any vendor.

Moving the address is the job of gratuitous ARP. A normal ARP request asks who owns an IP; a gratuitous ARP asks nothing and announces its own IP-to-MAC pairing unprompted. The switch notes which port the frame arrived on and updates its CAM table accordingly. As far as the switch is concerned, the same MAC simply moved to a different physical port.

Those two mechanisms explain the entire failover. Leave the address alone and swap only the silicon holding it, and nothing else has to change — not the router's port forward, not the visitor's browser, not the TLS certificate.

Using ESP-NOW as a Control Plane

ESP-NOW is Espressif's connectionless wireless protocol. It exchanges frames using MAC addresses alone — no access point, no handshake, no IP. This project runs encrypted 1 Hz heartbeats over it.

The point is why wireless. If the election rode the Ethernet, then the moment a switch or a cable failed, the election itself would seize up and split the brain — two groups that cannot see each other, each electing its own leader. This project moved the election off the wire and removed that possibility structurally. In the author's words, "the election never rides the wire it protects."

Heartbeat frame structure

L4 TCP Splicing

A conventional reverse proxy works at L7. It parses the HTTP request, reads the headers, and reassembles it for the backend. To do that, the proxy has to terminate TLS and hold the certificate and key.

L4 splicing is different. It accepts the TCP connection and then passes bytes through to another socket without interpreting them. The proxy never learns whether it is carrying HTTPS or anything else.

At MCU scale this is effectively the only workable choice. Because the backend blade terminates TLS, the splicer needs no certificate and carries no crypto load. The bottleneck of piling every handshake onto one leader never forms.

The splicer's request path

Where W5500 Fits

In most projects, an Ethernet controller's MAC address is set once and forgotten. Here it is a resource that migrates between boards at runtime.

The W5500 sits on SPI at 20 MHz through the espressif/ethernet_init component, pinned to v1.3.0. The pin assignment differs per board, though the GPIO set is the same.

SignalLILYGO T-ETH-Lite S3Waveshare ESP32-S3-POE-ETH
SCLKGPIO10GPIO13
MOSIGPIO12GPIO11
MISOGPIO11GPIO12
CSGPIO9GPIO14
INTGPIO13GPIO10
RSTGPIO14GPIO9

The moment the mask changes hands is five lines of code.

esp_eth_stop(s_handle);
esp_eth_ioctl(s_handle, ETH_CMD_S_MAC_ADDR, (void *)mac);   /* rewrite the W5500 SHAR */
esp_netif_set_mac(s_netif, (uint8_t *)mac);
esp_netif_set_hostname(s_netif, hostname);
esp_eth_start(s_handle);

Step by step, this is what happens.

  1. The winning blade brings the link down. No frames may cross while the address changes.
  2. It overwrites the virtual MAC into the W5500's SHAR register. ETH_CMD_S_MAC_ADDR does this. The chip's source hardware address changes at this instant.
  3. The link comes back up and the static address is applied. lwIP emits a gratuitous ARP, and the switch CAM learns the new port.
  4. The server stays alive across the swap. The HTTP server listens on 0.0.0.0, so it keeps serving while the interface changes underneath it. The MCU goes back to accepting connections and relaying bytes to whichever blade is least busy.

Worth Stating — This Project Does Not Use Hardwired TCP/IP

The W5500 is usually associated with its hardwired TCP/IP engine (TOE), where the chip holds the TCP sessions itself and the MCU only moves payload. This project does not work that way.

Here the W5500 runs in MACRAW mode — a plain NIC (MAC + PHY) — and TCP/IP is handled by lwIP on the ESP32. That is why, in step 3 above, the gratuitous ARP comes from lwIP rather than from the chip.

This is deliberate. Everything the system needs exists only on top of a software stack.

  • TLS 1.3 termination — the backend blade handles it directly with mbedTLS
  • L4 TCP splicing — passing bytes between two sockets requires a socket API
  • Socket count — lwIP is configured for up to 48. Hardwired sockets top out at eight

The cost is equally clear. It is the MACRAW 16 KB buffer problem covered in the next section. The W5500 supports both modes, and this project chose the freedom of sockets over the comfort of the buffer.

Think of the W5500 as a service window whose nameplate can be swapped. The clerk behind it changes, but the name on the plate and the window number stay the same, so the people in the queue never learn that anything happened. The only thing that changed is who is sitting there.

Lifecycle of the mask — claiming it, holding it, giving it back

The Safety Rule — Never Strand Yourself

This design carries one danger: the only route to the fleet is that same wire. Set the address wrong and there is no way back into the board. The repository states the rules explicitly.

  • Every node boots into its own identity — its own MAC, its own IP. That means it is always OTA-able.
  • Followers keep their own identity. Only the leader wears the mask.
  • If a mask claim gets no address within the timeout, the node reverts to its own identity.

Without This Chip

The ESP32-S3 has no integrated Ethernet MAC. Espressif's own documentation says so — "ESP32-S3 has no integrated Ethernet MAC ... can only be used with an external ethernet interface such as an SPI-Ethernet device." The EMAC that the original ESP32 carried is gone in the S3. On this board, wired Ethernet has exactly one path: an external SPI Ethernet chip.

The reason for choosing the W5500 specifically is more concrete. The W5500's MAC address is a value written by software into the SHAR (Source Hardware Address Register). WIZnet's documentation notes that a software reset clears the MAC and network settings inside the chip, so they must be saved and restored when one is performed. The address is not fixed in silicon — software decides it every time.

Usually this is treated as an inconvenience. This project made it the axis of the architecture.

→ Reference: https://docs.espressif.com/projects/esp-idf/en/latest/esp32s3/api-reference/network/esp_eth.html
→ W5500 documentation: https://docs.wiznet.io/Product/Chip/Ethernet/W5500

The Price of MACRAW

The most valuable part of this repository is not the boasting — it is the record of what failed. And that failure follows directly from the choice in the previous section — MACRAW plus lwIP instead of hardwired TCP/IP.

Expecting the whole page to fly out at once, the author raised the TCP receive window to 23,040 bytes (16 × MSS). The result went the other way.

A single request slowed from 798 ms to 2,263 ms — three times worse. Yet both cores sat 80% idle. That is the signature of retransmission, not work.

The repository records the cause directly.

"The W5500 holds 16 KB TOTAL in MACRAW mode, so a 23 KB burst overruns the NIC itself and TCP recovers on RTO timescales. The stock window is not a legacy default here; it is matched to the buffer this hardware actually has."

A warning to whoever revisits it follows — "measure single-request latency FIRST; throughput at concurrency hides the stall."

The same file records an mbedTLS dynamic-buffer experiment tried and reverted the same day. The theory was sound; the measurements did not back it. Enabling PSRAM, by contrast, worked. Moving the TLS record buffers out of internal SRAM raised the heap floor from roughly 67 KB to 8.34 MB, eliminated HEAPLOW events entirely, and let the keep-alive socket count go from four to eight.

High-Availability Design

Moving the W5500's address is only the last step of failover. The rules in front of it are what actually keep the system alive.

  • The election is preemptive arithmetic"the lowest healthy, serving id leads." Nothing to negotiate, nothing to time out, no state machine to wedge.
  • The load balancer is an elected role, not a box. Every blade carries both programs. The backend set is rebuilt every second from radio gossip, and no configuration file names a backend.
  • Health is proven twice, from both sides. Every blade TCP-connects to its own public port every 7 seconds. After two failures it reports itself not-serving and drops out of the election; after six — roughly 42 seconds of a dead server — it captures a coredump and reboots. The splicer independently benches backends by the bytes that actually come back, because "a self-reported flag can lie."
  • One hiccup does not cost two handovers. A blade that stood down stays down for a 60-second hold-down. The election has no stickiness, so without it a leader whose pool merely drained would take the mask straight back.
  • If everything but one dies, the survivor splices to itself over loopback and keeps serving.
The self-healing flow

Deployment follows the same principle. OTA is key-gated and variant-guarded, and it rolls followers first, leader last. A blade's new image must survive 90 seconds of continuous serving to be marked valid, and if it fails to within five minutes the bootloader rolls back to the previous slot.

The two displays, which have no Ethernet, take a different path. They use an on-demand WiFi OTA window opened by touch, and their trial is set separately at 60 seconds alive within a five-minute deadline. Blades never need that window because they have a standing admin plane. The rollback contract is the same in principle; the trial times differ by role.

The deployment pipeline

Observability is not an afterthought. Every node keeps a 4,096-record blackbox ring and a coredump partition in flash, and every boot's cause rides the heartbeat to be named in the touch panel's chronicle. In the author's words, "no reboot goes unexplained."

The glass and the rail — one pixel per member

Measured Performance

Numbers measured and published by the author.

MetricValue
New TLS connections2.9–3.0/s with nothing dropped
P-256 handshakeabout 650 ms — 99% of a new connection's cost
TTFBabout 325 ms
Warm keep-alive requests40–60 ms, 40–50 RPS fleet-wide
Single splicer blade ceilingabout 100 RPS
Failover recoveryunder 4 seconds
Total power drawabout 22 W

Under severe backpressure the site is designed not to cascade — it serves at its ceiling and sheds the excess gradually.

The OTA window — the update opening, revealed by touch

Similar Projects on WIZnet Makers

How to Register an ESP32-S3 as a Kubernetes Node with W5500 Ethernet? — it also makes an ESP32-S3 with a W5500 a full member of a distributed system, with the wire as the price of admission. The difference is who coordinates. That project registers as a node with an external orchestrator; this one has no orchestrator at all.
→ Project link: https://maker.wiznet.io/viktor/projects/how-to-register-an-esp32-s3-as-a-kubernetes-node-with-w5500-ethernet/

ESP32 dual-W5500 transparent Ethernet bridge — it also works the W5500 at the L2 frame layer rather than above it. But that project passes frames through two chips, while this one shares one chip's address across four boards.
→ Project link: https://maker.wiznet.io/gunn/projects/esp32-dual-w5500-transparent-ethernet-bridge/

ESP32-S3 Hybrid Aeroponic Controller with W5500, RS485 and ESP-NOW — it runs wired W5500 and wireless ESP-NOW in one system. The roles are reversed: there, wireless is a way to add sensor nodes; here, wireless is the control plane that watches the wire.
→ Project link: https://maker.wiznet.io/josephsr/projects/esp32-s3-hybrid-aeroponic-controller-with-w5500-rs485-and-esp-now/

ItemThis projectKubernetes Nodedual-W5500 bridgeAeroponic
W5500 count1 per node (4 total)121
Layer the W5500 is worked atthe MAC address itselfIP and aboveL2 framesIP and above
Wireless alongsidecontrol planenonenonesensor expansion
Who coordinatesself-contained electionexternal K8snonecentral controller

All three use the W5500 as a means of connection. Treating the MAC address itself as a migrating resource is what this project does first.

Business Value

External View — Where This Applies

  • High availability at the edge — 22 W buys a redundant serving point where no server can be installed. A local dashboard or equipment status page on a plant floor — something that must not go down but cannot justify a server — fits this shape.
  • Redundancy that must keep one address — where upstream systems are registered against a fixed IP or MAC and the address cannot change, you can duplicate the hardware and hold the address still. Many industrial protocols land here.
  • Power-constrained sites — remote installations running on solar or battery can build redundancy for less power than a single Linux server draws.

Internal View — What This Means for WIZnet

The value of this case is not that it is another project using a W5500. It is that it shows which property of the W5500 made the design possible.

  • A software MAC is design headroom, not a limitation. The absence of a factory-burned address usually reads as a drawback; here it was the one condition that made address-migration failover possible. The same pattern generalizes to industrial redundancy.
  • In the ESP32-S3 ecosystem, SPI Ethernet is a premise, not a choice, because the original ESP32's EMAC is gone in the S3. As the ESP32 family moves to the S3, demand for this position grows structurally.
  • Third-party measurements of a chip's limits are hard for a vendor to produce. A document showing the relationship between the 16 KB MACRAW buffer and the TCP window as 798 ms versus 2,263 ms has real value as an application note.

Limitations and Future Improvements

This is a fine implementation, but there are things to know before carrying the structure elsewhere.

  • The TLS handshake is the ceiling. About 650 ms per new connection, and past 3 conn/s connections begin to drop. The author does not hide it — he publishes the number and asks visitors to be gentle. This is not the shape for a high-traffic site.
  • It is fast recovery, not zero downtime. New connections fail for four seconds. In-flight connections are drained deliberately before the mask comes off.
  • The vMAC is L2-scoped. It is valid only within the same broadcast domain and cannot move past a router.
  • The switch has to cooperate. A switch that ignores gratuitous ARP, or that pins MACs through port security, may block the move outright. The author verified this on one switch model, the SG108. [Inferred — no record of other switches being tested]
  • External validation is thin. With 3 forks and 0 issues this is effectively a one-person project. The author himself reports a defect where a C break bound to the wrong loop, after which he added the 7-second self-probe and the stand-down after two consecutive failures.
  • The concurrency ceiling shows up in the wild. A Hacker News commenter reported the site unreachable. Under traffic this limit is real, and the author anticipated it by publishing the numbers.
  • The r/esp32 discussion could not be cross-checked. The original post was confirmed, but the comments were behind blocked access. The reception described above is what was verified on Hacker News.

FAQ

Q. How hard is this to reproduce? The hardware is not the hard part — an ESP32-S3 board with a W5500 and an eight-port switch will do. The firmware is. The election, the splicer, the mask, the OTA path and the blackbox interlock, so lifting one piece out is not simple. That said, the repository is MIT-licensed and documented from assembly through firmware.

Q. Will it work on any switch? No guarantee. The technique assumes the switch accepts a gratuitous ARP and updates its CAM table. On a managed switch with port security or MAC pinning enabled, the move itself may be blocked. Test a MAC migration on your actual switch before committing to this.

Q. Wouldn't a single Raspberry Pi do the job? On performance alone, yes. The author concedes that "going with a Raspberry Pi would have probably been the prudent way of doing it." But what this project demonstrates is not performance — it is structure. One Pi is a single point of failure; this fleet keeps serving when any one member dies. And it builds that redundancy on 22 W.

Q. Could this approach go into a commercial product? The failover technique is a combination of standard mechanisms — locally administered MACs and gratuitous ARP are both in spec. You would need to design around the switch dependency and the L2 scope, and if TLS load is significant, reconsider where termination happens. For a role that needs redundancy without changing the address, it is worth evaluating.

Documents
  • github

  • fabscene

Comments Write