Edge AI & TinyML on the ESP32, Jetson Nano, and Hailo-8L: Choosing the Right Tier
Edge AI splits into three tiers defined by memory and power budgets, and the right chip depends on the application
Author
Avik ArefinPublished
September 29, 2026
Edge AI & TinyML on the ESP32, Jetson Nano, and Hailo-8L: Choosing the Right Tier
Three class of devices, three different jobs. TinyML on the ESP32-S3, the Hailo-8L accelerator class, and the NVIDIA Jetson family Orin Nano Super for you to pick the silicon that fits the workload instead of the marketing.
Edge AI is not one thing
Running a neural network on a camera feed and running one on a vibration sensor are both "edge AI." They have almost nothing else in common. The dividing line is the memory and power budget.
| Tier | Example | Memory | Power | Model scale |
|---|---|---|---|---|
| TinyML MCU | ESP32-S3 | 512 KB SRAM, 8 MB PSRAM | under 1 W | tens to hundreds of KB |
| Accelerator + host | Hailo-8L on a Pi 5 | host DRAM + 8 MB-ish on-chip context | 1–1.5 W for the NPU | INT8 vision models |
| Embedded GPU | Jetson Orin Nano Super | 8 GB LPDDR5, 102 GB/s | 7–25 W | transformers, LLMs, VLMs |
TinyML is the discipline of running inference on microcontrollers under roughly 256 KB of RAM and 50 mW of active power, with latency targets near 50 ms (emergentmind). A canonical TinyML deployment fits in under 1 MB of flash, under 256 KB of SRAM, and burns 1–100 µJ per inference. That is a different universe from a single-board computer with multi-megabyte DRAM and a gigahertz-class CPU. If you blur those two, you will pick the wrong chip.
Tier 1: TinyML on the ESP32
The constraint is memory, not ambition
Espressif's own comparison puts the ESP32-S3 at 512 KB of SRAM and 8 MB of PSRAM against a cloud server's gigabytes (Espressif). The ESP32-S3 uses dual-core Xtensa LX7 at 240 MHz and draws under 1 W. Everything about the workflow follows from that gap.
The first lever is quantization. Moving weights from 32-bit floats to 8-bit integers cuts memory use by about 4x and enables integer-only multiply-accumulate on the MCU (emergentmind). The second lever is the instruction set. The ESP32-S3 adds 128-bit SIMD vector instructions to the LX7 core; Espressif reports roughly an 18x speedup over unoptimized code for INT8 operations.
flowchart LR
A[Train in PyTorch / TF] --> B[Export to ONNX]
B --> C[Quantize with ESP-PPQ FP32 to INT8]
C --> D[.espdl model file]
D --> E[ESP-DL runtime on device]
The ESP-DL and ESP-WHO toolchain
ESP-DL is Espressif's inference engine. It loads a FlatBuffers-based .espdl model, supports zero-copy deserialization, and runs operators on the SoC's SIMD instructions (Espressif). Two details matter for real projects. A static memory planner places each layer in either internal SRAM or PSRAM, so you control the trade-off between speed and capacity. Heavy operators like Conv2D are split across both cores automatically.
ESP-WHO sits on top and adds the camera pipeline, display integration, and FreeRTOS scaffolding. It ships pre-quantized models for face detection, face recognition, gesture recognition, and YOLO11-based object detection. The same application code runs on the ESP32-S3-EYE and the ESP32-P4 board by switching the board support package.
What actually fits
On the ESP32-S3, image-based models take hundreds of milliseconds per inference (electroniccomponent). That suits periodic sensing, not real-time video. The workloads that shine are the small ones: keyword spotting at roughly 12 KB per model, gas sensing under 2 ms, and vibration-based predictive maintenance at 32 KB of weights (emergentmind).
If a vision workload still needs more headroom, the ESP32-P4 is the current step up. Its dual RISC-V cores at 400 MHz add PIE extensions that fuse multiply-accumulate-shift into one cycle on INT8 vectors, which Espressif reports as roughly 2.7–3x faster than the S3 on typical vision models. It also adds MIPI-CSI and up to 32 MB of PSRAM.
Tier 2: the Hailo-8L and the accelerator class
Not every job should run on the main CPU. A dedicated neural accelerator offloads one model class cheaply and efficiently, and this is where the Hailo-8L earns its place as the most relevant edge AI device in 2026.
The Hailo-8L delivers about 13 TOPS of INT8 compute at 1–1.5 W, usually bought as the Raspberry Pi AI Kit or AI HAT+ in an M.2 form factor (edgeaistack). It supports a broader, newer set of models through the actively maintained HailoRT toolchain (mustafa.net).
The natural comparison is the Google Coral Edge TPU, which was the default for years. Coral runs fully INT8-quantized TensorFlow Lite models within roughly 8 MB of on-chip weight storage, at 4 TOPS and about 2 W (edgeaistack). Google has shipped no new Coral hardware since 2022, and its ecosystem is winding down. For an existing fleet with a stable model set, Coral still works. For a new build, the Hailo-8L is the safer bet at a similar price.
Both are detection-focused accelerators, and that shapes where they belong. They are excellent at real-time object detection on camera streams. They are not general-purpose ML backends; neither one accelerates a photo-stack face-recognition pipeline that expects CUDA or OpenVINO. Pick the accelerator for the job it was built for.
Tier 3: the Jetson line, from Nano to Orin Nano Super
When you need to run a transformer, a vision-language model, or several video streams at once, you leave the MCU and accelerator tiers behind and move to an embedded GPU.
The original Jetson Nano defined the entry point: a 128-core Maxwell GPU at 472 GFLOPS, a 4-core CPU, 4 GB of LPDDR4, and a 5–10 W power window (NVNexus). It was the $99 board that put CUDA-accelerated inference in reach of hobbyists. It is now a baseline to compare against, not a current default.
The Jetson Orin Nano Super is where that line landed. NVIDIA lists 67 INT8 TOPS — a 1.7x gain over the 40 TOPS of its predecessor — with an Ampere GPU of 1024 CUDA cores and 32 tensor cores, a 6-core Arm Cortex-A78AE CPU, 8 GB of LPDDR5 at 102 GB/s, and a 7–25 W power range (NVIDIA). Existing Orin Nano Developer Kits get the same boost from a software update alone. At this tier you can run large language models, vision-language models, and vision transformers locally.
Which tier do you need?
The deciding question is not accuracy. It is whether the model, its activations, and the input pipeline fit the memory and power budget at the required latency. Work backward from the largest number, because that one is usually the camera or the language model.
flowchart TD
S[Start] --> Q1{Does the model fit under 256 KB RAM and run under 50 ms?}
Q1 -- Yes --> T1[ESP32-S3 / ESP32-P4 TinyML]
Q1 -- No --> Q2{Is the job a fixed INT8 vision model on camera streams?}
Q2 -- Yes --> T2[Hailo-8L accelerator with a host SBC]
Q2 -- No --> Q3{Do you need transformers, LLMs, or multiple streams?}
Q3 -- Yes --> T3[Jetson Orin Nano Super]
Q3 -- No --> T2
A useful rule from the accelerator comparison: plan one full-rate 1080p detection stream per Coral-class device, and treat the accelerator as a co-processor, not the system. Video decode stays on the host.
What this looked like at CortexTech
We hit this decision on a client project, a mixed factory-and-retail deployment that started life as one do-everything camera box. The brief was to flag safety-vest violations at a loading bay and to catch a bearing failure on a conveyor from vibration and there was preference for no cloud round-trip.
Instead of a Single 25W Jetson Nano, the architecture was split into two parts. The vibration monitor became an ESP32-S3 running an 8-bit 1D CNN — small weights, millisecond inference, single-digit microjoules per sample. It spends the day asleep and wakes on an accelerometer threshold. The camera stream moved to a Raspberry Pi 5 with a Hailo-8L doing the vest detection, because that workload is exactly the fixed INT8 vision job the accelerator is built for. We kept the Jetson Orin Nano Super, but only for the shift-report summarizer that runs a small language model on the day's flagged events, and it sleeps between runs.
The lesson was not "the Jetson is overkill" or "TinyML is enough." It was that we had been matching devices to a product name instead of to a memory and power budget. The ESP32-S3, the Hailo-8L, and the Orin Nano Super each ended up doing exactly one thing.
Sources
- Edge-AI with ESP32-S3 Workshop: Introduction — Espressif Developer Portal
- TinyML: Edge AI on Microcontrollers — Emergent Mind
- Jetson Orin Nano Super Developer Kit — NVIDIA
- NVIDIA Jetson Nano 4GB — NVNexus
- Google Coral Edge TPU — Specs, Sizing & Deployment Fit — EdgeAIStack
- Hailo-8L vs Google Coral: The Homelab AI Accelerator Comparison (2026) — Mustafa.net
- Embedded AI: Running TinyML on ESP32-S3 and Other Microcontrollers — Electronic Component
- Why isn't this being updated? — google-coral/edgetpu issue #842