Sep 29, 2026 / Avik Arefin

Edge AI and TinyML Hardware: ESP32-S3 vs Jetson Orin Nano vs Hailo-8L

The ESP32-S3 fits sensor models under a watt, the Hailo-8L runs real-time camera detection at about 1.5 W, and Jetson Orin Nano runs GPU model at 7–25 W.

Edge AI and TinyML Hardware: ESP32-S3 vs Jetson Orin Nano vs Hailo-8L

Author

Avik Arefin

Published

September 29, 2026

Edge AI and TinyML Hardware: ESP32-S3 vs Jetson Orin Nano vs Hailo-8L

Pick the ESP32-S3 for sound and sensor models on a microcontroller, and the Hailo-8L for real-time camera detection on a Raspberry Pi 5. Pick the Jetson Orin Nano when the model needs a full GPU framework and the device can spend 7–25 W. The sections below explain each board, then the trade-offs between them.

Edge AI and TinyML: what the terms mean

Edge AI means running a trained model on the device that collects the data, instead of sending the data to a server. TinyML is the smallest form of edge AI: models that run on microcontrollers, which are chips with no operating system and memory measured in kilobytes. The three boards in this article sit at three different sizes on that scale, so the right choice depends on what your model reads: a sensor, a microphone, or a camera.

flowchart TD
    A[What does the model read?] -->|Audio or sensor data| B[ESP32-S3]
    A -->|Camera video| C{Need a GPU framework?}
    C -->|No, a supported detector is enough| D[Hailo-8L on Raspberry Pi 5]
    C -->|Yes| E[Jetson Orin Nano]

ESP32-S3: TinyML in 512 KB

The ESP32-S3 is a small, low-cost Wi-Fi and Bluetooth chip from Espressif that runs models without an operating system. Espressif lists 512 KB of internal SRAM (fast working memory) and vector instructions, which process several numbers in one step to speed up neural-network and signal-processing math. Developers reach those instructions through Espressif's ESP-NN and ESP-DSP libraries, and the chip supports TensorFlow Lite Micro, Google's runtime for microcontroller models.

512 KB holds a keyword-spotting or vibration-anomaly model, but not a full camera object detector. A third-party guide reports 5–10 frames per second for image classification with small MobileNet or FOMO models on ESP-NN; that figure is not from Espressif, so treat it as an estimate. Classification answers "is there a person in this frame", while detection also draws a box around each object, and detection is the workload that exceeds this chip.

Choose the ESP32-S3 when the input is audio or a sensor stream, the device runs on battery, and a yes/no or class label is enough.

Jetson Orin Nano: a full GPU at 7–25 W

The Jetson Orin Nano is a small Linux computer from NVIDIA with a graphics processor (GPU) that runs standard deep-learning frameworks such as PyTorch. NVIDIA rates the Jetson Orin Nano Super Developer Kit at 67 TOPS for sparse INT8 math, or 33 TOPS dense. It pairs a 512-core Ampere GPU with 16 Tensor Cores, a 6-core Arm Cortex-A78AE CPU, and 8 GB of LPDDR5 memory at 102 GB/s. The kit launched in December 2024 at $249 and runs in 7 W, 15 W, or 25 W power modes.

8 GB of shared memory holds larger detectors, several models at once, or a small language model, which the ESP32-S3 and Hailo-8L cannot. Its software, JetPack 6.2, runs on Ubuntu 22.04 with CUDA 12.6, TensorRT 10.7, and cuDNN 9.6, NVIDIA's libraries for running and optimizing models on the GPU. The cost is power: the lowest mode, 7 W, is several times the Hailo-8L's 1.5 W typical draw.

The "Super" figures come from a software update, not new silicon: JetPack 6.2 added a 25 W mode and an uncapped MAXN SUPER mode to existing Orin Nano modules. An Orin Nano bought before December 2024 gains the higher figures after the update.

Choose the Jetson Orin Nano when the model uses layers an accelerator compiler will not accept, or when one device must run several models, and the power supply can deliver 7–25 W.

Hailo-8L: 13 TOPS at about 1.5 W

The Hailo-8L is a dedicated AI accelerator chip that does only neural-network math and needs a host computer to feed it. Hailo rates it at 13 TOPS (trillions of operations per second, a rough measure of AI speed) at a typical 1.5 W. Raspberry Pi sold it from June 2024 as the AI Kit, an M.2 module on a HAT+ board that connects to the Raspberry Pi 5's single PCIe lane, for $70 at launch.

That power budget makes real-time camera detection practical. One published test reports YOLOv8s, a common small object detector, at 25–35 frames per second at 640×640 pixels on the Hailo-8L, against 2–3 frames per second on the Pi 5's CPU alone. The trade-off is the toolchain: models must be compiled into Hailo's own format, so you cannot load an arbitrary PyTorch model the way you can on a GPU. The larger Hailo-8 (26 TOPS) handles bigger models or several at once.

Choose the Hailo-8L when a Raspberry Pi 5 is acceptable as the host and the model is a supported detector or classifier.

ESP32-S3 vs Jetson Orin Nano vs Hailo-8L: side by side

ESP32-S3 Jetson Orin Nano Super Hailo-8L (on Pi 5)
What it is Microcontroller Linux computer with GPU Accelerator chip, needs a host
Compute (vendor figure) Vector instructions, no TOPS rating 67 sparse INT8 TOPS (33 dense) 13 TOPS
Power (vendor figure) Not compared here 7 W, 15 W, or 25 W modes ~1.5 W typical, plus the Pi
Software TensorFlow Lite Micro, ESP-NN JetPack 6.2, Ubuntu 22.04, CUDA 12.6 Hailo compiler and runtime
Fits Audio, sensors, small classifiers Any GPU model, larger detectors, several models at once Real-time camera detection

NVIDIA's 67 TOPS counts sparse operations, which assume half the model's weights are zero; the dense 33 TOPS is the figure closer to Hailo's 13 TOPS rating.

CortexTech's perspective

Over at CortexTech, our most interesting edge AI build used two Espressif chips, an ESP32-P4 and an ESP32-C6, instead of a Hailo-8L or a Jetson Orin Nano. The device was a battery-powered camera that counts objects in a scene that changes over minutes, not milliseconds. A detection every few seconds was enough, so the Hailo-8L's 25–35 frames per second, and the Pi 5 it needs as a host, bought speed we would not use.

The ESP32-S3 was too slow for camera detection. One published comparison reports 6–26 seconds per frame for YOLO models on the S3. We moved to the ESP32-P4, Espressif's fastest chip: Espressif lists a dual-core 400 MHz RISC-V processor, 768 KB of internal SRAM, a MIPI-CSI camera interface, and support for external PSRAM. An open-source YOLO port on the P4 reports 1.7–2.4 seconds per frame with ESP-DL, Espressif's deep-learning library, which fits a scene that changes over minutes.

The problem was that the ESP32-P4 has no radio, so it cannot send its counts over Wi-Fi on its own. We followed the design Espressif uses on its own ESP32-P4-Function-EV-Board: a second chip, the ESP32-C6, handles Wi-Fi 6 and Bluetooth LE. The C6 runs Espressif's ESP-Hosted firmware and talks to the P4 over SDIO, a fast chip-to-chip link, so the P4 uses Wi-Fi as if it had its own radio.

The split also let each chip do one job. The P4 spends its cycles on the camera and the detector, while the C6 runs the network stack and sends only counts, never images, so little data leaves the device. The lesson we reuse is to match the frame rate to how fast the scene changes. When a detection every few seconds is enough, two microcontrollers can replace an accelerator and its Linux host.

Sources

Want to read more?

More engineering deep dives from the Cortex R&D team.

All Articles