ST has recently released the STM32N6 family: a high-performance MCU that integrates onto a single silicon a Cortex-M55 with Helium extensions (128-bit SIMD), a Neural-ART NPU rated at 600 GOPS and 3 TOPS/W, an ISP, hardware media encoders and over 4 MB of internal RAM. On paper, it enables a class of products that until yesterday required two separate chips or a Linux-based SoC: imaging with on-board processing, deterministic, low-power.
For us, it’s interesting because we work on medical imaging with embedded processing, and this is exactly the trajectory we want to follow: bringing classical processing and neural inference into the same MCU, without resorting to Linux. STM32N6 is the first ST silicon that makes this plausible.
Why start with a UVC webcam
When a new silicon arrives, the temptation is to run a synthetic benchmark — something that exercises a single feature and produces a number. We chose the opposite road: the most representative bench possible of a real pipeline, even at the cost of not having “clean” numbers to publish.
A UVC webcam touches two of the silicon’s most demanding subsystems: the hardware H.264 encoder and the USB high-speed transport. Both have to work together with serious timing: the encoder must produce compressed frames at a steady rhythm, the USB endpoint must deliver them without losing packets. Upstream, there’s a real sensor connected via MIPI-CSI2, with an ISP that handles debayering and color processing on-chip. Three subsystems talking to each other: if one loses sync, the whole throughput collapses.
Above all, a complete imaging pipeline modifiable to add processing — classical or NPU — is an infinitely more useful base than a benchmark that doesn’t reproduce a real scenario.
The pipeline
The end-to-end flow is linear:
IMX335 → MIPI-CSI2 → ISP → RAM buffer → H.264 HW encoder → USB UVC
Orchestration is zero-copy with DMA, scheduled by FreeRTOS. Raw frames from the sensor arrive in RAM via MIPI-CSI2 and DMA; the ISP processes them in place (debayering, gain, AWB); the hardware H.264 encoder reads from the system buffer and writes the compressed stream into a second buffer; the stream is finally consumed and served as a UVC endpoint to the PC through USBX, the USB stack from Microsoft Azure RTOS integrated into ST CubeMX. No hidden memcpys: each stage writes where the next one reads.
On the silicon, in this demo, we actually use ISP, hardware H.264 encoder and DMA-2D (for block movements in buffer orchestration). NPU and Helium SIMD stay off: they aren’t needed for a webcam, they’re territory for the next phase.
Effective throughput in the current setup: 720p at 100 fps, glass-to-USB latency of about 50 ms, H.264 bitrate of about 3 Mbit/s. These are numbers that give a sense of scale — they aren’t the silicon’s limit, they’re the numbers of the configuration we used. The Nucleo has no external PSRAM, so we work entirely within the 4.2 MB of internal RAM, which is the constraint that decides how many frame buffers we can keep in flight at once.
What was easy and what wasn’t
No technical surprises in how the silicon behaves: the official ST demo runs on EVAL and works well. Our work was adapting it to Nucleo — a more accessible board, more limited peripherals, no external PSRAM. What ST documents about the silicon, the ISP and the encoder, in production behaves as advertised.
The real difficulty was elsewhere: there are no official ST examples for the UVC H.264 case on Nucleo specifically. For the EVAL there’s a complete reference demo; for the Nucleo you have to rebuild the peripheral configuration from pieces scattered across various example projects, and change pinning where the Nucleo doesn’t expose the same lines. The work is largely carpentry — figuring out what the original demo takes for granted that has to be reconfigured on the new board — rather than silicon exploration. It’s the kind of work that’s worth doing once, though, and it leaves you with a solid base for everything you build on Nucleo-N657X0-Q afterwards.
What comes next
This repo is a working MVP: it covers the UVC H.264 case end-to-end, it’s verifiable in minutes by opening it as a webcam in OBS or Zoom, and it serves as a base for the explorations that follow. The logical next step is turning on the Neural-ART NPU for on-camera AI pipelines, on a real medical-imaging use case — not in the abstract.
What decided the result
You evaluate new silicon not by reading the datasheet but by finding the simplest bench that exercises its real capabilities. For STM32N6 that bench was a UVC webcam: it’s the minimum exercise that puts ISP, hardware encoder and DMA in motion at the same time — the three things that distinguish this silicon from the rest of the STM32 family. Everything we’ll want to build on top of it, NPU included, goes through those three blocks. Verify they actually work before everything else.
The firmware is open source, MIT-licensed, and lives on GitHub: x-cube-n6-camera-capture-nucleo. For the full technical documentation, see the dedicated wiki page.
Are you evaluating STM32N6 for an embedded imaging project, or working on imaging pipelines with on-device AI? Let’s talk.