Image Processing Lab
Three image-processing tools in one: apply convolution kernels, explore Bayer sensor demosaicing and analyse histogram and exposure — on your own image or a sample. A demonstrative tool, born from our work on CMOS sensors and embedded camera pipelines.
Image processing lab tool
1 · Image
Three operations, one pipeline
The three modules of this tool are not disconnected: they are three stages of the pipeline that turns light into a usable image. A CMOS sensor captures a raw Bayer mosaic (Demosaic module); demosaicing converts it to RGB; then convolution filters reduce noise or enhance detail (Convolution module); finally you check the histogram and exposure to see whether the scene was captured correctly (Histogram module). These are the same steps that govern the quality of an embedded camera.
What is 2D convolution?
Convolution slides a small matrix of weights — the kernel — over each pixel: the new value is the weighted sum of the pixel and its neighbours. It is the basic operation of nearly every spatial filter and of the first layers of a convolutional network. The R, G and B channels are filtered separately; the alpha channel is left unchanged.
Convolution slides a small matrix of weights — the kernel — over each pixel: a pixel’s new value is the weighted sum of itself and its neighbours. It is the basic operation of nearly every spatial filter and of the first layers of a convolutional network. With an odd-sized kernel K (3×3, 5×5…) and an image I, the output pixel at (x, y) is:
where a = (size−1)/2. The R, G and B channels are filtered separately; the alpha channel is left unchanged.
What are the main kernels?
Identity leaves the image unchanged; box and Gaussian blur average the neighbourhood, reducing noise and detail; sharpen amplifies local contrast; the Laplacian and Sobel X/Y respond to derivatives and highlight contours; emboss simulates a side-lit relief. Zero-sum kernels are not normalized: an offset of 128 is added to make them visible.
- Identity \u2014 centre 1, everything else 0: the image is unchanged. The starting point for building a custom kernel.
- Box blur / Gaussian blur \u2014 all-positive weights that average the neighbourhood: they reduce noise and detail. The Gaussian weights the centre more (1\u20132\u20131 / 1\u20134\u20136\u20134\u20131) and blurs more naturally than the box.
- Sharpen \u2014 high centre, negative edges: amplifies local contrast, enhancing detail. It is an \u201cunsharp\u201d in kernel form.
- Edge detect (Laplacian) \u2014 zero-sum, responds to the second derivative: highlights contours in every direction.
- Sobel X / Sobel Y \u2014 directional gradient (first derivative): detects vertical or horizontal edges. The root of the sum of their squares gives the gradient magnitude, the basis of edge detection.
- Emboss \u2014 an asymmetric kernel that simulates a side-lit relief, turning edges into an embossed effect.
Borders and normalization
At the image borders the kernel overflows: there the coordinates are clamped, i.e. the edge pixel is extended (mirror or periodic modes also exist). Normalization divides the result by the sum of the coefficients, so a blur preserves the average luminance. Zero-sum kernels (edge, Sobel) are not normalized: to make them visible an offset is added (128 here), mapping zero to mid grey.
What is the Bayer sensor?
A colour CMOS sensor does not measure RGB at every pixel: in front of the photosites sits a colour-filter array (CFA) laid out in the Bayer pattern, a 2×2 grid with two greens, one red and one blue. The greens are double because the eye is more sensitive to green. RGGB, BGGR, GRBG and GBRG are the four possible phases.
A colour CMOS sensor does not measure RGB at every pixel. In front of the photosites sits a colour-filter array (CFA) laid out in the Bayer pattern: a repeating 2×2 grid with two greens, one red and one blue. The greens are double because the human eye is more sensitive to green, which carries most of the luminance information. RGGB, BGGR, GRBG and GBRG are the four possible phases of the same scheme: only the starting pixel of the grid changes.
Demosaicing: nearest and bilinear
Since each pixel records a single channel, the two missing ones must be reconstructed by interpolating same-colour neighbours: that is demosaicing. Nearest neighbor copies the closest sample — very fast, but with aliasing and visible 2×2 blocks. Bilinear averages same-colour neighbours — smoother, but on sharp edges it can produce two typical artefacts: zippering (a zig-zag along contours) and colour fringing. The algorithms cameras actually use are more sophisticated: instead of averaging blindly, they detect the direction of contours and interpolate along edges rather than across them, reducing these artefacts.
Histogram and luma
The histogram counts how many pixels fall in each level (0–255), separately for the three channels and for luma. Luma Y′ combines the channels with the Rec.601 weights on the values AS THEY SIT in the file, that is gamma-encoded: it is not photometric luminance, which requires linearizing the channels first. On pure red the difference is 51 levels out of 255.
A cluster on the left means a dark scene, on the right a bright one; a distribution spanning the whole scale signals good contrast.
Exposure and clipping
When the histogram presses against 0 or 255 there is clipping: detail in the shadows or highlights is lost and cannot be recovered in post. The module’s indicators report the percentage of shadow- and highlight-clipped pixels: this is the check made at capture time to set the sensor’s exposure and gain correctly, before any image processing.
From the field to the code
These are not abstract exercises: they are the building blocks of bringing up an embedded camera. In our work on CMOS sensors — from register-level sensor bring-up to video-stream analysis — reading a histogram, judging the demosaicing and applying a sharpening or noise-reduction filter are everyday tasks. The tool this one grew out of (see the related repo) exists precisely to bring a sensor to life and inspect its image.
What are the tool’s limitations?
It is a demonstrative tool: the algorithms (bilinear demosaicing, basic kernels) are didactic, not those of a real ISP. The image is resized to keep processing responsive. It always works on already-demosaiced, compressed data (JPEG/PNG): the Bayer mosaic is simulated from RGB, not a sensor RAW. Convolution runs at 8-bit integer precision with clamping. The kernel is applied without index reversal, so the operation is formally a cross-correlation: that is the image-processing convention, and for symmetric kernels the two results coincide. For antisymmetric kernels the gradient sign flips — with Sobel, light and dark edges swap. A true convolution needs the kernel rotated by 180°.
- Demonstrative tool: the algorithms (bilinear demosaicing, basic kernels) are didactic, not those of a real ISP.
- The image is resized to keep processing responsive; very large images lose detail.
- It always works on already-demosaiced, compressed data (JPEG/PNG): the Bayer mosaic is simulated from RGB, not a sensor RAW.
- Convolution runs at 8-bit integer precision with clamping; it does not replace full-precision floating-point processing.