Image processing: convolution, Bayer, histogram
The fundamentals behind the Image Processing Lab tool: 2D convolution and kernels, the Bayer pattern of CMOS sensors and demosaicing, histogram, luminance and exposure.
Code on GitHub ↗What this wiki covers
The three techniques behind the Image Processing Lab tool: convolution with kernels, demosaicing of a Bayer sensor, and reading the histogram and exposure. They are the elementary building blocks of any image pipeline, from desktop software to an embedded camera.
The 2D convolution
Convolution is the basic operation of spatial filters. A small matrix of weights — the kernel — slides over every pixel; the output value is the weighted sum of the pixel and its neighbours. But a distinction has to be made here that most applied literature skips, and that this tool has a duty to declare, because on antisymmetric kernels it changes the result. With an odd-sided kernel K (a = (side − 1)/2) and an image I there are two operations:
They differ by the sign of the indices: convolution flips the kernel, correlation does not. This tool — like most image-processing libraries, OpenCV’s filter2D included — implements correlation, and calls it convolution out of settled habit.
On symmetric kernels (blur, gaussian, 3×3 Laplacian) the two coincide exactly and the distinction is academic. On antisymmetric ones they do not. On a horizontal ramp, Sobel X gives:
Sobel X kernel [[-1, 0, 1], [-2, 0, 2], [-1, 0, 1]]
correlation (here) 168 208 208 208 208 208 208 168
true convolution 88 48 48 48 48 48 48 88
The two results are mirrored about the offset of 128: 208 − 128 = +80, 128 − 48 = +80. The gradient magnitude is identical, the sign is inverted — that is, a rising edge reads as a falling one. For the magnitude (√(Gx² + Gy²)) nothing changes; for the gradient direction, for optical flow, for any use where the sign matters, everything does. If true convolution is needed, rotate the kernel by 180°.
The R, G, B channels are filtered separately. Two practical details change the result:
- Borders. Where the kernel overflows the image, coordinates are clamped (the edge pixel is extended). Mirror and periodic modes also exist.
- Normalization. Dividing by the sum of the coefficients keeps the average luminance constant: essential for blurs. Zero-sum kernels (edges, gradients) are not normalized; to make them visible an offset (typically 128) is added, mapping zero to mid grey.
Notable kernels
| Kernel | Sum | Effect |
|---|---|---|
| Identity | 1 | no effect (centre 1) |
| Box / Gaussian blur | >0 | averages the neighbourhood → reduces noise and detail |
| Sharpen | 1 | high centre, negative edges → enhances local contrast |
| Laplacian | 0 | second derivative → contours in every direction |
| Sobel X / Y | 0 | directional gradient → vertical / horizontal edges |
| Emboss | 1 | side-lit relief |
The two Sobel operators estimate the horizontal and vertical derivatives; combined, they give the gradient magnitude, the heart of edge detection:
Three 3×3 kernels written out
sharpen Sobel X box blur (×1/9)
0 -1 0 -1 0 1 1 1 1
-1 5 -1 -2 0 2 1 1 1
0 -1 0 -1 0 1 1 1 1
The sharpen has sum 1 (it preserves luminance) but a high centre and negative edges: it amplifies the difference between a pixel and its neighbours. The Sobel X has sum 0 and responds only where intensity changes horizontally (Sobel Y is its transpose). The box blur, normalized to 1/9, is the pure neighbourhood average.
The Bayer sensor and demosaicing
A colour CMOS sensor does not measure RGB at every pixel. In front of the photosites sits a colour-filter array (CFA) laid out in the Bayer pattern: a repeating 2×2 grid with two greens, one red and one blue. The greens are double because the human eye is more sensitive to green, which carries most of the luminance information. RGGB, BGGR, GRBG and GBRG are the four phases of the same scheme.
Since each pixel records a single channel, the other two must be reconstructed: that is demosaicing.
- Nearest neighbor — copies the closest same-colour sample. Very fast, but with aliasing and visible 2×2 blocks.
- Bilinear — averages same-colour neighbours. Smoother, but on sharp edges it can produce two typical artefacts: zippering (a zig-zag along contours) and colour fringing.
The algorithms cameras actually use are more sophisticated: instead of averaging blindly, they detect the direction of contours and interpolate along edges rather than across them, reducing the artefacts.
How much resolution the mosaic costs
The arithmetic is simple and worth doing, because it explains demosaicing artefacts better than any qualitative description. In a 2×2 RGGB block each channel is sampled at a different fraction:
| Channel | Samples in a 2×2 block | Effective pitch | Nyquist (cycles/pixel) |
|---|---|---|---|
| R and B | 1 / 4 = 25 % | 2.00 px | 0.250 |
| G | 2 / 4 = 50 % | 1.41 px | 0.354 |
| monochrome sensor | 4 / 4 = 100 % | 1.00 px | 0.500 |
Green is sampled at twice the rate of the others (it is the channel perceived luminance leans on) but on a diagonal lattice, so its effective pitch is √2 pixels and its Nyquist is 0.354 cycles/pixel instead of 0.5. Red and blue sit at 0.25, half a monochrome sensor.
This is where the typical artefacts come from: a fine coloured detail above 0.25 cycles/pixel is undersampled in R and B, and demosaicing can only interpolate — producing zipper patterns on edges and false colours on regular textures. It is also why an optical anti-aliasing filter sits in front of a Bayer sensor: giving up a little sharpness costs less than having coloured moiré. A monochrome sensor, having no mosaic, has neither the problem nor the filter — which is why many machine-vision applications choose black and white.
Histogram, luma and exposure
The histogram counts how many pixels fall in each level (0–255), separately for the three channels and for a fourth quantity that deserves to be called by its name: luma. It is not luminance, and the difference is not terminological.
Luma Y′ is computed on the RGB values exactly as they sit in the file, that is gamma-encoded: it is a shortcut born in analogue television, where doing the weighted sum before gamma correction saved a circuit. Luminance Y is a photometric quantity: it requires linearizing the three channels first (for sRGB, the EOTF with its 2.4 exponent and the linear segment near zero) and uses the Rec.709 primary weights.
How different are they? On a grey, not at all. On a saturated colour, a lot:
| Colour | Luma Y′ (601) | Relative luminance Y | Y′ of the equal-luminance grey | Gap |
|---|---|---|---|---|
| pure red | 76.2 | 21.3 % | 127.1 | +50.9 |
| pure green | 149.7 | 71.5 % | 219.9 | +70.2 |
| pure blue | 29.1 | 7.2 % | 76.0 | +46.9 |
| yellow | 225.9 | 92.8 % | 246.7 | +20.8 |
| cyan | 178.8 | 78.7 % | 229.5 | +50.7 |
| magenta | 105.3 | 28.5 % | 145.4 | +40.1 |
| grey 128 | 128.0 | 21.6 % | 128.0 | 0 |
| white | 255.0 | 100.0 % | 255.0 | 0 |
The “gap” column is the most direct reading: a pure red has luma 76 but the same luminance as a grey 127. Fifty-one levels out of 255, that is, saturated red shows up in the luma histogram at a fifth of the scale while in luminance it sits at half. On pure green the gap reaches 70 levels. Only the grey diagonal is exact.
The same table shows something else that surprises: grey 128, the “middle” of the digital scale, is 21.6 % of white’s luminance, not 50 %. That is gamma, and it is why “brighten by 50 %” means different things depending on the space you work in.
There is then a third ambiguity, independent of gamma: the weights. Rec.601 (0.299 / 0.587 / 0.114) are the standard-definition television ones; Rec.709, used for HD, weighs differently (0.2126 / 0.7152 / 0.0722). Applied to the same gamma-encoded pixels:
| Colour | Y′ with 601 weights | Y′ with 709 weights | Difference |
|---|---|---|---|
| pure red | 76.2 | 54.2 | −22.0 |
| pure green | 149.7 | 182.4 | +32.7 |
| pure blue | 29.1 | 18.4 | −10.7 |
| yellow | 225.9 | 236.6 | +10.7 |
| cyan | 178.8 | 200.8 | +22.0 |
| magenta | 105.3 | 72.6 | −32.7 |
| grey 128 | 128.0 | 128.0 | 0 |
| white | 255.0 | 255.0 | 0 |
Up to 33 levels of difference on green and magenta — and neither is “wrong”, they are two conventions for two different systems. This tool uses Rec.601 on gamma-encoded values, that is luma Y′ in the strict sense: perfectly suited to its purpose (assessing exposure and clipping at capture time), not to be confused with a photometric measurement.
The histogram shape tells the image’s story: a cluster on the left → a dark scene; on the right → a bright one; a wide spread → good contrast. When the counts press against 0 or 255 there is clipping: detail in the shadows or highlights is lost and cannot be recovered in post. This is the check made at capture time to set the sensor’s exposure and gain correctly — before processing the image.
From the sensor to the code
These are not abstract notions: they are the everyday operations of bringing up an embedded camera. During CMOS sensor bring-up you read the histogram to tune exposure, judge demosaicing quality and apply a sharpening or noise-reduction filter. The tool this one grew out of, arducam-evk-gui, exists precisely to bring a sensor to life at register level and inspect its image.
Limitations
The tool is demonstrative: the algorithms are didactic, not those of a real ISP; the image is resized to stay responsive; the Bayer mosaic is simulated from an already-compressed RGB image, not a sensor RAW; convolution runs at 8-bit integer precision with clamping.
References
- B. E. Bayer, Color Filter Array patent (1976) — the origin of the Bayer pattern.
- I. Sobel, gradient operator for edge detection.
- ITU-R BT.601 / BT.709 — luminance coefficients.
- Try it: Image Processing Lab.
A similar project?
Acoustics, embedded, calculation tools: if you have a related use case, let’s talk.