// WIKI

Bandwidth of a MIPI CSI-2 camera interface

How to size the link between a camera sensor and a SoC: from pixel rate to data rate, D-PHY lane utilization, pixel formats and packetization overhead.

Published on Updated on MIPICSI-2D-PHYCameraImaging
Code on GitHub ↗

When you connect a camera sensor to a microcontroller or a SoC, before choosing how many lanes to use and how fast to run them you must answer one simple question: how many bits per second does the sensor produce, and can the interface carry them? This tool answers that question; this page explains the calculation.

What MIPI CSI-2 is

CSI-2 (Camera Serial Interface 2) is the protocol by which, in the vast majority of modern devices, an image sensor talks to the processor. It defines how pixels are packetized and marked (line start/end, frame start/end, virtual channels), but not how they physically travel on the wire: that is the job of the physical layer, usually D-PHY.

On D-PHY data flows over differential lanes in parallel — typically 1, 2 or 4 — alongside a clock lane. Each lane runs at a given lane rate (e.g. 800, 1500, 2500 Mbps). The overall bandwidth is the sum of the lanes.

How the bandwidth is computed

The raw video stream is a continuous jet of pixels. The required bandwidth is the product of how many pixels per second and how much each pixel weighs:

An example: 1920×1080 at 30 fps in RAW10. The pixel rate is 1920 · 1080 · 30 ≈ 62.2 Mpixel/s; at 10 bits per pixel the data rate is ≈ 622 Mbps. On 2 lanes at 1500 Mbps the capacity is 3000 Mbps, so utilization is ≈ 21%: plenty of headroom.

How to read the utilization:

  • below ~80% — there is headroom;
  • 80–100% — it is tight (little room for blanking, overhead, tolerances);
  • above 100% — the lanes cannot carry the stream: you need more lanes, a higher lane rate or a lighter format.

That 80 % threshold is neither arbitrary nor taken from the specification: it is the sensor’s blanking read backwards, as shown further down.

The opposite case: 3840×2160 at 60 fps in RAW12. The pixel rate is 3840·2160·60 ≈ 498 Mpixel/s; at 12 bits/pixel the data rate is ≈ 6.0 Gbps. On the same 2 lanes at 1500 Mbps (3000 Mbps) utilization would be ~199% — impossible; you need 4 lanes at 1500 Mbps (6000 Mbps), and even then you are at ~99.5%, at the edge. This is the typical point where you raise the lane rate or move to a lighter format.

Pixel formats and bits/pixel

For a given resolution and frame rate, the format is what changes the bandwidth:

Format bits/pixel Notes
RAW8 / 10 / 12 / 14 8 / 10 / 12 / 14 sensor Bayer output, before the ISP
YUV422 16 luma + subsampled chroma
RGB888 24 full RGB, after the ISP

Going from RAW10 to RGB888 nearly triples the bandwidth (from 10 to 24 bits/pixel): this is why high-resolution pipelines prefer to transfer the RAW and process it downstream.

Overhead: two different mechanisms, often conflated

The pixel data rate is the floor, not the ceiling. But the two things that raise the ceiling are different in nature, and worth keeping apart because they weigh very differently.

Packetization, computed: less than 1 %

A CSI-2 long packet carries a 4-byte Packet Header (Data Identifier, a two-byte Word Count, ECC) and a 2-byte Packet Footer checksum; to those add the Frame Start and Frame End short packets (4 bytes each) and, if enabled, the Line Start and Line End ones. A line’s payload is width · bpp / 8 bytes, so:

The control bytes are a per-line constant while the payload grows with width: the overhead is therefore inversely proportional to horizontal resolution, and on real formats it is tiny.

Resolution Format Line payload Overhead without LS/LE With LS/LE
160×120 RAW8 160 B 3.792 % 8.792 %
640×480 RAW8 640 B 0.940 % 2.190 %
1280×720 RAW10 1600 B 0.376 % 0.876 %
1920×1080 RAW10 2400 B 0.250 % 0.584 %
1920×1080 RGB888 5760 B 0.104 % 0.243 %
2592×1944 RAW10 3240 B 0.185 % 0.432 %
3840×2160 RAW12 5760 B 0.104 % 0.243 %

On 1080p RAW10 packetization is worth 0.25 % — or 0.58 % with line sync enabled. It only becomes significant where the payload is tiny: at 160×120 RAW8 it is 3.8 %, and 8.8 % with LS/LE. That is the case of low-resolution high-frame-rate vision sensors, where the bottleneck is usually the LP time between packets rather than the bytes.

So: packetization is not 10–20 %. Anyone quoting that figure is talking about something else.

Blanking, which is the real 10–30 %

Blanking does not add bytes: it squeezes the same bytes into a shorter window. A sensor transmits a line during the active time, then goes quiet for the horizontal blanking; at the end of the frame it goes quiet for the vertical blanking. The link, however, has to carry the instantaneous rate during the active line, not the average over the frame.

With b_h and b_v the relative blanking figures (the sensor’s line_length_pck and frame_length_lines registers divided by the active area, minus one), the active line lasts 1/((1+b_h)(1+b_v)) of the time it would take without blanking, so:

Typical CMOS sensor values sit between 5 % and 30 % per axis:

Horizontal blanking Vertical blanking Peak / average Maximum average utilization
5 % 3 % 1.082 92.5 %
10 % 5 % 1.155 86.6 %
12 % 10 % 1.232 81.2 %
15 % 8 % 1.242 80.5 %
14 % 13 % 1.288 77.6 %
20 % 15 % 1.380 72.5 %
30 % 20 % 1.560 64.1 %

And this is where the 80 % threshold stops being a convention: 80 % average utilization corresponds exactly to a peak/average ratio of 1.25, that is about 12 % blanking on each axis — an utterly ordinary sensor. The threshold is not generic caution: it is typical blanking, read backwards. With a wider-blanking sensor (14 %/13 %, ratio 1.288) the ceiling drops to 77.6 %, and with a narrow one (5 %/3 %) it rises to 92.5 %.

What to put in the overhead field

The tool exposes the net figure (pixels only) and, if you set an overhead, the effective one too, and it bases utilization on the latter. So the number to put there is the blanking, not a packetization guess: (1+b_h)(1+b_v) − 1, which for the 14 %/13 % case is 28.8 %. Packetization, if you want to be pedantic, goes on top and is worth a few tenths of a point.

When you need more lanes

If utilization exceeds the threshold, you have three levers:

  1. More lanes — doubling from 2 to 4 lanes doubles the capacity. It is the most direct lever, if the sensor and SoC expose them.
  2. Higher lane rate — newer D-PHY reach higher lane rates; often, though, it is the sensor or the PCB that sets the limit. And the lane rate is not continuous: it comes from a PLL with integer dividers, so the values actually selectable are a grid. Asking for “1500 Mbps” may mean getting 1440 or 1584, which is one more reason not to design at the edge of capacity.
  3. A lighter format — dropping bit depth or subsampling reduces the data rate at the source.

The tool also computes the minimum number of lanes to stay below the threshold at that lane rate, so you immediately see how much interface you need.

A few cases, net and with blanking

The same links as above, computed by the engine, with and without the 28.8 % blanking of the 14 %/13 % case:

Case Net With 14/13 % blanking Capacity Utilization Verdict
1920×1080 30 fps RAW10, 2 lane @ 1500 622 Mbps 801 Mbps 3000 Mbps 26.7 % ok
1920×1080 60 fps RAW10, 2 lane @ 1500 1244 Mbps 1603 Mbps 3000 Mbps 53.4 % ok
1280×720 100 fps RAW10, 2 lane @ 1500 922 Mbps 1187 Mbps 3000 Mbps 39.6 % ok
2592×1944 30 fps RAW10, 2 lane @ 1500 1512 Mbps 1947 Mbps 3000 Mbps 64.9 % ok
3840×2160 30 fps RAW12, 4 lane @ 2500 2986 Mbps 3847 Mbps 10000 Mbps 38.5 % ok
3840×2160 60 fps RAW12, 4 lane @ 2500 5972 Mbps 7693 Mbps 10000 Mbps 76.9 % tight

The last row is the interesting one: 4K at 60 fps in RAW12 over 4 lanes at 2500 Mbps sits at 59.7 % on the net figure — reassuring — but at 76.9 % with blanking, one step away from the threshold. That is exactly the case where reasoning on the net figure alone leads to a design that does not work.

The limits of the calculation

It is a bandwidth calculation, not a timing one. It tells whether the stream fits in the lanes; it does not validate timing, clock training, lane skew, equalization, nor the distinction between continuous and discontinuous clock. It does not model the ISP nor compression, and it uses only the active area: blanking is not inferred from the sensor registers but has to be entered by hand in the overhead field, as explained above — and it accounts for 20–30 %. For real camera-link sizing the sensor and SoC datasheets and the MIPI specification remain indispensable.

And if the PHY is C-PHY

The calculation assumes D-PHY, where each lane carries one bit per bit time and the capacity is the trivial sum of the lanes. C-PHY works differently: it operates on trios rather than differential pairs and encodes more than one bit per symbol — 16 bits every 7 symbols, that is 2.286 bits/symbol. A trio at a given symbol rate therefore carries more than the number suggests:

C-PHY trios Symbol rate Equivalent bit rate
1 2500 Msps 5714 Mbps
2 2500 Msps 11 429 Mbps
3 2500 Msps 17 143 Mbps
2 5700 Msps 26 057 Mbps

To use this tool with a C-PHY link, just convert: number of trios × symbol rate × 16/7 gives the equivalent capacity in Mbps, to be compared against the data rate. Multi-virtual-channel cases stay out of scope, where several streams share the same lanes and the rates have to be summed before the comparison.

References

Last updated: · Spotted an error or stale figure? Let us know

← Back to the Wiki index

A similar project?

Acoustics, embedded, calculation tools: if you have a related use case, let’s talk.