Data formats and endianness
How the same bytes become different numbers — big/little endianness, two's-complement integers, IEEE-754 single/double floats (sign, biased exponent, mantissa) and the practical cases of reading registers and binary frames.
A buffer of bytes has, by itself, no numeric meaning: it acquires one only when we decide how to read it — how many bytes to group, in what order, signed or unsigned, as an integer or as a float. The same 40 49 0F DB is an unsigned integer, a signed integer or π in floating point, depending on the convention. This tool shows every interpretation in parallel; this page explains the rules behind them.
Endianness: big and little
A value larger than one byte spans several memory cells, and their order has to be decided. In big-endian the most significant byte (MSB) comes first — the same order in which we write the digits of a number. In little-endian it comes last: the least significant byte opens the sequence.
The difference is concrete. The sequence 01 00 read as a 16-bit integer is:
- big-endian: 0x0100 = 256
- little-endian: 0x0001 = 1
Same two bytes, two different numbers. Neither is “right” in absolute terms: it is a convention, and it has to be agreed between writer and reader.
In practice the x86 and ARM architectures (in their typical configuration) are little-endian. Many network protocols transmit in big-endian — so much so that big-endian is also called network byte order; numerous binary file formats and peripheral datasheets specify the order explicitly. Getting endianness wrong is one of the most common — and most silent — mistakes in binary data parsing.
The symptom you expect is not the one you get
You often hear that a float read with the wrong endianness “turns into NaN”. That would be lucky, because a NaN is caught by the first check. Tried on 200,000 values with magnitudes from 10⁻² to 10⁴ — the typical range of a measurement — reversing the four bytes of a float32 gives:
| What you get | Frequency |
|---|---|
| a plausible value | 51.69 % |
| huge magnitude (> 10²⁰) | 24.07 % |
| tiny magnitude (< 10⁻²⁰) | 23.83 % |
| NaN | 0.41 % |
| infinity or zero | 0.00 % |
NaN turns up four times in a thousand. In more than half the cases you get a number that looks entirely innocuous, passes any validity check and propagates downstream. That is the real reason an endianness error is “silent”: not because it leaves no trace, but because the trace it leaves looks like data.
On individual values you can see the result has no regularity at all:
| Expected value | Bytes reversed | Words reversed (middle-endian) |
|---|---|---|
| 1 | 4.6006e−41 | 2.2780e−41 |
| 3.14159 | −9.6158e+9 | 2.0535e−29 |
| 0.1 | −4.2949e+8 | −1.0761e+8 |
| 25.5 | 7.3272e−41 | 2.3603e−41 |
| 1000 | 4.3861e−41 | 2.4565e−41 |
| 0.001 | 4.5343e+28 | 7.5487e−28 |
| 1000000 | 3.3478e−39 | 2.7818e−17 |
And there is middle-endian too
The last column of the table is a third case, which anyone working with industrial devices meets sooner or later. Many Modbus slaves expose a 32-bit value as two 16-bit registers: each register is big-endian by specification, but the order of the two registers is not always. The result is a byte order of 2 3 0 1 — neither big nor little, often called middle-endian or byte-swapped — and neither classic interpretation decodes it.
This is the case where showing the interpretations side by side is not enough: if neither big nor little gives a sensible value but the two individual registers look plausible, the answer is almost always the word order. The tool covers big and little, not middle-endian: the practical fix is to swap the two byte pairs by hand before pasting them.
Integers: signed and unsigned, two’s complement
A group of N bits can be read as an unsigned integer — a non-negative value in [0, 2ᴺ−1] — or as a signed one. The universal signed encoding is two’s complement: the most significant bit is not added, it is subtracted. For N bits:
The negative of a number is obtained by inverting all bits and adding 1. The advantage of two’s complement is that addition and subtraction work with the same hardware as unsigned integers: there is no separate “−0” to handle.
A practical consequence: the same 8 bits 0xFF are −1 read as int8 and 255 read as uint8. 0x80 is −128 signed and 128 unsigned. This is why a 16-bit register that “should be positive” but arrives as a huge number is almost always an int read as a uint (or vice versa).
IEEE 754: sign, exponent, mantissa
Floating-point numbers follow the IEEE 754 standard. A value is encoded in three fields: a sign bit s, an exponent e stored with a bias, and a mantissa m (the significant digits). For normal values:
The two common formats:
| sign bits | exponent bits | mantissa bits | bias | |
|---|---|---|---|---|
| single (float32) | 1 | 8 | 23 | 127 |
| double (float64) | 1 | 11 | 52 | 1023 |
Two details make IEEE 754 efficient:
- exponent bias: the exponent is stored as a non-negative number, from which the bias is subtracted to get the real one. So a raw exponent of 128 in the single is 128 − 127 = +1, and one of 126 is −1. This allows positive and negative exponents to be represented without a second sign bit.
- hidden bit: for normal numbers the mantissa has an implicit leading 1 before the point, which is not stored. So the 23 mantissa bits of the single actually give 24 bits of precision.
The exponent edge values are reserved:
- all-zero exponent: zero (zero mantissa, with ±0 distinguished by the sign bit) and subnormals (non-zero mantissa, without the implicit 1 — they fill the gap around zero).
- all-one exponent: infinity (zero mantissa) and NaN (non-zero mantissa — undefined results such as 0/0).
The tool decomposes floats by showing the three fields as coloured bit groups, reports the raw exponent, the bias and the actual one, reconstructs the value from the fields and states its class (normal, subnormal, zero, infinity, NaN).
An example: 40 49 0F DB
The four bytes 40 49 0F DB, read as a big-endian float32, are 0x40490FDB. The sign bit is 0 (positive); the 8 exponent bits are 0x80 = 128, giving the real exponent 128 − 127 = +1; the 23 mantissa bits, with the implicit 1, give 1.5708. The value is therefore (+1)·1.5708·2¹ = 3.14159…, i.e. π. Read in little-endian, the same bytes become 0xDB0F4940: a completely different float, and a negative one at that (the sign bit now falls in the 0xDB byte). Same sequence, unrecognisable result — the practical proof of how much endianness changes everything.
How much precision there really is
The two formats have a fixed number of mantissa bits, so the step between two consecutive representable numbers grows with the value. That is not a detail: it is the source of most numerical surprises.
| Magnitude | float32 step | float64 step |
|---|---|---|
| 1 | 1.19e−7 | 2.22e−16 |
| 100 | 7.63e−6 | 1.42e−14 |
| 10⁴ | 9.77e−4 | 1.82e−12 |
| 10⁶ | 6.25e−2 | 1.16e−10 |
| 1.68·10⁷ | 2 | 3.73e−9 |
| 10⁹ | 64 | 1.19e−7 |
| 10¹⁵ | 3.36e+7 | 0.125 |
float32 has 24 effective bits of precision, so from 1.68·10⁷ upwards the step exceeds unity: two consecutive integers are no longer distinguishable. The exact boundary is 2²⁴ = 16,777,216, past which 16777216 and 16777217 are the same float32. For the double the boundary is 2⁵³ = 9,007,199,254,740,992.
Two frequent operational consequences follow:
- a millisecond counter in float32 loses millisecond resolution after about 4.7 hours (16.7 million ms) — which is why timestamps are not kept in float32;
- a 64-bit identifier does not survive a trip through a
double, which is JavaScript’s numeric type (Numberis IEEE-754 binary64). JSON, by contrast, defines a decimal number syntax without mandating an internal representation — but in practice it has to be treated as binary64, because that is what the most common implementations use: above 2⁵³ IDs are silently rounded. This is why large IDs are transmitted as strings.
| Format | Min subnormal | Min normal | Max | ε (eps) | Reliable / round-trip digits |
|---|---|---|---|---|---|
| float32 | 1.40e−45 | 1.18e−38 | 3.40e+38 | 1.19e−7 | 6 / 9 |
| float64 | 4.94e−324 | 2.23e−308 | 1.80e+308 | 2.22e−16 | 15 / 17 |
The “reliable digits” are the ones you can read and rewrite without losing anything (6 for single, 15 for double); the “round-trip digits” are the ones needed to guarantee the re-read value is bit for bit the same (9 and 17). They are two different numbers for two different purposes: the first for printing, the second for serializing.
The error accumulates
The worst case is not the single rounding but the sum of many. Adding 0.01 a million times, where the exact result is 10,000:
- in float64: 10,000.000000172 — an error of 1.7·10⁻⁷, negligible;
- in float32: 9,865.22 — an error of 134.78, that is −1.35 %.
float32 is off by more than one percent not because of a bug but because, once the accumulator has passed a thousand, the representation step has grown larger than the increment being added: some of the addends are lost. This is why integrators, IIR filters and totalizers are written in double precision or fixed point, even when the input data is float32.
And the classic, which here gets an explanation rather than just an observation: 0.1 + 0.2 gives 0.30000000000000004 in double because neither 0.1 nor 0.2 has a finite binary representation — just as 1/3 has none in decimal. In float32 the same sum gives 0.30000001192092896.
Practical cases
Knowing how to read bytes is an everyday task in instrumentation and protocols:
- Reading registers: a datasheet defines a 16- or 32-bit register, signed or unsigned, and a given endian. To interpret it you have to apply exactly that convention.
- Parsing binary frames: Modbus, CAN, a UDP payload, a proprietary file format — each field has a width, a type and a byte order. The first step in debugging a parser is figuring out which type+endian combination produces the expected value.
- Round-trip and data exchange: a float saved by a microcontroller and read back on the PC arrives “wrong” if the two ends disagree on endianness. Seeing the interpretations side by side makes it immediate to find the correct combination.
The tool’s inverse mode does the opposite path: given a value, a type and an endian, it returns the hex bytes — handy for building a test frame or verifying a round-trip.
References
- IEEE 754 — standard for floating-point arithmetic (single/double formats, special cases, rounding).
- Related tool — the data format and endianness inspector puts this page into practice: paste the bytes, choose the offset and read every interpretation in big and little endian, with the IEEE-754 decomposition.
A similar project?
Acoustics, embedded, calculation tools: if you have a related use case, let’s talk.