Every new board goes through the same ritual: you power it up for the first time, and almost always something refuses to start. In the bring-up of an MCU — be it an ARM Cortex-M or a RISC-V core — two failures recur with an almost boring regularity: the crystal oscillator that won’t oscillate (or drifts), and the UART that spits out random characters. Both are bugs you can diagnose on paper, if you know where to look. Here’s how we approach them, why they keep happening, and why the cause is almost always a wrong number from the design phase, not a hardware fault.
The crystal that won’t start: it’s about load capacitance
The symptom is unmistakable: the system clock won’t lock, the PLL won’t stabilize, or the frequency is hundreds of ppm out of spec and everything time-dependent — UART, USB, timers — drifts. The instinctive first reaction is to suspect a defective crystal. It almost never is.
The usual culprit is load capacitance. A crystal isn’t a part that just “oscillates”: it oscillates at its nominal frequency only when it sees a specific capacitance across its terminals, the load capacitance CL stated on the datasheet (typically 8, 12, 18, 20 pF). That value is the capacitance the external circuit must present to the crystal. If it’s wrong, the crystal oscillates at a slightly different frequency (pulling), or the oscillation loop doesn’t have enough gain and won’t start at all.
The thing many people get wrong is computing the two capacitors C1 and C2 at the crystal pins. Two equal capacitors set to CL won’t do: the crystal sees the series of C1 and C2, plus the parasitic capacitances of the PCB and the package pins. The correct formula is:
CL = (C1 · C2) / (C1 + C2) + Cstray
where Cstray collects the parasitics of traces, pads and the on-chip oscillator inputs — typically 3-5 pF, not negligible. With C1 = C2, the series term is C/2: to obtain CL = 18 pF with Cstray = 4 pF you need C1 = C2 ≈ 28 pF, not 18 pF and not 36 pF either. Getting this wrong is mistake number one: you end up with capacitances too high (the crystal struggles to start) or too low (frequency out of spec). To avoid redoing the math by hand every time we put a load-capacitance calculator online; the full reasoning, with the edge cases, is in the dedicated wiki.
The unreadable UART: baud-rate error
The second classic shows up right after the clock comes alive: you open the terminal, expect the boot banner, and read garbage. If the characters are completely random, it’s usually a grossly wrong baud rate. But there’s a more insidious case — characters almost right, some bytes correct and some corrupted, especially on long strings — that betrays a baud-rate error of just a few percent.
An asynchronous UART generates its own transmit clock by dividing a peripheral clock by an integer (or fractional, on newer MCUs) divisor. The actual baud rate is therefore almost always slightly different from the nominal one, because the nearest integer divisor rarely hits it exactly. The percentage difference between actual and nominal baud rate is the baud error.
The problem is that transmitter and receiver sample bits at predefined instants inside the frame. If the two clocks diverge too much, the offset accumulates bit after bit and, toward the end of the frame (start + 8 data + stop), the receiver is already sampling the next bit. The rule of thumb we use: keep the combined error of the two sides within about ±2-3%. Beyond that threshold frames start to corrupt, first sporadically, then systematically. It’s a particularly sneaky bug because it’s content-dependent: characters with certain bit patterns get through, others don’t.
The diagnosis is purely numerical: take the actual peripheral clock (careful: after the PLL, not the crystal’s nominal — and this is why the two bugs are cousins), the configured divisor, and compute the effective baud and the percentage error. If it’s over threshold, pick a friendlier peripheral clock (many projects adopt the classic 14.7456 or 18.432 MHz precisely because they divide into standard bauds with no error) or a different baud rate. Here too we have a UART baud-error calculator and its wiki with the offset table.
The arithmetic before the screwdriver
These two bugs share a moral: in bring-up, the problem is almost always a number decided at design time — not a broken part. The crystal that won’t start and the unreadable serial are both children of a wrong clock calculation, made weeks earlier at schematic time. You diagnose them with pen and paper, before even reaching for the oscilloscope. Checking them ahead of time, with the right numbers, saves hours of probing and testing on a board that “won’t work for no clear reason”.
Are you bringing up a board and something doesn’t add up on the clock or the serial? Let’s talk.