How Computers Represent Numbers: Floating Point and Its Surprises

•5 min read•

Open a console in almost any programming language and type 0.1 + 0.2. You will not get 0.3. You will get 0.30000000000000004. New programmers assume this is a bug in the language. It is not. The same answer comes back in JavaScript, Python, Java, C, and Rust, because they all store decimal numbers the same way, and that way cannot represent 0.1 exactly. Understanding why is one of those pieces of knowledge that quietly separates people who get burned by number handling from people who do not.

Integers are the easy part

Computers store everything in binary, as sequences of bits that are either zero or one. Whole numbers map onto binary cleanly. The bits are place values that are powers of two instead of powers of ten, so 1011 in binary is 8 plus 0 plus 2 plus 1, which is 11. Nothing is lost.

Negative integers use a scheme called two's complement[1], which flips the bits of the positive value and adds one. The clever result is that ordinary binary addition works for negative numbers too, with no special case, which is why hardware adders are simple and fast. The only real limit on integers is range. A 32 bit integer can hold a bit over four billion values[2], and a 64 bit integer holds far more than you will usually need. As long as you stay inside the range, integer math is exact.

Fractions are where it gets hard

The trouble starts with fractions. We want to store numbers like 0.1, and we want them to fit in a fixed number of bits so the hardware can work with them quickly. The format nearly every computer uses is called IEEE 754, standardized in 1985[3], and it works like scientific notation in binary.

A 64 bit floating point number, the kind most languages call a double[4], splits its bits into three parts. One bit is the sign. Eleven bits are the exponent, which says where the binary point goes, giving the format its enormous range. The remaining fifty-two bits are the significand, the actual digits of the number. The value is the significand times two raised to the exponent, with the sign applied. This is a brilliant design. It represents numbers from the subatomic to the astronomical in the same sixty-four bits, with roughly fifteen to seventeen significant decimal digits of precision.

Why 0.1 cannot be stored

Here is the catch. In base ten, some fractions have no finite representation. One third is 0.3333... repeating forever, and you have to round it to write it down. The same thing happens in base two, just for different fractions. The number 0.1 in binary is 0.0001100110011..., a pattern that repeats forever.[5]

Since the significand has only fifty-two bits, the computer cannot store the infinite pattern. It stores the closest value it can and rounds off the rest. So the number sitting in memory is not exactly 0.1. It is a hair above or below. Store 0.1 and 0.2, both slightly off, add them, and the two small errors combine into something that is visibly not 0.3.

0.1 + 0.2            // 0.30000000000000004
0.1 + 0.2 === 0.3    // false
(0.1 + 0.2).toFixed(2) // "0.30"

The number was never wrong. It was rounded, exactly as the format requires, and the rounding simply became visible.

The rules that follow from this

Once you accept that floating point values are approximations, several practical rules fall out.

Never compare floating point numbers for exact equality. Instead, check whether the difference between them is smaller than a tiny tolerance, often called an epsilon. Two calculations that should produce the same value can land a fraction apart.

Be careful adding numbers of wildly different sizes. Add a tiny number to a huge one and the tiny number can vanish entirely, because there are no bits left to record it. Summing a long list of small values into a running total can accumulate error, which is why numerical libraries use smarter summation algorithms.

Watch for special values. IEEE 754 defines positive and negative infinity, which you get from overflow or dividing by zero in floating point, and a value called NaN, meaning not a number, which comes from undefined operations like zero divided by zero. NaN has a famous property: it is not equal to anything, including itself.[6] A check of x !== x being true is the old trick for spotting one.

Money is the classic trap

The place this bites hardest is money. Representing dollars and cents as floating point is asking for one cent discrepancies that accountants will find and you will not enjoy explaining. There are two standard fixes.

The simplest is to store money as integers in the smallest unit, cents instead of dollars. An item that costs 19.99 is stored as the integer 1999, all arithmetic is exact integer arithmetic, and you divide by one hundred only when you display it. The alternative is a decimal type, which many languages and databases provide[7], that stores numbers in base ten and does exact decimal arithmetic at some cost in speed. Use integers or a decimal type for anything involving currency, and keep floating point for the measurements and scientific values it was designed for.

The takeaway

Floating point is not broken, and 0.1 + 0.2 is not a defect. It is the honest result of fitting an infinite decimal into sixty-four bits and rounding the remainder, the same way you round one third when you write it as a decimal. Integers are exact within their range, floating point trades exactness for enormous range and is right to about fifteen digits, and money belongs in integers or a decimal type. Knowing which representation you are standing on tells you exactly when to trust the last few digits and when to stop.

Sources (7)
  1. Wikipedia: Two's complement
  2. Wikipedia: Integer (computer science)
  3. Wikipedia: IEEE 754
  4. Wikipedia: Double-precision floating-point format
  5. Wikipedia: Floating-point arithmetic
  6. Wikipedia: NaN
  7. Wikipedia: Decimal data type