Floating-point numbers · 浮点数
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| floating-point/ˈfləʊtɪŋ pɔɪnt/ | 浮点 | fú diǎn |
| mantissa/mænˈtɪsə/ | 尾数 | wěi shù |
| exponent/ekˈspəʊnənt/ | 指数 | zhǐ shù |
| normalised/ˈnɔːməlaɪzd/ | 规格化 | guī gé huà |
| precision/prɪˈsɪʒn/ | 精度 | jīng dù |
| rounding errors/ˈraʊndɪŋ ˈerəz/ | 舍入误差 | shě rù wù chā |
| fixed-point/fɪkst pɔɪnt/ | 定点 | dìng diǎn |
| overflow/ˌəʊvəˈfləʊ/ | 溢出 | yì chū |
| underflow/ˌʌndəˈfləʊ/ | 下溢 | xià yì |
A missile that missed by 600 metres
- On 25 February 1991 a Patriot battery at Dhahran failed to intercept an incoming missile. Twenty-eight soldiers died.
- The system counted time in tenths of a second, and 0.1 has no exact binary representation: it is a repeating fraction, truncated to the register's width. Each tick lost a fraction of a microsecond.
- After a hundred hours running, the accumulated error was a third of a second, and in a third of a second the target moved 600 metres. The arithmetic was correct at every single step.
- This lesson is the floating-point 浮点 format, normalisation 规格化, and the approximation errors that make cases like this possible.
偏了 600 米的拦截
- 1991 年 2 月 25 日,达兰的一个爱国者导弹连没能拦截来袭导弹。二十八名士兵死亡。
- 系统以十分之一秒计时,而 0.1 没有精确的二进制表示:它是一个循环小数,被截断到寄存器的宽度。每一次计时都丢掉零点几微秒。
- 连续运行一百小时后,累积误差达到三分之一秒,而在三分之一秒里目标移动了 600 米。每一步的算术都是对的。
- 这一课讲浮点(floating-point)格式、规格化(normalisation),以及让这类事故成为可能的近似误差。
The format
- To store real numbers of very different sizes, a computer uses a binary form of scientific notation with two fields: a mantissa 尾数, the significant digits, and an exponent 指数, the power of 2 to multiply by. Both are stored as two's complement.
- Read the mantissa as a binary fraction: the first bit after the point is worth $\tfrac12$, the next $\tfrac14$, then $\tfrac18$. So
0.1010000is $\tfrac12 + \tfrac18 = 0.625$. - With exponent
00000010, that is $+2$, the value is $0.625 \times 2^2 = 2.5$.
A fraction and a power of two
格式
- 为了存储大小相差悬殊的实数,计算机使用二进制形式的科学记数法,有两个字段:尾数(mantissa),即有效数字;指数(exponent),即要乘的 2 的幂。两者都以补码存储。
- 把尾数读作二进制小数:小数点后第一位值 $\tfrac12$,下一位 $\tfrac14$,再下一位 $\tfrac18$。所以
0.1010000是 $\tfrac12 + \tfrac18 = 0.625$。 - 配上指数
00000010,即 $+2$,值就是 $0.625 \times 2^2 = 2.5$。

一个小数和一个 2 的幂
Build a floating-point number · 构建一个浮点数
Flip the mantissa and exponent bits to make a value, and check whether it is normalised. · 翻转尾数和指数的位来构建一个值,并检查它是否已规范化。
A floating-point number is stored as: · 一个浮点数被存储为:
value = mantissa × 2^exponent — binary scientific notation, with both parts stored in two's complement. · 值 = 尾数 × 2^指数——二进制科学记数法,两部分都存为补码。
Match each part of floating-point representation to its role. · 把浮点表示的每个部分与它的作用配对。
value = mantissa × 2^exponent; normalising maximises precision; money uses fixed-point/BCD to avoid rounding. · 值 = 尾数 × 2^指数;规范化最大化精度;货币用定点/BCD 避免舍入。
Worked example: binary to denary
- A number has mantissa
10110000and exponent00000011. Find its denary value. - The exponent is $+3$. The mantissa begins with 1, so it is negative and is read by two's complement rules: as
1.0110000the sign bit is worth $-1$ and the fraction bits add $\tfrac14 + \tfrac18 = 0.375$, so the mantissa is $-1 + 0.375 = -0.625$. - $\text{number} = -0.625 \times 2^3 = -5.0$.
- The commonest error is reading a negative mantissa as if it were positive. Check the first bit before anything else.
例题:二进制转十进制
- 一个数的尾数是
10110000,指数是00000011。求它的十进制值。 - 指数是 $+3$。尾数以 1 开头,所以它是负的,按补码规则读:作为
1.0110000,符号位值 $-1$,小数位加起来 $\tfrac14 + \tfrac18 = 0.375$,所以尾数是 $-1 + 0.375 = -0.625$。 - $\text{number} = -0.625 \times 2^3 = -5.0$。
- 最常见的错误是把负尾数当成正的读。先看第一位,再做别的。
A floating-point number has mantissa 10110000 and exponent 00000011. What is its denary value? · 一个浮点数的尾数是 10110000,指数是 00000011。它的十进制值是多少?
The mantissa starts with 1, so it is negative: −1 + 1/4 + 1/8 = −0.625. Times 2^3 gives −5.0. Reading it as positive is the usual error. · 尾数以 1 开头,所以是负的:−1 + 1/4 + 1/8 = −0.625。乘以 2^3 得 −5.0。把它当成正的读是常见错误。
Worked example: denary to binary
- Store $+2.5$ in the same 8-bit mantissa, 8-bit exponent format.
- In binary, $2.5 = 10.1$. Written as a fraction times a power of two: $2.5 = 0.101 \times 2^2$.
- So the mantissa is
01010000, a sign bit of 0 then.101padded with zeros, and the exponent is00000010. - Always write it in the normalised form first; then the two fields read straight off.
例题:十进制转二进制
- 用同样的 8 位尾数、8 位指数格式存储 $+2.5$。
- 二进制里 $2.5 = 10.1$。写成小数乘 2 的幂:$2.5 = 0.101 \times 2^2$。
- 所以尾数是
01010000,即符号位 0 再加上补零的.101;指数是00000010。 - 先写成规格化形式,两个字段就能直接读出来。
Written as a normalised fraction times a power of two, 2.5 = 0.101 x 2^. Give the exponent. · 写成规格化小数乘 2 的幂,2.5 = 0.101 x 2^。给出指数。
2.5 is 10.1 in binary; sliding the point two places left gives 0.101, so the exponent is 2 and the mantissa is 01010000. · 2.5 的二进制是 10.1;把小数点左移两位得 0.101,所以指数是 2,尾数是 01010000。
Normalisation
- A number is normalised when the first significant bit sits immediately after the binary point, so there are no wasted leading zeros. For a positive number the mantissa starts
0.1; for a negative one,1.0. - Why: it maximises precision 精度, because every mantissa bit then carries information rather than a leading zero.
- To normalise, shift the mantissa left and decrease the exponent by the same number of places, or shift right and increase it. The value is unchanged; only its representation is.
Shift the bits, adjust the exponent, keep the value
规格化
- 当第一个有效位紧挨在二进制小数点之后、没有浪费的前导零时,数就是规格化的。正数的尾数以
0.1开头;负数以1.0开头。 - 为什么:它最大化精度(precision),因为这样每个尾数位都携带信息,而不是一个前导零。
- 规格化的做法是把尾数左移,并把指数减少相同的位数;或右移并增加指数。值不变,变的只是它的表示。

移动这些位,调整指数,值保持不变
Trading mantissa bits against exponent bits
- The total number of bits is fixed, so the two fields compete.
- More mantissa bits means greater precision: each value is stored more exactly, with a smaller rounding error.
- More exponent bits means greater range: much larger and much smaller magnitudes can be represented, but each one less precisely.
- A question asking for the effect of moving a bit from one field to the other wants exactly this trade: range against precision.
尾数位与指数位的取舍
- 总位数是固定的,所以两个字段互相争夺。
- 尾数位更多意味着更高的精度:每个值存得更精确,舍入误差更小。
- 指数位更多意味着更大的范围:能表示大得多和小得多的量级,但每一个都不那么精确。
- 问"把一位从一个字段挪到另一个有什么影响"的题,要的正是这个取舍:范围换精度。
Normalising a floating-point number · 规范化一个浮点数
Step through normalisation. Shifting the mantissa to remove wasted leading zeros — and adjusting the exponent to match — keeps the value the same but spends every bit on precision. · 逐步走过规范化。移动尾数以去掉浪费的前导零——并调整指数来匹配——保持值不变,但把每个位都用在精度上。
Normalising a floating-point number: · 规范化一个浮点数:
Normalisation shifts the mantissa so the first significant bit follows the point — every bit then carries information, and the value is unchanged. · 规范化移动尾数,使第一个有效位跟在小数点之后——那时每个位都携带信息,而值不变。
Which are true of normalising a floating-point number? Select all · 所有 that apply. · 关于浮点数的规格化,哪些是对的?选出所有适用的。
Normalising changes only the representation. The value is unchanged; that is what adjusting the exponent guarantees. · 规格化只改变表示。值不变;这正是调整指数所保证的。
Approximation and its consequences
- Many denary reals cannot be stored exactly in binary. $0.1$ is the repeating fraction $0.000110011\ldots$, so it must be truncated: the stored value is close to $0.1$ but never equal to it.
- Rounding errors 舍入误差 accumulate over many operations, which is why
0.1 + 0.2is not exactly0.3and why the Patriot's clock drifted. - So never test two reals for equality: write
ABS(x - 0.3) < 1e-9instead ofx = 0.3. And beware subtracting two nearly equal values, which throws away most of the significant digits. - For values that must be exact, currency above all, use fixed-point 定点 or BCD instead.
近似及其后果
- 许多十进制实数在二进制里无法精确存储。$0.1$ 是循环小数 $0.000110011\ldots$,所以必须截断:存下的值接近 $0.1$,却永远不等于它。
- 舍入误差(rounding errors)在多次运算中累积,这就是
0.1 + 0.2不精确等于0.3、以及爱国者系统时钟漂移的原因。 - 所以绝不要检验两个实数是否相等:写
ABS(x - 0.3) < 1e-9而不是x = 0.3。也要当心相减两个几乎相等的值,那会丢掉大部分有效数字。 - 对必须精确的值——首先是货币——改用定点(fixed-point)或 BCD。
Match each change in the bit allocation to its effect. · 把位分配的每种改变与它的影响配对。
The word size is fixed, so precision and range trade against each other; overflow and underflow are the exponent field running out at each end. · 字长固定,所以精度和范围互相取舍;上溢和下溢是指数字段在两端用尽。
Overflow and underflow
- Overflow 溢出 happens when a result is too large for the exponent's range, so it cannot be represented at all.
- Underflow 下溢 happens when a result is too small, so close to zero that the exponent cannot go low enough, and it rounds to zero.
- Both are properties of the exponent field's size, which is the other half of the trade-off above.
上溢与下溢
- 当结果对指数的范围来说太大、根本无法表示时,发生上溢(overflow)。
- 当结果太小、太接近零以致指数降不到那么低时,发生下溢(underflow),它会舍入为零。
- 两者都是指数字段大小的性质,而那正是上面那个取舍的另一半。
0.1 cannot be stored exactly in binary (it is a repeating fraction), so 0.1 + 0.2 does not give exactly 0.3 on a computer. · 0.1 不能在二进制中被精确存储(它是一个循环小数),所以 0.1 + 0.2 在一台计算机上不给出恰好 0.3。
The tiny approximation errors add up — which is why money is handled with fixed-point or BCD, not floating-point. · 微小的近似错误累积起来——这就是为什么货币用定点或 BCD 处理,而不是浮点。
For exact money calculations you should use: · 对精确的货币计算,你应该用:
Floating-point rounding errors are unacceptable for currency; fixed-point or BCD store decimal values exactly. · 浮点舍入错误对货币是不可接受的;定点或 BCD 精确地存储十进制值。
Marks that slip away
- The mantissa is a fraction, not an integer: the first bit after the point is a half.
- A mantissa starting with 1 is negative and follows two's complement rules. Check that bit first.
- Normalisation maximises precision; it does not change the value, and it does not make the number bigger.
- More mantissa means precision, more exponent means range. Overflow is too big, underflow is too small.
容易丢掉的分
- 尾数是小数,不是整数:小数点后第一位是二分之一。
- 以 1 开头的尾数是负的,遵循补码规则。先看那一位。
- 规格化最大化精度;它不改变值,也不会让数变大。
- 尾数更多意味着精度,指数更多意味着范围。上溢是太大,下溢是太小。
You've got it
- floating point stores $\text{mantissa} \times 2^{\text{exponent}}$, both in two's complement; the mantissa is a binary fraction
- convert by reading the mantissa as a fraction (two's complement if it starts with 1) and multiplying by two to the exponent
- normalised means the first significant bit is immediately after the point, which maximises precision; shifting left lowers the exponent
- more mantissa bits give precision, more exponent bits give range; binary cannot store many reals exactly, so expect rounding errors, compare with a tolerance, and use fixed-point or BCD for currency
你掌握了
- 浮点存储 $\text{mantissa} \times 2^{\text{exponent}}$,两者都用补码;尾数是二进制小数
- 转换时把尾数读作小数(以 1 开头就按补码),再乘以 2 的指数次幂
- 规格化意味着第一个有效位紧挨小数点之后,这最大化精度;左移会降低指数
- 尾数位更多给精度,指数位更多给范围;二进制无法精确存储许多实数,所以要预期舍入误差、用容差比较,货币用定点或 BCD