Floating-point numbers · Números de punto flotante
| English | Español |
|---|---|
| floating-point/ˈfləʊtɪŋ pɔɪnt/ | punto flotante |
| mantissa/mænˈtɪsə/ | mantisa |
| exponent/ekˈspəʊnənt/ | exponente |
| normalised/ˈnɔːməlaɪzd/ | normalizado |
| precision/prɪˈsɪʒn/ | precisión |
| rounding errors/ˈraʊndɪŋ ˈerəz/ | errores de redondeo |
| fixed-point/fɪkst pɔɪnt/ | punto fijo |
| overflow/ˌəʊvəˈfləʊ/ | desbordamiento |
| underflow/ˌʌndəˈfləʊ/ | subdesbordamiento |
A missile that missed by 600 metres
- On 25 February 1991 a Patriot battery at Dhahran failed to intercept an incoming missile. Twenty-eight soldiers died.
- The system counted time in tenths of a second, and 0.1 has no exact binary representation: it is a repeating fraction, truncated to the register's width. Each tick lost a fraction of a microsecond.
- After a hundred hours running, the accumulated error was a third of a second, and in a third of a second the target moved 600 metres. The arithmetic was correct at every single step.
- This lesson is the floating-point 浮点 format, normalisation 规格化, and the approximation errors that make cases like this possible.
Un misil que falló por 600 metros
- El 25 de febrero de 1991, una batería Patriot en Dhahran no interceptó un misil entrante. Veintiocho soldados murieron.
- El sistema contaba el tiempo en décimas de segundo, y 0.1 no tiene representación binaria exacta: es una fracción repetitiva, truncada al ancho del registro. Cada tick perdió una fracción de un microsegundo.
- Después de cien horas funcionando, el error acumulado fue de un tercio de segundo, y en un tercio de segundo el objetivo se movió 600 metros. La aritmética era correcta en cada paso individual.
- Esta lección trata sobre el formato de punto flotante 浮点, la normalización 规格化, y los errores de aproximación que hacen posibles casos como este.
The format
- To store real numbers of very different sizes, a computer uses a binary form of scientific notation with two fields: a mantissa 尾数, the significant digits, and an exponent 指数, the power of 2 to multiply by. Both are stored as two's complement.
- Read the mantissa as a binary fraction: the first bit after the point is worth $\tfrac12$, the next $\tfrac14$, then $\tfrac18$. So
0.1010000is $\tfrac12 + \tfrac18 = 0.625$. - With exponent
00000010, that is $+2$, the value is $0.625 \times 2^2 = 2.5$.
A fraction and a power of two
El formato
- Para almacenar números reales de tamaños muy diferentes, un ordenador utiliza una forma binaria de notación científica con dos campos: una mantisa 尾数, los dígitos significativos, y un exponente 指数, la potencia de 2 por la que multiplicar. Ambos se almacenan como complemento a dos.
- Lee la mantisa como una fracción binaria: el primer bit después del punto vale $\tfrac12$, el siguiente $\tfrac14$, luego $\tfrac18$. Por lo tanto,
0.1010000es $\tfrac12 + \tfrac18 = 0.625$. - Con exponente
00000010, eso es $+2$, el valor es $0.625 \times 2^2 = 2.5$.

Una fracción y una potencia de dos
Build a floating-point number · Construir un número de punto flotante
Flip the mantissa and exponent bits to make a value, and check whether it is normalised. · Invertir los bits de la mantisa y del exponente para obtener un valor, y verificar si está normalizado.
A floating-point number is stored as: · Un número de punto flotante se almacena como:
value = mantissa × 2^exponent — binary scientific notation, with both parts stored in two's complement. · valor = mantisa × 2^exponente — notación científica binaria, con ambas partes almacenadas en complemento a dos.
Match each part of floating-point representation to its role. · Asocia cada parte de la representación de punto flotante con su función.
value = mantissa × 2^exponent; normalising maximises precision; money uses fixed-point/BCD to avoid rounding. · valor = mantisa × 2^exponente; la normalización maximiza la precisión; el dinero usa fijo/BCD para evitar redondeos.
Worked example: binary to denary
- A number has mantissa
10110000and exponent00000011. Find its denary value. - The exponent is $+3$. The mantissa begins with 1, so it is negative and is read by two's complement rules: as
1.0110000the sign bit is worth $-1$ and the fraction bits add $\tfrac14 + \tfrac18 = 0.375$, so the mantissa is $-1 + 0.375 = -0.625$. - $\text{number} = -0.625 \times 2^3 = -5.0$.
- The commonest error is reading a negative mantissa as if it were positive. Check the first bit before anything else.
Ejercicio resuelto: binario a decimal
- Un número tiene mantisa
10110000y exponente00000011. Encuentra su valor decimal. - El exponente es $+3$. La mantisa comienza con 1, por lo que es negativa y se lee siguiendo las reglas del complemento a dos: como
1.0110000el bit de signo vale $-1$ y los bits de fracción suman $\tfrac14 + \tfrac18 = 0.375$, por lo que la mantisa es $-1 + 0.375 = -0.625$. - $\text{number} = -0.625 \times 2^3 = -5.0$.
- El error más común es leer una mantisa negativa como si fuera positiva. Comprueba el primer bit antes de cualquier otra cosa.
A floating-point number has mantissa 10110000 and exponent 00000011. What is its denary value? · Un número de punto flotante tiene mantisa 10110000 y exponente 00000011. ¿Cuál es su valor decimal?
The mantissa starts with 1, so it is negative: −1 + 1/4 + 1/8 = −0.625. Times 2^3 gives −5.0. Reading it as positive is the usual error. · La mantisa comienza con 1, así que es negativa: −1 + 1/4 + 1/8 = −0.625. Multiplicado por 2^3 da −5.0. Leerlo como positivo es el error habitual.
Worked example: denary to binary
- Store $+2.5$ in the same 8-bit mantissa, 8-bit exponent format.
- In binary, $2.5 = 10.1$. Written as a fraction times a power of two: $2.5 = 0.101 \times 2^2$.
- So the mantissa is
01010000, a sign bit of 0 then.101padded with zeros, and the exponent is00000010. - Always write it in the normalised form first; then the two fields read straight off.
Ejercicio resuelto: decimal a binario
- Almacena $+2.5$ en el mismo formato de mantisa de 8 bits y exponente de 8 bits.
- En binario, $2.5 = 10.1$. Escrito como una fracción multiplicada por una potencia de dos: $2.5 = 0.101 \times 2^2$.
- Por lo tanto, la mantisa es
01010000, un bit de signo de 0 seguido de.101rellenado con ceros, y el exponente es00000010. - Siempre escríbelo primero en la forma normalizada; luego los dos campos se leen directamente.
Written as a normalised fraction times a power of two, 2.5 = 0.101 x 2^. Give the exponent. · Escrito como una fracción normalizada multiplicada por una potencia de dos, 2.5 = 0.101 x 2^. Indica el exponente.
2.5 is 10.1 in binary; sliding the point two places left gives 0.101, so the exponent is 2 and the mantissa is 01010000. · 2.5 es 10.1 en binario; desplazar la coma dos lugares a la izquierda da 0.101, así que el exponente es 2 y la mantisa es 01010000.
Normalisation
- A number is normalised when the first significant bit sits immediately after the binary point, so there are no wasted leading zeros. For a positive number the mantissa starts
0.1; for a negative one,1.0. - Why: it maximises precision 精度, because every mantissa bit then carries information rather than a leading zero.
- To normalise, shift the mantissa left and decrease the exponent by the same number of places, or shift right and increase it. The value is unchanged; only its representation is.
Shift the bits, adjust the exponent, keep the value
Normalización
- Un número está normalizado cuando el primer bit significativo se sitúa inmediatamente después del punto binario, de modo que no hay ceros iniciales desperdiciados. Para un número positivo, la mantisa comienza con
0.1; para uno negativo, con1.0. - Por qué: maximiza la precisión 精度, porque cada bit de la mantisa transporta información en lugar de un cero inicial.
- Para normalizar, desplaza la mantisa hacia la izquierda y disminuye el exponente en el mismo número de posiciones, o desplázala hacia la derecha y aumentalo. El valor no cambia; solo su representación.

Desplazar los bits, ajustar el exponente, mantener el valor
Trading mantissa bits against exponent bits
- The total number of bits is fixed, so the two fields compete.
- More mantissa bits means greater precision: each value is stored more exactly, with a smaller rounding error.
- More exponent bits means greater range: much larger and much smaller magnitudes can be represented, but each one less precisely.
- A question asking for the effect of moving a bit from one field to the other wants exactly this trade: range against precision.
Intercambio de bits de mantisa contra bits de exponente
- El número total de bits es fijo, por lo que los dos campos compiten.
- Más bits de mantisa significa mayor precisión: cada valor se almacena con más exactitud, con un error de redondeo menor.
- Más bits de exponente significa mayor rango: se pueden representar magnitudes mucho mayores y mucho menores, pero cada una con menos precisión.
- Una pregunta sobre el efecto de mover un bit de un campo a otro busca exactamente este intercambio: rango frente a precisión.
Normalising a floating-point number · Normalización de un número de punto flotante
Step through normalisation. Shifting the mantissa to remove wasted leading zeros — and adjusting the exponent to match — keeps the value the same but spends every bit on precision. · Paso a paso por la normalización. Desplazar la mantisa para eliminar ceros iniciales innecesarios —y ajustar el exponente en consecuencia— mantiene el valor sin cambios pero utiliza cada bit para la precisión.
Normalising a floating-point number: · Normalizar un número de punto flotante:
Normalisation shifts the mantissa so the first significant bit follows the point — every bit then carries information, and the value is unchanged. · La normalización desplaza la mantisa para que el primer bit significativo siga al punto; entonces cada bit aporta información y el valor no cambia.
Which are true of normalising a floating-point number? Select all · todos that apply. · ¿Cuáles son ciertas sobre la normalización de un número de punto flotante? Selecciona todas las que correspondan.
Normalising changes only the representation. The value is unchanged; that is what adjusting the exponent guarantees. · La normalización solo cambia la representación. El valor permanece inalterado; eso es lo que garantiza el ajuste del exponente.
Approximation and its consequences
- Many denary reals cannot be stored exactly in binary. $0.1$ is the repeating fraction $0.000110011\ldots$, so it must be truncated: the stored value is close to $0.1$ but never equal to it.
- Rounding errors 舍入误差 accumulate over many operations, which is why
0.1 + 0.2is not exactly0.3and why the Patriot's clock drifted. - So never test two reals for equality: write
ABS(x - 0.3) < 1e-9instead ofx = 0.3. And beware subtracting two nearly equal values, which throws away most of the significant digits. - For values that must be exact, currency above all, use fixed-point 定点 or BCD instead.
Aproximación y sus consecuencias
- Muchas reales decimales no se pueden almacenar exactamente en binario. $0.1$ es la fracción recurrente $0.000110011\ldots$, por lo que debe truncarse: el valor almacenado es cercano a $0.1$ pero nunca igual a él.
- Los errores de redondeo 舍入误差 se acumulan a lo largo de muchas operaciones, por eso
0.1 + 0.2no es exactamente0.3y por qué el reloj del Patriot se desvió. - Así que nunca pruebes la igualdad de dos reales: escribe
ABS(x - 0.3) < 1e-9en lugar dex = 0.3. Y cuidado al restar dos valores casi iguales, ya que esto elimina la mayoría de los dígitos significativos. - Para valores que deben ser exactos, especialmente la moneda, utiliza punto fijo 定点 o BCD en su lugar.
Match each change in the bit allocation to its effect. · Asocia cada cambio en la asignación de bits con su efecto.
The word size is fixed, so precision and range trade against each other; overflow and underflow are the exponent field running out at each end. · El tamaño de palabra es fijo, así que la precisión y el rango se contraponen; el desbordamiento y el subdesbordamiento ocurren cuando el campo del exponente se agota en ambos extremos.
Overflow and underflow
- Overflow 溢出 happens when a result is too large for the exponent's range, so it cannot be represented at all.
- Underflow 下溢 happens when a result is too small, so close to zero that the exponent cannot go low enough, and it rounds to zero.
- Both are properties of the exponent field's size, which is the other half of the trade-off above.
Desbordamiento e insuficiencia
- Desbordamiento 溢出 ocurre cuando un resultado es demasiado grande para el rango del exponente, por lo que no puede representarse en absoluto.
- Insuficiencia 下溢 ocurre cuando un resultado es demasiado pequeño, tan cercano a cero que el exponente no puede bajar lo suficiente, y se redondea a cero.
- Ambas son propiedades del tamaño del campo del exponente, que es la otra mitad del intercambio anterior.
0.1 cannot be stored exactly in binary (it is a repeating fraction), so 0.1 + 0.2 does not give exactly 0.3 on a computer. · 0.1 no puede almacenarse exactamente en binario (es una fracción periódica), por lo que 0.1 + 0.2 no da exactamente 0.3 en un ordenador.
The tiny approximation errors add up — which is why money is handled with fixed-point or BCD, not floating-point. · Los pequeños errores de aproximación se acumulan —por eso el dinero se maneja con representaciones de punto fijo o BCD, no con punto flotante.
For exact money calculations you should use: · Para cálculos monetarios exactos debes usar:
Floating-point rounding errors are unacceptable for currency; fixed-point or BCD store decimal values exactly. · Los errores de redondeo de punto flotante son inaceptables para moneda; el punto fijo o BCD almacenan valores decimales con exactitud.
Marks that slip away
- The mantissa is a fraction, not an integer: the first bit after the point is a half.
- A mantissa starting with 1 is negative and follows two's complement rules. Check that bit first.
- Normalisation maximises precision; it does not change the value, and it does not make the number bigger.
- More mantissa means precision, more exponent means range. Overflow is too big, underflow is too small.
Puntos que se pierden fácilmente
- La mantisa es una fracción, no un entero: el primer bit después del punto es un medio.
- Una mantisa que comienza con 1 es negativa y sigue las reglas del complemento a dos. Verifica ese bit primero.
- La normalización maximiza la precisión; no cambia el valor, ni hace que el número sea más grande.
- Más mantisa significa precisión, más exponente significa rango. El desbordamiento es demasiado grande, la insuficiencia es demasiado pequeña.
You've got it
- floating point stores $\text{mantissa} \times 2^{\text{exponent}}$, both in two's complement; the mantissa is a binary fraction
- convert by reading the mantissa as a fraction (two's complement if it starts with 1) and multiplying by two to the exponent
- normalised means the first significant bit is immediately after the point, which maximises precision; shifting left lowers the exponent
- more mantissa bits give precision, more exponent bits give range; binary cannot store many reals exactly, so expect rounding errors, compare with a tolerance, and use fixed-point or BCD for currency
Lo has entendido
- el punto flotante almacena $\text{mantissa} \times 2^{\text{exponent}}$, ambos en complemento a dos; la mantisa es una fracción binaria fracción
- convertir leyendo la mantisa como una fracción (complemento a dos si comienza con 1) y multiplicando por dos elevado al exponente
- normalizado significa que el primer bit significativo está inmediatamente después del punto, lo que maximiza la precisión; desplazar a la izquierda reduce el exponente
- más bits de mantisa dan precisión, más bits de exponente dan rango; el binario no puede almacenar muchas reales exactamente, así que espera errores de redondeo, compara con una tolerancia y usa punto fijo o BCD para la moneda