Fault Tolerance · 容错
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| fault tolerant/fɒlt ˈtɒlərənt/ | 容错的 | róng cuò de |
| redundancy/rɪˈdʌndənsi/ | 冗余 | rǒng yú |
| single point of failure/ˈsɪŋɡl pɔɪnt ɒv ˈfeɪlɪə/ | 单点故障 | dān diǎn gù zhàng |
Keeping working when parts fail
- A system is fault tolerant 容错的 when it keeps working even if part of it fails.
- The Internet is the famous example — no single failure can bring the whole thing down.
- This is a design goal, not luck.
- The question is how you build it in.
部分失败时保持工作
- 一个系统是容错的(fault tolerant),当它即使部分失败也保持工作。
- 互联网是著名的例子——没有单一失败能让整个系统崩溃。
- 这是一个设计目标,不是运气。
- 问题是如何把它建进去。
A fault-tolerant system: · 一个容错系统:
No single failure brings the whole system down. · 没有单一失败让整个系统崩溃。
Having more than one path or copy so a backup takes over is called . · 有不止一条路径或副本,让一个备份接管,叫。
Redundancy is the main way to achieve fault tolerance. · 冗余是实现容错的主要方式。
Redundancy is the key
- The main way to achieve fault tolerance is redundancy 冗余.
- That means more than one path, or more than one copy, so a backup takes over.
- Because the network has many routes, packets take an alternate route around a broken link.
- This is exactly why the Internet is fault tolerant.
冗余是关键
- 实现容错的主要方式是冗余(redundancy)。
- 那意味着不止一条路径,或不止一份副本,于是一个备份接管。
- 因为网络有许多路线,数据包绕过断掉的链路走备用路线。
- 这正是互联网容错的原因。
Fault tolerant, or a single point of failure? · 容错,还是单点故障?
Redundancy (extra paths or copies) makes a system fault tolerant; relying on one part with no backup creates a single point of failure. · 冗余(额外的路径或副本)让系统容错;依赖一个没有备份的部分造成单点故障。
A single point of failure is: · 一个单点故障是:
A network with one central router has a single point of failure. · 一个只有一个中央路由器的网络有一个单点故障。
Because the Internet has many routes, packets can go around a broken link. · 因为互联网有许多路线,数据包能绕过断掉的链路。
Alternate routes are what make it fault tolerant. · 备用路线是让它容错的东西。
Single point of failure
- The danger to avoid is a single point of failure 单点故障 — one part whose failure stops everything.
- A network with only one central router has a single point of failure.
- If that router dies, the whole system stops.
- Redundancy removes single points of failure by adding backups.
单点故障
- 要避免的危险是单点故障(single point of failure)——一个其失败会停下一切的部分。
- 一个只有一个中央路由器的网络有一个单点故障。
- 如果那个路由器死了,整个系统停止。
- 冗余通过添加备份移除单点故障。
Why might a small blog accept a less reliable setup than a hospital? · 为什么一个小博客可能接受比医院更不可靠的配置?
Designers weigh the cost of redundancy against the reliability needed. · 设计者把冗余的成本与所需的可靠性相权衡。
Running a website on two servers in different cities removes the single point of failure of one server. · 在两个不同城市的两台服务器上运行网站,移除了一台服务器的单点故障。
If either server fails, the other keeps the site up. · 如果任一服务器失败,另一台保持网站在线。
Redundancy has a cost
- Redundancy is not free — extra cables, servers, and backups all cost money.
- Designers weigh the cost of redundancy against the reliability it buys.
One server vs two. A company runs its site on a single server — a single point of failure; if it crashes, the site goes offline. Adding a second server in another city means the site stays up if either fails. It costs twice as much, but for a business losing money every minute offline, the reliability is worth it.
冗余有成本
- 冗余不是免费的——额外的电缆、服务器和备份都花钱。
- 设计者把冗余的成本与它买来的可靠性相权衡。
一台服务器对两台。 一家公司在一台服务器上运行它的网站——一个单点故障;如果它崩溃,网站下线。在另一个城市加第二台服务器意味着任一失败网站仍在线。它花费两倍,但对一家每分钟下线都亏钱的企业,这个可靠性是值得的。
A fault tolerant system keeps working when a part fails, achieved mainly through redundancy — extra paths or copies so a backup takes over. Avoid a single point of failure, one part that stops everything. Redundancy costs money, so designers weigh cost against the reliability it buys.
一个容错的系统在一个部分失败时保持工作,主要通过冗余实现——额外的路径或副本,让一个备份接管。避免单点故障,一个停下一切的部分。冗余花钱,所以设计者把成本与它买来的可靠性相权衡。