Sound — sampling and file size · 声音——采样与文件大小
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| analogue data/ˈænəlɒɡ ˈdeɪtə/ | 模拟数据 | mó nǐ shù jù |
| digital data/ˈdɪdʒɪtl ˈdeɪtə/ | 数字数据 | shù zì shù jù |
| sampling/ˈsæmplɪŋ/ | 采样 | cǎi yàng |
| amplitude/ˈæmplɪtjuːd/ | 振幅 | zhèn fú |
| sampling rate/ˈsæmplɪŋ reɪt/ | 采样率 | cǎi yàng lǜ |
| sampling resolution/ˈsæmplɪŋ ˌrezəˈluːʃn/ | 采样分辨率 | cǎi yàng fēn biàn lǜ |
| quantisation noise/ˌkwɒntaɪˈzeɪʃn nɔɪz/ | 量化噪声 | liàng huà zào shēng |
Why the phone makes you sound thin
- A telephone call samples your voice 8,000 times a second. A CD samples the same voice 44,100 times a second. Your voice has not changed; the phone simply throws away everything above about 3.4 kHz.
- That is why a phone call is perfectly intelligible and unmistakably thin: speech lives below 3.4 kHz, but the sparkle of a cymbal or a hard
ssound does not. - The choice was not laziness. It was a calculation: sample fast enough to carry speech, and no faster, because every extra sample costs bandwidth on every call in the country.
- This lesson is how a continuous sound becomes numbers, the two settings that control quality and size, and the calculation the exam asks for.
电话为什么让你的声音变薄
- 一通电话每秒对你的声音采样 8000 次。CD 对同一个声音每秒采样 44,100 次。你的声音没变;电话只是把约 3.4 kHz 以上的一切都扔掉了。
- 这就是为什么电话通话完全听得懂,却又明显发薄:语音住在 3.4 kHz 以下,而镲片的亮泽或一个硬的
s音不在那里。 - 这个选择不是偷懒。它是一次计算:采样快到足以承载语音,而不再更快,因为每多一个采样,全国每一通电话都要多付带宽。
- 这一课讲连续的声音怎样变成数字、控制质量和大小的两个设置,以及考试要求的计算。
Analogue and digital
- Sound in the real world is a continuous pressure wave: analogue data 模拟数据, which varies smoothly and takes infinitely many values.
- A computer stores digital data 数字数据: discrete numbers. Converting one to the other is sampling 采样.
- Sampling measures the wave's amplitude 振幅, its height, at regular time intervals, and stores each measurement as a binary number.
The smooth wave becomes a row of numbers
模拟与数字
- 真实世界中的声音是连续的压力波:模拟数据(analogue data),平滑变化,取无穷多个值。
- 计算机存储数字数据(digital data):离散的数。把一种转成另一种就是采样(sampling)。
- 采样在等间隔的时刻测量波的振幅(amplitude),即它的高度,并把每次测量存成一个二进制数。

平滑的波变成一排数字
What does sampling do? · 采样做什么?
Sampling turns a continuous analogue wave into discrete digital values. Removing inaudible frequencies is lossy compression, a different step. · 采样把连续的模拟波变成离散的数字值。去掉听不见的频率是有损压缩,是另一步。
The two settings
- Sampling rate 采样率 is the number of samples taken per second, measured in hertz. CD quality is 44.1 kHz; a telephone uses 8 kHz.
- Sampling resolution 采样分辨率 (sample bit depth) is the number of bits used to store each sample's amplitude. CD quality is 16 bits, giving 65,536 possible levels.
- The rate controls which pitches survive; the resolution controls how precisely each measurement is stored.
两个设置
- 采样率(sampling rate)是每秒采样的次数,以赫兹计。CD 品质是 44.1 kHz;电话用 8 kHz。
- 采样分辨率(sampling resolution,采样位深)是存储每个采样振幅所用的位数。CD 品质是 16 位,给出 65,536 个可能的电平。
- 采样率控制哪些音高能保留下来;分辨率控制每次测量存得多精确。
Sound sampling · 声音采样
y = a sin(bt + c)
Sampling measures a sound wave · 声波 at regular intervals — a higher rate copies it more truly. · 采样在规则的间隔测量一个声波——一个更高的率更真实地复制它。
The sampling rate of a digital sound is: · 一个数字声音的采样率是:
Sampling rate = samples per second (hertz). The bits per sample is the sample resolution. · 采样率 = 每秒的样本数(赫兹)。每个样本的位数是采样分辨率。
Put the steps of turning a sound into a digital file in order. · 把把声音变成数字文件的步骤按顺序排列。
Rate decides how often you measure; resolution decides how finely each measurement is rounded, and the rounding is where quantisation noise comes from. · 采样率决定测量的频率;分辨率决定每次测量舍入得多细,而量化噪声正是从这个舍入来的。
Sound file size
- Mono is one channel; stereo is two, which doubles the size.
- A 5-second mono clip at 8 kHz and 8 bits: $8000 \times 8 \times 5 \times 1 = 320\,000$ bits, which is 40,000 bytes.
声音文件大小
以位计的大小按下式计算:
- 单声道是一个声道;立体声是两个,大小翻倍。
- 一段 5 秒、8 kHz、8 位的单声道片段:$8000 \times 8 \times 5 \times 1 = 320\,000$ 位,即 40,000 字节。
A 5-second mono clip is sampled at $8000\ \text{Hz}$ with 8-bit resolution. What is its size in bits? · 一个 5 秒的单声道片段以 $8000\ \text{Hz}$、8 位分辨率采样。它的大小以位为单位是多少?
size · 大小 $= 8000 \times 8 \times 5 \times 1 = 320\,000$ bits. · 大小 $= 8000 \times 8 \times 5 \times 1 = 320\,000$ 位。
Sound file size in bits = sampling rate x sampling resolution x duration x ____. · 以位计的声音文件大小 = 采样率 x 采样分辨率 x 时长 x ____。
Mono is one channel, stereo two. Leaving the channel count out halves every stereo answer. · 单声道是一个声道,立体声是两个。漏掉声道数会让每个立体声答案减半。
Worked example: a CD track
- A 10-second stereo clip is recorded at CD quality: 44.1 kHz, 16-bit. Find its size in MiB.
- $44\,100 \times 16 \times 10 \times 2 = 14\,112\,000$ bits.
- Divide by 8 for bytes: $1\,764\,000$. Divide by $1024^2$: 1.68 MiB.
- Every mark is in the working: the four factors, one division by 8, and a stated prefix. Forgetting the two channels halves the answer.
例题:一段 CD 音轨
- 一段 10 秒的立体声片段以 CD 品质录制:44.1 kHz、16 位。求它以 MiB 计的大小。
- $44\,100 \times 16 \times 10 \times 2 = 14\,112\,000$ 位。
- 除以 8 得字节:$1\,764\,000$。除以 $1024^2$:1.68 MiB。
- 每一分都在过程里:四个因子、一次除以 8、说明前缀。漏掉两个声道会让答案减半。
A 10-second stereo clip is sampled at 44100 Hz with 16-bit resolution. How many bits does it need? · 一段 10 秒的立体声片段以 44100 Hz、16 位分辨率采样。需要多少位?
44100 x 16 x 10 x 2 = 14,112,000 bits, about 1.68 MiB. Leaving out the two channels halves the answer. · 44100 x 16 x 10 x 2 = 14,112,000 位,约 1.68 MiB。漏掉两个声道会让答案减半。
A stereo recording is twice the size of the same mono recording at the same rate and resolution. · 在相同采样率和分辨率下,立体声录音是同样单声道录音的两倍大。
Channels multiply the size: two channels, twice the samples. The channel count belongs in the formula. · 声道数乘进大小:两个声道,两倍的采样。声道数属于公式的一部分。
Changing the settings
- Higher sampling rate: higher pitches are captured, so the sound is more accurate, but the file is larger.
- Higher sampling resolution: each amplitude is stored in finer steps, so there is less quantisation noise 量化噪声, the roughness introduced when a smooth value is rounded to the nearest level; the file is larger.
- Lower of either: a smaller file and an audible loss of quality. As with images, answer in these terms, not "it sounds bad".
改变设置
- 提高采样率:更高的音高被捕捉,所以声音更准确,但文件更大。
- 提高采样分辨率:每个振幅以更细的台阶存储,所以量化噪声(quantisation noise)更少——把平滑的值舍入到最近电平时引入的粗糙感;文件更大。
- 两者降低:文件更小,质量有可听出的损失。和图像一样,用这些术语回答,不要说"听起来很糟"。
Increasing the sampling rate will: · 增加采样率将:
A higher sampling rate captures higher frequencies (better quality) but increases the file size. · 一个更高的采样率捕获更高的频率(更好的质量)但增加文件大小。
A higher sample resolution (more bits per sample) reduces quantisation noise. · 一个更高的采样分辨率(每个样本更多的位)减少量化噪声。
More bits give finer amplitude steps, so each sample is closer to the true value — less quantisation noise (but a larger file). · 更多的位给出更细的振幅步进,所以每个样本更接近真实值——更少的量化噪声(但更大的文件)。
Match each change to its main effect. · 把每种改变与它的主要影响配对。
Rate decides which pitches survive; resolution decides how precisely each sample is stored. · 采样率决定哪些音高能保留;分辨率决定每个采样存得多精确。
Worked example: how fast must you sample?
- The sampling rate must be at least twice the highest frequency you want to keep.
- Human hearing reaches about 20 kHz. What is the minimum rate to capture it? $2 \times 20\,000 = 40\,000$ Hz, which is why CD chose 44.1 kHz, a little above the minimum.
- A telephone samples at 8 kHz. What is the highest frequency it can carry? Half of 8 kHz, so 4 kHz, and in practice about 3.4 kHz. Speech fits; music does not.
例题:必须采样多快?
- 采样率必须至少是你想保留的最高频率的两倍。
- 人的听觉到约 20 kHz。捕捉它的最低采样率是多少?$2 \times 20\,000 = 40\,000$ Hz,这就是 CD 选 44.1 kHz 的原因,略高于最低要求。
- *电话以 8 kHz 采样。它能承载的最高频率是多少?*8 kHz 的一半,即 4 kHz,实际约 3.4 kHz。语音装得下;音乐装不下。
To capture frequencies up to $20\,000\ \text{Hz}$, what is the minimum sampling rate, in Hz? · 要捕获高达 $20\,000\ \text{Hz}$ 的频率,最小采样率以 Hz 为单位是多少?
The sampling rate must be at least twice the highest frequency: $2 \times 20\,000 = 40\,000\ \text{Hz}$. · 采样率必须至少是最高频率的两倍:$2 \times 20\,000 = 40\,000\ \text{Hz}$。
A telephone samples at 8000 Hz. What is the highest frequency, in Hz, it can in principle carry? · 电话以 8000 Hz 采样。它原则上能承载的最高频率是多少赫兹?
Half the sampling rate, since the rate must be at least twice the highest frequency kept. In practice a phone carries about 3.4 kHz. · 采样率的一半,因为采样率必须至少是所保留最高频率的两倍。实际上电话承载约 3.4 kHz。
Trading quality against size
- Speech for a voice note: a low rate and resolution are enough, and the file is small.
- Music for distribution: CD settings or better, because listeners hear the difference.
- Live streaming: the rate and resolution are chosen to fit the available bandwidth, then lossy compression removes more.
- A "justify" answer names the setting, the effect on quality and the effect on size, tied to the use.
用质量换大小
- 语音留言的语音:低采样率和低分辨率就够了,文件也小。
- 用于发行的音乐:CD 设置或更好,因为听众听得出差别。
- 直播流:采样率和分辨率按可用带宽选择,再由有损压缩去掉更多。
- "论证"的答案要说出设置、它对质量的影响和对大小的影响,并与用途联系起来。
Marks that slip away
- Rate is samples per second; resolution is bits per sample. The question names one.
- Stereo is two channels. Leaving the channel count out of the formula halves every stereo answer.
- The minimum rate is twice the highest frequency wanted, not the same as it.
- Say quantisation noise, not "it sounds fuzzy". The syllabus term carries the mark.
容易丢掉的分
- 采样率是每秒的采样数;分辨率是每个采样的位数。题目会点名其一。
- 立体声是两个声道。公式里漏掉声道数会让每个立体声答案减半。
- 最低采样率是想保留的最高频率的两倍,不是与它相同。
- 要说量化噪声,不要说"听起来模糊"。大纲术语才带分。
You've got it
- sampling converts continuous analogue data into digital data by measuring the amplitude at regular intervals
- sampling rate = samples per second (CD 44.1 kHz), sampling resolution = bits per sample (CD 16)
- size in bits = rate $\times$ resolution $\times$ duration $\times$ channels; stereo doubles it
- higher rate or resolution means better accuracy, less quantisation noise and a larger file; the rate must be at least twice the highest frequency kept
你掌握了
- 采样通过等间隔测量振幅把连续的模拟数据转换成数字数据
- 采样率 = 每秒采样数(CD 44.1 kHz),采样分辨率 = 每采样位数(CD 16)
- 以位计的大小 = 采样率 $\times$ 分辨率 $\times$ 时长 $\times$ 声道数;立体声翻倍
- 提高采样率或分辨率意味着更准确、更少量化噪声和更大的文件;采样率必须至少是所保留最高频率的两倍