Skip to content

Inference for Quantitative Data: Means

AP Statistics Topic 7 7:38 English narration · English + 中文 subtitles burned in

space play · ←/→ 5s · j/l 10s · f fullscreen · ,/. speed

Chapters

Transcript
You already know the heart of inference. 你已经掌握了推断的核心。
Take a sample, find its mean, then do it again and again. 抽一个样本,算出它的平均数,然后一次又一次地重复。
Even when the population is lopsided, those sample means pile up into a normal curve, centred on the true mean. 即使总体是偏斜的,这些样本平均数也会堆成一条正态曲线,以真实平均数为中心。
That is the Central Limit Theorem. 这就是中心极限定理。
But there is a problem. 但有一个问题。
To use that curve you need the spread of the whole population, and nobody ever tells you what it is. 要用这条曲线,你需要知道整个总体的离散程度,而没有人会告诉你它是多少。
So we use the spread of our own sample instead. 于是我们改用自己样本的离散程度。
Every sample gives a different spread — sometimes too small, sometimes too big. 每个样本给出的离散程度都不一样——有时偏小,有时偏大。
Dividing by a number that wobbles makes extreme results more common. 用一个会摇摆的数去除,就让极端结果更容易出现。
The pile is no longer normal: it has a lower peak and heavier tails. 这堆数据不再是正态的: 它的峰更低,尾更厚。
We call this shape the t-distribution, and every mean in this unit lives on it. 我们把这个形状叫做 t 分布,本单元里每一个平均数都建立在它上面。
So unit seven is inference for means: an interval that estimates a mean, a test that weighs a claim, and the same two tools again for two groups. 所以第七单元讲的是关于平均数的推断:用一个区间去估计平均数,用一个检验去衡量论断, 然后用同样这两件工具去比较两个组。
Let's begin. 让我们开始吧。
Here are the two curves together. 把两条曲线放在一起看。
The blue one is the normal curve. The orange one is a t-curve. 蓝色的是正态曲线,橙色的是 t 曲线。
Same centre, same symmetry, but a lower peak and heavier tails, so it reaches further out. 中心相同,也同样对称, 但它的峰更低、尾更厚,所以伸得更远。
How heavy the tails are depends on the degrees of freedom, which is just the sample size minus one. 尾巴有多厚,取决于自由度, 而自由度就是样本量减一。
A small sample gives fat tails; as the sample grows, the t-curve settles back onto the normal. 样本小,尾巴就厚;样本变大,t 曲线就慢慢贴回正态曲线。
Here is the interval for a mean. 这就是平均数的置信区间。
Start at the sample mean, then add and subtract a margin of error. 从样本平均数出发,再加上和减去一个误差幅度。
The margin is a critical value times the standard error, and the standard error is the sample's own spread divided by the square root of the sample size. 误差幅度等于临界值乘以标准误,而标准误就是样本自身的离散程度除以样本量的平方根。
The critical value now comes from the t-curve, with degrees of freedom one less than the sample size. 现在临界值来自 t 曲线,自由度比样本量小一。
First, three conditions: a random sample, at most ten percent of the population, and a roughly normal shape — the Normal/Large Sample condition: population normal, or a large sample of at least thirty. 首先要检查三个条件:样本随机, 不超过总体的百分之十,以及大致正态形状——正态或大样本条件:总体正态, 或者大样本至少三十。
Let's do one. 我们来做一道。
Twenty five batteries are tested at random; their mean life is fifty hours, and the sample spread is eight hours. 随机测试二十五节电池,平均寿命是五十小时,样本离散程度是八小时。
Find a ninety five percent interval. 求百分之九十五的置信区间。
First the standard error: eight over the square root of twenty five, which is one point six. 先算标准误:八除以二十五的平方根,等于一点六。
Next the critical value: with twenty four degrees of freedom it is two point zero six four. 再看临界值:自由度为二十四时,它是二点零六四。
Multiply, and the margin of error is three point three. 相乘,误差幅度就是三点三。
So the interval runs from forty six point seven to fifty three point three hours. 所以这个区间从四十六点七到五十三点三小时。
Say it properly: we are ninety five percent confident that the true mean life of these batteries is between forty six point seven and fifty three point three hours. 要这样规范表述:我们有百分之九十五的把握,认为这些电池的真实平均寿命 在四十六点七到五十三点三小时之间。
But what does that confidence mean? 但这份把握到底是什么意思?
Every new sample gives a new interval. 每一个新样本都会给出一个新区间。
Most capture the true mean; a few miss completely. 大多数抓住了真实平均数,少数完全落空。
The ninety five percent describes the method over many samples, not your one interval. 这百分之九十五描述的是这套方法在大量样本上的表现,而不是你手上这一个区间。
An interval also settles arguments. 区间还能用来判定论断。
The maker claims the true mean life is forty five hours. 厂家声称真实平均寿命是四十五小时。
Forty five sits outside our interval, so the data give evidence against it. 四十五落在我们的区间之外,所以数据是反对它的证据。
Another engineer says forty eight; that one is inside, so it stays plausible. 另一位工程师说是四十八; 这个值在区间之内,所以它仍然合理。
And remember two rules: a bigger sample makes the interval narrower, more confidence makes it wider. 还要记住两条规律:样本越大,区间越窄; 要求的把握越高,区间就越宽。
Now test that claim properly. 现在正式检验这个论断。
The null hypothesis says the true mean is forty five hours; the alternative says it is not. 原假设说真实平均寿命是四十五小时,备择假设说不是。
The conditions are the same three as before. 条件和前面那三条一样。
The statistic counts how many standard errors our sample mean sits from the claimed value. 这个统计量数的是:我们的样本平均数离被声称的值有多少个标准误。
On top, fifty minus forty five is five. Underneath sits the standard error, one point six. 分子上,五十减四十五等于五;分母是标准误,一点六。
Five divided by one point six is three point one three, with twenty four degrees of freedom. 五除以一点六等于三点一三, 自由度是二十四。
How surprising is that? 这有多让人意外?
Suppose the claim were true. Results would scatter like this curve, and ours sits far out here. 假设那个论断为真,结果就会像这条曲线这样散布, 而我们的结果落在很远的这里。
The p-value is the area in the tail beyond our result: the chance of a result at least this extreme, if the claim were true. P 值就是我们的结果之外那块尾部面积: 在论断为真的前提下,出现至少这么极端结果的概率。
Our test is two-sided, so we count both tails — under one percent. 我们的检验是双侧的,所以两条尾巴都要算——总共不到百分之一。
Now write it the way the exam wants. 现在按考试的要求把它写出来。
Compare the p-value with the significance level, usually five percent. 把 P 值和显著性水平比较,通常是百分之五。
Ours is under one percent, well below it, so we reject the null hypothesis. 不到百分之一更小,所以我们拒绝原假设。
Then the sentence, in context: there is convincing evidence that the true mean life is not forty five hours. 然后写出结合情境的结论句: 有令人信服的证据表明,真实平均寿命不是四十五小时。
Notice that forty five sat outside our interval too, so the interval and the test agree. 注意,四十五也落在我们的区间之外, 所以区间和检验的答案一致。
And never write that you accept the null hypothesis. 另外,永远不要写你接受原假设。
But we could still be wrong, and that is not bad arithmetic — it is the randomness in the data. 但我们仍然可能出错,而这不是算错了——是数据本身的随机性。
Rejecting a claim that was actually true is a false alarm. Failing to reject a claim that was actually false is a missed effect. 拒绝了一个本来为真的论断,是一次虚惊;没能拒绝一个本来为假的论断,是一次漏检。
Wherever you draw the line, you trade one against the other. 无论你把界线画在哪里,都是在用一种错误去换另一种。
Most exam questions compare two groups. 大多数考题要比较两个组。
First, the design: two separate random samples, or one group measured twice? 先看设计:是两个独立的随机样本,还是同一组被测了两次?
For two independent samples, the estimate is the difference of the sample means, and the standard errors combine — square each, add, then take the square root. 如果是两个独立样本,估计值就是两个样本平均数之差,标准误要合并—— 各自平方,相加,再开平方。
Variances add; they never subtract. 方差是相加的,绝不相减。
Here the interval runs from three point seven four to ten point two six, entirely above zero, so brand one lasts longer. 这里的区间从三点七四到十点二六,整个都在零的上方,所以第一个品牌更耐用。
The test works the same way. 检验的做法一样。
The null hypothesis says the two means are equal, and the statistic is that difference divided by the same combined standard error. 原假设说两个平均数相等,统计量就是这个差除以同样那个合并的标准误。
Here the means are eighty five and seventy eight, so the top is seven, the standard error is about one point four one, and the statistic is about four point nine five. 这里两个平均数是八十五和七十八,所以分子是七,标准误大约是一点四一, 统计量约为四点九五。
The p-value is near zero point zero zero one, so we reject. P 值接近零点零零一,所以我们拒绝原假设。
Do not pool the spreads for means, and let technology give the degrees of freedom. 做平均数时不要合并离散程度,自由度交给计算器或软件给出。
One design catches students out every year. 有一种设计每年都让学生栽跟头。
That design is paired data. 那种设计就是配对数据。
If the same subjects are measured twice — before and after a course, say — the two columns are not independent. Each person is a pair. 如果同一批对象被测了两次—— 比如课程前和课程后——那两列数据并不独立,每个人本身就是一对。
So do not reach for the two-sample test. 所以不要去用双样本检验。
Take the difference inside each pair first, then run the ordinary one-sample test on that single column of differences. 先在每一对内部求出差值,然后对这一列差值做普通的单样本检验。
Look for the natural pairing: the same person, the same day, matched twins. 要留意天然的配对:同一个人、同一天、配对的双胞胎。
The hardest exam skill is choosing the procedure before you calculate anything. 考试中最难的技能,是在动笔计算之前先选对方法。
Three questions decide it. The goal: to estimate, which needs an interval, or to weigh a claim, which needs a test? 三个问题就能决定: 目标是什么——是估计,那就用区间;还是衡量一个论断,那就用检验?
How many groups: one, two, or one group measured twice? 有几个组——一个、两个,还是同一组测两次?
And is it a mean or a proportion — because a mean means t. 变量是平均数还是比例——平均数就用 t。
Three marks students throw away every year. 学生每年都会丢的三分。
First, a mean uses t, not z, and you must state the degrees of freedom. 第一,平均数用 t,不用 z,而且必须写出自由度。
Second, check the conditions with the numbers from this question, not as a memorised list. 第二,用这道题里的数字去检查条件,而不是背一串条件。
Third, look hard for pairing before you pick a two-sample test. 第三,在选双样本检验之前,仔细找有没有配对。
Get those right, and this unit is yours. 把这三点做对,这个单元就是你的了。

Log in or create account

IGCSE, A-Level & AP