Probability & Statistics 2
A-Level Mathematics Topic 6 10:06 English narration · English + 中文 subtitles burned in
Chapters
Transcript
Cambridge, the nineteen-twenties.
剑桥,二十世纪二十年代。
At a garden party, a lady says something surprising.
在一次花园聚会上,一位女士说了句令人惊讶的话。
She can taste, she claims, whether the milk was poured into the cup before the tea, or after.
她声称,她能尝出牛奶是在倒茶之前放入杯中,还是在倒茶之后。
Everyone laughs.
众人大笑。
But one scientist does not.
但有一位科学家没有笑。
Instead of arguing, he asks a better question: how could we actually test that?
他没有争辩,而是提出了一个更好的问题:我们究竟该如何检验这一点?
That question is the heart of today's topic: Probability and Statistics, part two.
这个问题正是今天主题的核心:概率统计,第二部分。
We will learn to model random events, to describe uncertainty with numbers, and finally to test a claim against the evidence.
我们将学习为随机事件建立模型, 用数字描述不确定性,最后用证据去检验一个说法。
Let's begin.
让我们开始吧。
Some events are rare, random, and independent.
有些事件是稀有的、随机的、彼此独立的。
Think of calls to a help desk, or particles from a decaying atom.
想想打到客服台的电话,或衰变原子放出的粒子。
The Poisson distribution counts how many happen in a fixed interval.
泊松分布数的是在一段固定时间内发生了多少次。
It needs just one number, the average rate.
它只需要一个数字——平均发生率。
And here is its special feature: the mean and the variance are both equal to that rate.
它有一个特别之处:均值和方差都等于这个发生率。
Suppose the average rate is three. Then the chance of exactly two events is e to the minus three, times three squared, over two factorial — about zero point two two four.
设平均发生率为三, 那么恰好发生两次的概率是 e 的负三次方,乘以三的平方,除以二的阶乘——约为零点二二四。
The Poisson model has two useful friends.
泊松模型有两个有用的伙伴。
Take a binomial with a large number of trials, but a tiny chance of success each time. It is much easier to treat it as Poisson instead.
取一个二项分布,试验次数很大,但每次成功的概率很小, 这时把它当作泊松分布来处理会容易得多。
And when the average grows large, the Poisson becomes smooth and symmetric. Then you can use a normal approximation — a bell-shaped curve with a small continuity correction.
而当平均值变得很大时,泊松分布会变得平滑对称, 于是你可以用一条钟形的正态曲线来近似它,再加上一个小小的连续性修正。
Binomial, to Poisson, to normal — one family, chosen to fit the situation.
二项、到泊松、再到正态——它们是同一个家族,按具体情形来选用。
Often we build linear combinations of random variables.
我们常常用旧的变量构造出新的变量。
If you stretch or shift a single variable, the expectation follows the same rule, but the variance is multiplied by the scale factor squared.
如果你把一个变量拉伸或平移,均值按同样的规则变化, 但方差要乘以比例因子的平方。
For two independent variables added together, the means simply add.
对于两个相加的独立变量,均值直接相加。
And the variances always add too — never subtract, even when you are taking a difference.
而方差也总是相加——绝不相减,哪怕你算的是差。
Try one. A variable has mean five and variance four.
试一试:某变量的均值是五,方差是四。
Then three times it, minus one, has mean fourteen, and variance thirty-six.
那么它的三倍减一,均值是十四,方差是三十六。
Those rules gave you a mean and a variance — but not a distribution, and that is not the same thing.
刚才那些公式给了你均值和方差——但没有给分布,这是两回事。
You cannot standardise a variable, or look anything up in a table, until you know what shape it has.
在你知道它是什么形状之前,既不能标准化,也不能查表。
Two facts close that gap.
有两条结论补上了这个缺口。
First, a normal stays normal: if X is normal with mean mu and variance sigma squared, then a X plus b is normal too, with mean a mu plus b and variance a squared sigma squared.
第一,正态还是正态:如果 X 服从均值为 mu、方差为 sigma 平方的正态分布, 那么 a X 加 b 也是正态分布,均值是 a mu 加 b,方差是 a 平方乘 sigma 平方。
That is what lets you standardise a scaled or shifted normal and read the table as usual.
正因如此,你才能把一个被缩放或平移过的正态变量标准化,照常查表。
Second, independent Poissons add: if X is Poisson with rate lambda one and Y is Poisson with rate lambda two, then X plus Y is Poisson with rate lambda one plus lambda two — the rates simply add.
第二,独立的泊松可以相加:如果 X 服从参数为 lambda 一的泊松分布, Y 服从参数为 lambda 二的泊松分布,那么 X 加 Y 服从参数为 lambda 一加 lambda 二的泊松分布—— 速率直接相加。
So two independent sources of rare events can be treated as one.
所以两个独立的稀有事件来源可以当成一个来处理。
And a warning, because it is the obvious next guess: there is no such rule for the binomial in general.
还要提醒一句,因为这是最容易顺手一猜的:二项分布一般没有这样的结论。
Two binomials add only when they share the same probability of success.
两个二项分布只有在成功概率相同时才能相加。
Not everything comes in whole numbers.
并非所有量都取整数。
A continuous variable, like a height or a waiting time, is described by a smooth curve called a density function.
像身高、等待时间这样的连续型变量,用一条平滑的曲线来描述, 这条曲线叫概率密度函数。
Probability is not the height of the curve. It is the area under it, between two values.
概率不是曲线的高度,而是它在两个数值之间下方的面积。
Because the variable takes some value for certain, the total area is one.
因为变量必定取到某个值,所以曲线下方的总面积是一。
To find the mean, you weight each value by the density and integrate.
要求均值,就把每个值按密度加权后积分。
For example, take the density one half x, between zero and two. Its mean works out to four thirds.
例如,取密度二分之一 x,范围从零到二,它的均值算出来是三分之四。
Integrating the density from the left gives the cumulative distribution function: the median is where it reaches one half, and other percentiles where it reaches any other fraction.
把密度从左边开始积分,就得到累积分布函数: 中位数就是它达到二分之一的地方,其他百分位数则是它达到相应分数的地方。
The mean is only half the summary.
均值只是概括的一半。
The variance of a continuous variable comes from an integral too, and it has the same shape as before: the integral of x squared times the density, minus the square of the mean.
连续型随机变量的方差同样来自积分,而且形状和以前一样: x 平方乘以密度函数的积分,再减去均值的平方。
Compare that with the discrete formula and it is the same expression, with a sum replaced by an integral.
拿它和离散情形的公式比一比,就是同一个式子,只是把求和换成了积分。
So work the integral on the same density we just used, one half x between zero and two.
那就用刚才那个密度函数来算这个积分:零到二上的二分之一 x。
x squared times one half x is one half x cubed, which integrates to x to the fourth over eight; putting in the limits gives two.
x 平方乘二分之一 x 就是二分之一 x 立方,积分得到八分之 x 的四次方;代入上下限,得到二。
Now subtract the square of the mean.
现在减去均值的平方。
The mean was four thirds, so its square is sixteen ninths, and two minus sixteen ninths is two ninths.
均值是三分之四,平方就是九分之十六,二减九分之十六等于九分之二。
Take the square root and the standard deviation is about nought point four seven one.
开平方,标准差大约是零点四七一。
Notice what changed between the two calculations: nothing except one more factor of x inside the integral.
注意两次计算之间变了什么:除了积分里多乘一个 x,别的什么都没变。
If you can find the mean of a density, you can already find its variance.
只要你会求一个密度函数的均值,你就已经会求它的方差了。
A random sample needs randomness: every member of the population must have a fair chance of being chosen.
随机样本需要随机性:总体中的每一个成员都必须有公平的机会被选中。
Now sampling and estimation.
现在来讲抽样与估计。
We usually cannot measure a whole population — every voter, every light bulb.
我们通常无法测量整个总体——每一位选民、每一只灯泡。
So we study a sample instead.
于是我们改为研究一个样本。
For the sample to be trustworthy it must be random: every member needs a fair chance of being chosen.
样本要可信,就必须是随机的:每个成员都要有被选中的公平机会。
Ask only volunteers, or the first twenty people through the door: those methods are unsatisfactory, and you get a biased sample that leans one way.
只问志愿者, 或者只问最先进门的二十个人,你得到的就是一个偏向一方的有偏样本。
The cure is random numbers: number everyone, then let chance pick the sample.
解决办法是随机数:给每个成员编号,再让随机来挑选样本。
Here is where it becomes powerful.
精彩之处就在这里。
The sample mean is itself a random variable.
样本均值本身也是一个随机变量。
Take another sample, and you get a slightly different mean.
再取一个样本,你会得到略有不同的均值。
On average, the sample mean lands exactly on the population mean, so the sample gives unbiased estimates of the population mean and variance.
平均而言,样本均值正好落在总体均值上——它以真实均值为中心。
Its spread is much narrower, shrinking as the sample grows.
它的分布要窄得多, 而且随着样本变大而不断收窄。
And the Central Limit Theorem says something remarkable: whatever shape the population has, the sample mean is close to a normal curve, once the sample is large.
中心极限定理给出了一个了不起的结论: 无论总体是什么形状,只要样本足够大,样本均值都接近一条正态曲线。
And the normal curve keeps a promise.
正态曲线还信守一个承诺。
About sixty-eight percent of its area lies within one standard deviation of the centre. About ninety-five percent lies within two.
约百分之六十八的面积落在离中心一个标准差之内, 约百分之九十五落在两个标准差之内。
That ninety-five percent is the figure we lean on next — for estimates we can trust, and for tests we can defend.
接下来我们要依靠的正是这个百分之九十五—— 用它做出可信的估计,也用它支撑经得起推敲的检验。
A single sample mean is just one estimate.
单个样本均值只是一次估计。
A confidence interval turns it into a range that probably contains the truth.
置信区间把它变成一个很可能包含真值的范围。
For ninety-five percent confidence, it reaches out one point nine six standard errors on each side of the sample mean.
对于百分之九十五的置信度,它从样本均值向两侧各伸出一点九六个标准误。
More confidence means a wider interval; a bigger sample means a tighter one.
置信度越高,区间越宽;样本越大,区间越窄。
The same method gives an interval for a population proportion. An example: a sample of sixty-four has a mean of fifty, from a population with standard deviation eight.
举个例子:一个大小为六十四的样本, 均值为五十,来自一个标准差为八的总体。
The interval is fifty, plus or minus one point nine six — from forty-eight to fifty-two.
区间就是五十,加减一点九六——从四十八到五十二。
Now we can test a claim. We write two rival claims.
现在我们可以检验一个说法了。
The null hypothesis is the cautious one, usually that nothing has changed.
我们写下两个相互竞争的说法。 原假设是谨慎的那个, 通常是"没有变化"。
The alternative is what we suspect instead.
备择假设则是我们所怀疑的另一种情况。
We pick a significance level, often five percent.
我们选一个显著性水平,常取百分之五。
For a two-tailed test, that five percent is split between both tails of the normal curve, beyond plus or minus one point nine six. Then we compute a test statistic from the data.
对于双尾检验,这百分之五被分到正态曲线的两侧尾部,即超过正负一点九六之外。
If it lands in a tail — the rejection region, also called the critical region — we reject the null; anywhere else is the acceptance region.
然后我们从数据算出一个检验统计量。 如果它落在尾部——也就是拒绝域——我们就拒绝原假设。
Suppose the claim is a mean of fifty, but a sample gives fifty-two. The test statistic comes to two.
假设声称均值为五十,但样本给出五十二,检验统计量算出来是二。
Two is beyond one point nine six, so we reject: the mean really has changed.
二超过了一点九六, 所以我们拒绝:均值确实变了。
A test can be wrong in two ways.
一个检验可能以两种方式出错。
A Type one error is a false alarm: we reject the null when it was actually true.
第一类错误是虚惊一场:原假设其实为真,我们却拒绝了它。
A Type two error is a miss: we accept the null when it was actually false.
第二类错误是漏报:原假设其实为假,我们却接受了它。
We cannot remove both at once.
两者不能同时消除。
But we do control the first. The probability of a Type one error equals the significance level we chose.
但我们确实能控制第一类:犯第一类错误的概率,正好等于我们所选的显著性水平。
Choose five percent, and in the long run you raise a false alarm just five times in a hundred.
选百分之五,长期来看,你每一百次里只会虚惊五次。
So, back to the tea.
那么,回到那杯茶。
Eight cups: four with milk first, four with tea first, in a random order.
八只杯子:四只先放牛奶,四只先倒茶,顺序随机。
The null hypothesis is simple — she is only guessing.
原假设很简单—— 她只是在猜。
If that were true, the chance of sorting all eight correctly by luck alone is one in seventy, about one and a half percent.
如果真是这样,仅凭运气就把八只全部排对的概率是七十分之一,约百分之一点五。
She got all eight correct.
而她八只全部答对了。
One and a half percent is below five percent, so we reject the idea that she was guessing.
百分之一点五低于百分之五,所以我们拒绝"她在猜"这个说法。
The lady really could taste the difference.
这位女士确实尝得出差别。
That is statistics in a single cup: a claim, a model of chance, and a test.
这就是浓缩在一只杯子里的统计学:一个说法、一个随机模型、一次检验。
Before you go, four ways to keep your marks.
结束之前,四个保分要点。
First, reach for Poisson only for rare, random, independent events, and remember its mean equals the variance.
第一,只在稀有、随机、独立的事件上使用泊松分布, 并记住它的均值等于方差。
Second, when you combine independent variables, variances add — they never subtract.
第二,组合独立变量时,方差相加——绝不相减。
Third, in a hypothesis test, state both hypotheses, the significance level, and your conclusion always in context, not just the word reject.
第三,做假设检验时,写出两个假设、显著性水平,结论一定要结合情境, 而不只是"拒绝"两个字。
Fourth, in a confidence interval, use the right value — one point nine six for ninety-five percent — and say what it means in words.
第四,做置信区间时,用对数值——百分之九十五对应一点九六—— 并用文字说明它的含义。