Further Probability & Statistics
A-Level Further Mathematics Topic 4 8:08 English narration · English + 中文 subtitles burned in
Chapters
Transcript
Dublin, nineteen-oh-eight.
都柏林,一九〇八年。
Inside the Guinness brewery, a young chemist named William Gosset faces a very practical problem.
在健力士啤酒厂里,一位名叫威廉·戈塞特的年轻化学家遇到了一个很现实的问题。
He can test only a few batches of barley at a time, never thousands.
他每次只能检验少数几批大麦,绝不可能有成千上万批。
But the normal curve, the standard tool, assumed huge samples.
可当时通用的工具——正态曲线,却假设样本很大。
So Gosset built a sharper tool, made for small data.
于是戈塞特造出了一件更精准、专为小数据设计的工具。
His employer forbade publishing, so he signed his work with one modest word: Student.
雇主禁止他署名发表,于是他只用了一个谦逊的名字署名:Student(学生)。
That new curve is one of several sharper tools we meet today, in Further Probability and Statistics.
这条新曲线,只是我们今天在"进阶概率统计"中要认识的若干更精准工具之一。
We will treat continuous distributions with their cumulative functions, test claims from small or awkward samples, and pack a whole distribution into a single function.
我们将用累积分布函数来处理连续型分布,从小样本或不规则样本中检验各种说法, 还会把一整个分布装进一个函数里。
Let's begin.
让我们开始吧。
A continuous random variable is described by a density curve, and probability is the area under it.
连续型随机变量用一条密度曲线来描述,而概率就是它下方的面积。
Now meet its partner, the cumulative function.
现在来看它的伙伴——累积分布函数。
The density gives the rate; the cumulative gives the running total — the chance of being at most a value, climbing from zero up to one.
密度给出的是变化率,累积给出的是累加的总量——即取值不超过某个数的概率,从零一路攀升到一。
The two are linked: the cumulative is the area so far, and its slope brings back the density.
两者紧密相连:累积就是到目前为止的面积,而它的斜率又还原成密度。
This makes percentiles easy — the median is the fiftieth percentile, where the cumulative reaches zero point five.
这样求百分位数就很容易了—— 中位数就是第五十百分位,累积达到零点五的地方。
For the density one half x, the cumulative is a quarter x squared.
对于密度二分之一 x,累积是四分之一 x 的平方。
Set it equal to one half, and the median is root two, about one point four one.
令它等于二分之一,中位数就是根号二,约为一点四一。
Sometimes we need the distribution of a related variable built from the old one.
有时我们需要由旧变量构造出的相关变量的分布。
Say Y equals some function of X.
设 Y 等于 X 的某个函数。
The method is to work through the cumulative function: the chance Y is at most a value equals the chance X lands in the matching range.
方法是从累积分布函数入手:Y 不超过某个值的概率,等于 X 落入相应范围的概率。
Then differentiate the cumulative to get back the new density.
然后对累积求导,就得回新的密度。
Here is an example.
看一个例子。
Let Y be X squared, where X has density one half x on zero to two.
设 Y 等于 X 的平方,其中 X 在零到二上的密度是二分之一 x。
Working through, the cumulative of Y comes out as a quarter y, so its density is a constant, one quarter.
推导下来,Y 的累积是四分之一 y,所以它的密度是常数四分之一。
That means Y is uniform on zero to four.
这意味着 Y 在零到四上服从均匀分布。
Back to Student's curve.
回到"学生"的这条曲线。
When your sample is small, and you do not know the population's spread, the normal curve is too confident.
当样本很小、又不知道总体的离散程度时,正态曲线过于自信了。
The t-distribution fixes that.
t 分布纠正了这一点。
Compare them.
把它们放在一起比较。
The normal is tall and narrow.
正态又高又窄。
The t-distribution sits lower, with heavier tails — it admits more uncertainty.
t 分布则更低,尾部更重—— 它容纳了更多的不确定性。
So its critical values are a little larger, which makes your test properly cautious.
所以它的临界值稍大一些,使你的检验更为谨慎。
And it comes in a family: you pick the member using the degrees of freedom, which is one less than the sample size.
而且它是一整个家族:你用自由度来挑选其中一员,自由度等于样本量减一。
The t-distribution also gives a confidence interval for the mean.
t 分布同样能给出均值的置信区间。
It looks like the earlier one, but with a t-value in place of one point nine six, and the sample's own standard deviation.
它和之前那个很像,只是用一个 t 值取代了固定的一点九六, 并用样本自己的标准差。
It reaches out t standard errors on each side of the sample mean.
它从样本均值向两侧各伸出 t 个标准误。
Take a sample of ten, with mean fifty and standard deviation four.
取一个大小为十的样本, 均值为五十,标准差为四。
The t-value for nine degrees of freedom is two point two six two.
九个自由度对应的 t 值是二点二六二。
Work it through, and the interval runs from forty-seven point one to fifty-two point nine.
算下来, 区间从四十七点一到五十二点九。
Often we compare two groups, not one.
我们常常要比较两组,而不只是一组。
Choose the test that fits.
选择合适的检验。
For two independent groups, use a two-sample t-test.
对于两个独立的组,用双样本 t 检验。
For before-and-after pairs on the same subjects, use a paired t-test — a matched-pairs test that looks at each difference.
对于同一批对象的前后配对,用配对 t 检验——即匹配对检验,它看的是每一对的差。
For large samples, a normal test will do.
对于大样本,用正态检验就行。
And with the two-sample t, when the groups share a spread, you first form a pooled estimate of the shared variance — a weighted average of the two sample variances.
而用双样本 t 时,若两组的离散程度相同, 就要先给出合并估计的共同方差——即两个样本方差的加权平均。
Picking the right test is half the marks.
选对检验,就是一半的分数。
Next, a different question: does the data fit a model at all?
接下来是另一个问题:数据究竟符不符合某个模型?
The chi-squared goodness of fit test compares the counts you observed with the counts a theoretical distribution expected.
卡方拟合优度检验把你观测到的计数, 与理论分布所期望的计数作比较。
Square each gap, divide by the expected count, and add them up.
把每个差平方,除以期望计数,再全部相加。
Take four equal categories with counts twenty, thirty, twenty-five, twenty-five — each expected to be twenty-five.
取四个等可能的类别,计数为二十、三十、二十五、二十五——每个的期望都是二十五。
The chi-squared value works out to two.
算出的卡方值是二。
Now compare it against the chi-squared curve, which is skewed to the right, with the top five percent shaded as the rejection tail.
再把它和卡方曲线比较,这条曲线向右偏斜,顶端百分之五被标为拒绝尾。
The critical value at three degrees of freedom is seven point eight one five.
三个自由度对应的临界值是七点八一五。
Our statistic, two, sits far below it — so we do not reject the model.
我们的统计量二,远远低于它——所以我们不拒绝这个模型。
The same chi-squared idea also tests whether two variables are independent.
同样的卡方思路,也能检验两个变量是否相互独立。
Arrange the counts in a contingency table, with row and column totals.
把计数排进一张列联表, 并列出行合计与列合计。
Under independence, the expected count in each cell is its row total times its column total, divided by the grand total.
在独立的假设下,每个单元格的期望计数,等于它所在行的合计乘以所在列的合计, 再除以总计。
Then compare observed with expected exactly as before.
然后像之前一样,把观测与期望作比较。
The degrees of freedom are rows minus one, times columns minus one.
自由度等于行数减一,乘以列数减一。
What if the data clearly is not normal?
如果数据显然不是正态的,怎么办?
Then we reach for a non-parametric test, which makes no such assumption.
那我们就动用非参数检验,它不做那样的假设。
The sign test is the simplest: it just counts how many values fall above and below a proposed centre.
符号检验最简单:它只数有多少个值落在某个假定中心的上方和下方。
The Wilcoxon signed-rank test goes further, using the sizes of the differences, not only their signs.
威尔科克森符号秩检验更进一步, 不仅用差的符号,还用差的大小。
And the Wilcoxon rank-sum test compares two separate samples.
而威尔科克森秩和检验比较的是两个独立的样本。
One caution: the Wilcoxon tests are valid only when the distribution is symmetrical.
有一点要注意:威尔科克森检验只有在分布对称时才有效。
Let's run a sign test.
我们来做一次符号检验。
Is the median five?
中位数是五吗?
In a sample of ten, nine values lie above five and one lies below.
在十个值的样本里,九个在五以上,一个在五以下。
Count the pluses and minuses.
数一数正号和负号。
If the median really were five, each value is equally likely to fall above or below, so the number above follows a binomial, with ten trials and one-half.
如果中位数真的是五,每个值落在上方或下方的机会均等, 那么上方的个数就服从二项分布,试验十次,概率二分之一。
Nine or more above is a lot.
九个或更多在上方,是很极端的。
Its probability is eleven over one thousand and twenty-four, about zero point zero one one.
它的概率是一千零二十四分之十一,约为零点零一一。
For a two-tailed test we compare with zero point zero two five.
对于双尾检验,我们与零点零二五比较。
Since it is smaller, we reject: the median is not five.
因为它更小,我们拒绝:中位数不是五。
Our last tool is a clever piece of algebra: the probability generating function.
我们的最后一件工具是一段巧妙的代数:概率母函数。
It packs an entire discrete distribution into a single function of a helper variable t.
它把一整个离散分布, 装进一个关于辅助变量 t 的函数里。
Each probability rides on its own power of t.
每个概率都搭在它自己的 t 的幂次上。
Why bother?
何必如此?
Because its derivatives hand you the summary numbers.
因为它的导数会把汇总的数字直接交给你。
The mean is simply the first derivative at t equals one.
均值就是它在 t 等于一处的一阶导数。
And the variance comes from the second derivative at one, plus the first, minus the first squared.
而方差来自在一处的二阶导数,加上一阶导数,再减去一阶导数的平方。
Let's use it.
我们来用一用。
A variable takes zero, one, or two, with probabilities one-half, zero point three, and zero point two.
一个变量取零、一或二,概率分别是二分之一、零点三和零点二。
Build the generating function: one-half, plus zero point three t, plus zero point two t squared.
构造母函数:二分之一,加零点三 t,加零点二 t 的平方。
Differentiate: zero point three plus zero point four t.
求导:零点三加零点四 t。
Put t equal to one, and the mean is zero point seven.
令 t 等于一,均值就是零点七。
One last gift: for independent variables added together, you simply multiply their generating functions.
最后一份礼物:对于相加的独立变量, 你只需把它们的母函数相乘。
A hard sum becomes an easy product.
一个困难的求和,就变成了一个简单的乘积。
Before you go, four ways to keep your marks.
结束之前,四个保分要点。
First, a density must integrate to one, and the mean is the integral of x times the density.
第一,密度必须积分为一,而均值是 x 乘以密度的积分。
Second, use the t-distribution for a small sample with unknown spread, and always state the degrees of freedom.
第二,遇到方差未知的小样本,就用 t 分布,并且始终写出自由度。
Third, in a chi-squared test, combine any classes whose expected count is below five before you start.
第三,做卡方检验时,先把期望计数小于五的类别合并,再开始计算。
Fourth, for every test, state both hypotheses, and give your conclusion in context — never just the word reject.
第四,每一个检验,都要写出两个假设,并结合情境给出结论——绝不只是"拒绝"两个字。