Skip to content

Exploring Two-Variable Data

AP Statistics Topic 2 7:49 English narration · English + 中文 subtitles burned in

space play · ←/→ 5s · j/l 10s · f fullscreen · ,/. speed

Chapters

Transcript
Thirty children from one primary school. 一所小学的三十个孩子。
For each child, two numbers: shoe size, and a reading test score. 每个孩子有两个数字:鞋码,和一次阅读测验的成绩。
Plot one point per child. 每个孩子画一个点。
The pattern is hard to miss: bigger feet, better reading. 规律很难忽视:脚越大,读得越好。
The correlation is nought point eight seven, which is strong. 相关系数是零点八七,这已经很强了。
So, does a bigger foot make a child read better? 那么,脚更大就能让孩子读得更好吗?
Here is the plan. 本节的安排是这样的。
First, what relating two variables means. 首先,两个变量"相关"到底是什么意思。
Then two categorical variables, and the table that compares them. 然后是两个分类变量,以及用来比较它们的表格。
Then scatterplots and correlation. 接着是散点图与相关系数。
Then the regression line. 然后是回归线。
Then residuals and the fit. 再然后是残差,以及拟合好坏。
And last, the points that pull a line off course. 最后是那些把直线拉偏的点。
Two variables are associated when they change together. 当两个变量一起变化时,我们说它们是相关的。
The explanatory variable is what we predict from; the response variable is the outcome. 解释变量是我们用来做预测的那一个;响应变量是被预测的结果。
Study hours explain a test score, not the other way round. 是学习时间解释考试分数,而不是反过来。
The explanatory variable goes on the horizontal axis. 解释变量总是画在横轴上。
And a pattern in one sample may be real, or just a fluke. 另外,一个样本里的规律可能是真实的,也可能只是巧合。
When both variables are categories, we count them in a two-way table. 当两个变量都是分类变量时,我们用双向表来计数。
Fifty boys, fifty girls, with a pet or without. 五十个男生,五十个女生, 有宠物或没有宠物。
One table gives three relative frequencies, and the exam wants the right name. 一张表能给出三种相对频率,而考试要你叫对名字。
Joint: one cell over the grand total. 联合相对频率:一个单元格除以总计。
Marginal: a row or column total over the grand total. 边缘相对频率:一行或一列的合计除以总计。
Conditional: a cell over its own row total. 条件相对频率:一个单元格除以它自己那一行的合计。
Counts mislead when the groups differ in size, so turn each row into percentages. 当各组人数不同时,计数会误导你,所以要把每一行都换成百分比。
Divide every cell by its own row total. 用每个单元格除以它自己那一行的合计。
Among the boys, thirty of fifty: sixty percent. 男生里,五十人中有三十人:百分之六十。
Among the girls, twenty-five of fifty: fifty percent. 女生里,五十人中有二十五人:百分之五十。
Those two rows are the conditional distributions. 这两行就是条件分布。
Show them as segmented bar charts. 把它们画成分段条形图。
Matching bars would mean no association, and these differ, so the variables are associated. 如果两条完全一样,就说明没有关联; 这两条不同,所以两个变量是相关的。
When both variables are numbers we use a scatterplot, described with DUFS — direction, unusual features, form and strength — one point per individual. 当两个变量都是数值时,我们用散点图,每个个体一个点。
Describing it earns marks, and there are four things to say. 描述散点图是能拿分的,要说四件事。
Direction: positive or negative? 方向:正相关还是负相关?
Form: straight, or curved? 形态:是直线还是曲线?
Strength: tight, or loose? 强度:点贴得紧,还是散得开?
And unusual features: outliers or clusters. 还有异常特征:离群点或成簇的点。
Say all four, in context. 四点都要说,而且要结合语境。
Direction is the easiest mark on the paper. 方向是卷面上最容易拿的分。
Positive: as one variable goes up, the other tends to go up too. 正相关:一个变量增大时,另一个也倾向于增大。
Negative: as one goes up, the other goes down. 负相关:一个增大时,另一个减小。
Zero: no direction at all, just a formless cloud. 零:完全没有方向,只是一团散乱的点。
Notice the word tends - not every single point rises. 注意"倾向"这个词——并不是每一个点都在上升。
Strength gets a number: the correlation, r. 强度可以用一个数来表示:相关系数。
Near zero the cloud is round, with no linear pattern. 接近零时,点云是圆的,没有线性规律。
As it grows the points pull towards a line; near plus one they almost sit on it. 它变大时,点向一条直线靠拢;接近正一时,几乎就落在线上。
Near minus one they are just as tight, sloping down. 接近负一时同样紧密,只是向下倾斜。
It has no units, and always lies between minus one and plus one. 它没有单位,而且总在负一到正一之间。
And it measures how tightly the points hug a line, never how steep it is. 它衡量的是点贴合直线的紧密程度,而不是这条线有多陡。
Three warnings. 三个警告。
One: it only measures a straight-line relationship. 第一:它只衡量直线关系。
Two: it is not resistant, so one unusual point drags it a long way. 第二:它不稳健,一个异常点就能把它拉出很远。
Three, the biggest: a strong correlation never proves cause. 第三,也是最重要的:强相关绝不能证明因果。
Back to those children. 回到那些孩子身上。
Older children have bigger feet, and they also read better. 年龄大的孩子脚更大,同时也读得更好。
Age drove both. 是年龄在同时影响两者。
That third variable has a name: a lurking variable. 这第三个变量有个名字:潜在变量。
If the form is linear, we draw the line that fits best. 如果形态是线性的,我们就画出拟合得最好的那条直线。
The equation is y hat equals a plus b times x. 方程是:y 帽等于 a 加 b 乘以 x。
y hat is the predicted response, never the observed one. y 帽是预测的响应值,绝不是实际观测到的值。
The slope b is the predicted change in the response for each extra unit of x. 斜率 b 是 x 每增加一个单位,预测响应值的变化量。
The intercept a is the predicted response when x is zero, which sometimes has no sensible meaning. 截距 a 是 x 等于零时的预测响应值,有时它并没有合理的含义。
And predicting far outside your data is extrapolation — avoid it. 而在数据范围之外很远的地方做预测,就叫外推。
Now a full question. 现在做一道完整的题。
A study records revision hours and test score. 一项研究记录了复习时长和考试分数。
The least-squares line is y hat equals twenty plus three x. 最小二乘回归线是:y 帽等于二十加三 x。
Part a: interpret the slope in context. 第 a 问:结合语境解释斜率。
Part b: predict the score after five hours. 第 b 问:预测复习五小时后的分数。
Pause here and try both. 先暂停,两问都自己试一试。
The slope is three, so each extra hour of revision predicts three more points. 斜率是三,所以每多复习一小时,预测分数就提高三分。
For part b, put five in place of x: twenty plus three times five is thirty-five. 第 b 问,把五代入 x:二十加三乘以五,等于三十五。
A prediction is almost never exactly right, and the miss has a name. 预测几乎不可能完全准确,而这个偏差有个名字。
A residual is the actual value minus the predicted value: the vertical gap from the point to the line. 残差就是实际值减去预测值:从点到直线的竖直距离。
So why this line and not another? 那么,为什么是这条线,而不是别的线呢?
Watch the total in the top left. 看左上角那个总和。
Every line has a total squared residual, and least squares makes it as small as possible. 每条直线都有一个残差平方和,最小二乘线就是让这个总和尽可能小的那一条。
Back to our study. 回到我们那项研究。
The student who revised five hours actually scored forty. 那位复习了五小时的学生,实际得了四十分。
Find the residual, and say what it means. 求残差,并说明它的含义。
Pause and try it. 先暂停,自己试一试。
The predicted score was thirty-five. 预测分数是三十五。
The residual is actual minus predicted: forty minus thirty-five, which is plus five. 残差等于实际值减预测值:四十减三十五,等于正五。
A positive residual means the point sits above the line. 残差为正,说明这个点位于直线上方。
So the line under-predicted this student by five points. 所以这条直线低估了这位学生五分。
One residual is a single miss. 一个残差只是一次偏差。
Plot all of them against the explanatory variable and you get a residual plot. 把所有残差对解释变量作图,就得到残差图。
On the left is what you want: a formless cloud around zero. 左边这张是你想要的样子:围绕零的一团散乱的点。
A line fits. 说明直线是合适的。
The middle one curves, up then down, so the relationship was never linear. 中间那张先上后下地弯曲,说明这个关系从来就不是线性的。
The one on the right fans out: the misses grow, so the line predicts well at one end and badly at the other. 右边那张呈扇形张开:偏差越来越大,所以直线在一端预测得好,在另一端预测得差。
Here is why. 原因就在这里。
Four data sets, all with the same correlation, nought point eight two, and the same line. 四组数据,相关系数都是零点八二,回归线也完全相同。
Along the top, the scatterplots barely differ. 上面一排,散点图几乎看不出差别。
Underneath, the residual plots give them away. 下面一排,残差图把它们全暴露了。
The second is a clean arch: curved, so a line was never right. 第二组是一道干净的拱形:它是曲线,直线从来就不对。
The third has one point flung far from the rest. 第三组有一个点被甩得远远的。
The fourth bunches every point at one end, with a lone point dragging the line. 第四组把所有点都挤在一端, 只剩一个孤零零的点把整条直线拽了过去。
A random residual plot means a line fits; a curve means it does not. 残差图没有规律,说明直线合适;出现曲线,就说明不合适。
Then r squared, the coefficient of determination: the fraction of the variation in the response that the line explains. 接下来是 r 平方,决定系数:响应变量的变异中,被这条直线解释的比例。
Here, eighty-one percent of the variation in scores is explained by revision hours. 在这里,分数变异的百分之八十一由复习时长解释。
And s, the standard deviation of the residuals, is a typical prediction error. 而 s,也就是残差的标准差,是典型的预测误差。
Finally, the slope equals the correlation times the spread of y, over the spread of x. 最后,斜率等于相关系数乘以 y 的离散程度,再除以 x 的离散程度。
Not every point pulls equally. 并不是每个点的影响力都一样。
An influential point swings the line when you take it out. 影响点就是把它去掉后会让直线明显改变的点。
An outlier in regression sits far from the line vertically, so its residual is large. 回归中的离群点在竖直方向上远离直线,所以它的残差很大。
A high-leverage point sits at an extreme x value, far from the others, and that one has real power over the slope. 高杠杆点位于 x 的极端位置,远离其他所有点,它对斜率的影响力最大。
And if the pattern curves, transform a variable, such as the log of the response. 如果整体规律是弯曲的,就对变量做变换,例如取响应变量的对数。
Four marks that go missing every year. 每年都有四个地方在丢分。
First, answer in context, with units: not the slope is three, but each extra hour predicts three more points. 第一,结合语境作答,并写出单位: 不要只说"斜率是三",而要说"每多一小时,预测分数多三分"。
Second, correlation is not causation - name a lurking variable instead. 第二,相关不等于因果——要指出一个潜在变量。
Third, never extrapolate. 第三,绝不要外推。
Fourth, check the residual plot: no pattern, the line fits; a curve, and it does not. 第四,检查残差图:没有规律说明直线合适;出现曲线就说明不合适。

Log in or create account

IGCSE, A-Level & AP