Statistics
IGCSE Mathematics Topic 9 8:09 English narration · English + 中文 subtitles burned in
Chapters
Transcript
Every day the world collects numbers about you: how far you travel, how long you sleep, what you score in a test.
每天,世界都在收集关于你的数字:你走了多远,睡了多久,考试考了多少分。
A long list of numbers tells you almost nothing.
单看一长串数字,几乎说明不了什么。
But pour it into a shape, the way these beads drop through the board, and a pattern appears.
但把它倒进一个形状里, 就像这些小球落下来,规律就出现了。
Statistics turns a pile of statistical data into an answer.
统计学把一堆统计数据变成一个答案。
This is Topic Nine: statistics.
这是第九个专题:统计。
Six ideas, and every one does the same job — describe a set of data so somebody else can understand it.
六个概念,每一个都在做同一件事—— 把一组数据描述出来,让别人看懂。
Parts marked Extended are only for the Extended papers.
标注 Extended 的部分只在扩展卷中考查。
Here is the plan.
先看计划。
First the three averages and the range.
首先是三种平均数和极差。
Then averages from a frequency table.
然后是从频数表求平均数。
Then quartiles and box plots, which measure spread.
接着是四分位数和箱线图,它们衡量离散程度。
Then the charts: pie charts, cumulative frequency curves and histograms.
然后是各种图表:饼图、累积频数曲线和直方图。
And last, scatter diagrams and correlation.
最后是散点图与相关性。
Start with the three averages.
先看三种平均数。
Each picks a typical value differently.
它们各自用不同的方式找一个有代表性的值。
The mean adds every value up and divides by how many there are: these five total twenty, so the mean is four.
平均数是把所有数加起来再除以个数:这五个数加起来是二十,所以平均数是四。
The median is the middle value once the data is in order — here, three.
中位数是排好顺序后正中间的值,这里是三。
The mode is the most common value, also three.
众数是出现最多的值,也是三。
The range is not an average: largest minus smallest, seven take away two, five.
极差不是平均数:最大减最小,七减二,等于五。
It measures spread.
它衡量的是离散程度。
Now all four at once, on four, seven, seven, two, five.
现在四个一起求,用四、七、七、二、五。
The most important move is this: always put the data in order first.
最关键的一步是:一定要先把数据按顺序排好。
Two, four, five, seven, seven.
二、四、五、七、七。
Now everything reads straight off.
现在每个量都能直接读出来。
The mean is the total, twenty-five, divided by five, which is five.
平均数是总和二十五除以五,等于五。
The median is the middle one, also five.
中位数是中间那个,也是五。
The mode is seven, the only value appearing twice.
众数是七,只有七出现了两次。
The range is seven take away two, which is five.
极差是七减二,等于五。
Data often arrives as a frequency table instead of a list.
数据常常以频数表出现,而不是一串数字。
One appears four times, two five times, three once.
一出现四次,二出现五次,三出现一次。
To find the mean, multiply each value by its frequency: four, ten and three.
求平均数时,先把每个数值乘以它的频数:四、十和三。
Then add those products up to get seventeen, the total of all ten values.
再把这些乘积加起来得到十七,这是全部十个数据的总和。
Finally, divide by the total frequency, ten, giving one point seven.
最后除以总频数十,得到一点七。
Marks are lost here: divide by ten, not by the number of rows.
失分就在这里:除以十,而不是除以行数。
Now spread, and this is Extended only.
现在讲离散程度,这部分只属于扩展内容。
Put the data in order and cut it into four equal parts.
把数据排好顺序,切成四等份。
The three cuts are the quartiles.
三个切点就是四分位数。
The lower quartile is one quarter of the way up, here seven.
下四分位数在四分之一处,这里是七。
The median is halfway, ten.
中位数在中间,是十。
The upper quartile is fourteen.
上四分位数是十四。
The interquartile range is fourteen take away seven — the spread of the middle half.
四分位距是十四减七——也就是中间一半数据的离散程度。
Unlike the range, it ignores extreme values.
和极差不同,它不受极端值影响。
Watch those five numbers become a picture.
看这五个数怎样变成一张图。
Sort the data.
先排序。
Mark the median.
标出中位数。
Then the two quartiles, one each side.
然后是两个四分位数,中位数两边各一个。
The box runs from the lower quartile to the upper quartile, so the box is the middle half of the data, and the line inside it is the median.
箱子从下四分位数画到上四分位数,所以箱子就是中间一半的数据,里面那条线是中位数。
The whiskers reach out to the smallest and largest values.
两条须线伸向最小值和最大值。
Now the charts you must draw and read.
现在看你必须会画、会读的图。
Start organising data with a tally table — one mark per value, then count.
先用计数表整理数据——每个值画一笔,再数出来。
A pictogram uses one symbol to stand for a fixed number of items, so always check the key before you count.
象形图用一个符号代表固定数量的物品,所以数之前一定要先看图例。
A bar chart draws one bar per category, and the height is the frequency — leave gaps, because the categories are separate.
条形图为每个类别画一根柱子,高度就是频数—— 柱子之间要留空隙,因为类别是彼此独立的。
Also on your syllabus: stem-and-leaf diagrams, which always need a key, and frequency distributions.
考纲上还有:茎叶图,它一定要有键;以及频数分布表。
A pie chart shows how a whole is shared out.
饼图表示一个整体是怎样分配的。
The whole circle is three hundred and sixty degrees, so each slice is the same fraction of the circle as of the data.
整个圆是三百六十度, 所以每一块占圆的比例,就等于它在数据中所占的比例。
A hundred and twenty people were asked; thirty chose tea.
一共问了一百二十个人,其中三十人选了茶。
That is one quarter, and a quarter of three hundred and sixty is ninety degrees — the right angle you can see.
那正好是四分之一, 三百六十的四分之一是九十度,就是你看到的那个直角。
Always check your angles add up to three hundred and sixty.
最后一定要检查:所有的角加起来是三百六十度。
Extended again.
又是扩展内容。
Cumulative frequency is a running total: each entry is that class plus everything before it.
累积频数就是一个累加的总数:每一项等于这一组加上前面所有的组。
Plot the running total against the upper end of each class, and join the points with a smooth curve.
把这个总数对着每组的上端点画出来,再用光滑曲线把点连起来。
Now the curve reads values for you.
这条曲线能帮你读数。
The total is forty, so go across at twenty, then straight down: that is the median, about thirty.
总数是四十,所以在二十处横过去,再竖直向下: 那就是中位数,大约三十。
Do the same at ten and thirty for the quartiles, or at thirty-six for the ninetieth percentile.
在十和三十的位置同样操作,就得到两个四分位数; 在三十六处可读出第九十百分位数。
Histograms look like bar charts, but they are not.
直方图看上去像条形图,但并不是一回事。
Watch what happens when the classes have different widths.
看看各组宽度不同时会发生什么。
Plot frequency as the height, and the wide bar looks enormous — it wins only because it is wide, and the picture lies.
如果把频数当作高度,那根宽的柱子看起来大得吓人——它赢只是因为宽,这张图在骗你。
So a histogram plots frequency density instead: frequency divided by class width.
所以直方图改画频数密度:频数除以组距。
Now the heights are fair, and the area of each bar becomes its frequency.
现在高度公平了,而且每根柱子的面积就是它的频数。
You must calculate with it too, and the rule is short: frequency divided by the class width.
你还必须会用它来计算,法则很短:频数除以组距。
A class of width ten holds twenty-five values, so its density is twenty-five divided by ten, two point five.
一个组距为十的组里有二十五个数据,所以它的密度是二十五除以十,等于二点五。
Now backwards, because the exam loves this: to read a frequency back off a histogram, multiply height by width — the area of the bar.
现在反过来,因为考试很喜欢这个方向:要从直方图读回频数, 就用高度乘以宽度——那就是柱子的面积。
And a free mark: label the vertical axis frequency density.
还有一分白送:纵轴要写频数密度。
A scatter diagram plots pairs of values as points, to show whether two things are linked.
散点图把成对的数值画成点,用来显示两件事之间有没有联系。
That link is correlation, and there are three words for it.
这种联系叫相关性,有三个说法。
Positive: as one goes up, the other goes up.
正相关:一个变大,另一个也变大。
Negative: as one goes up, the other goes down.
负相关:一个变大,另一个变小。
Zero: no clear link, just a cloud.
零相关:看不出联系,只是一团点。
Mark points with small crosses, and describe the correlation in words.
画点要用小叉号,并且要用文字描述相关性。
If there is correlation, draw a line of best fit.
如果存在相关性,就画一条最佳拟合线。
It must be a single straight ruled line, drawn by eye, running across the whole data, with about as many points above as below.
它必须是用直尺画的一条直线,凭眼睛判断, 贯穿整组数据,线上方和下方的点数大致相同。
Here revision time against test score climbs steadily, so the correlation is positive.
这里复习时间对考试分数稳步上升,所以是正相关。
Use it to predict: up from five hours, across, and read about fifty-five marks.
用它来预测:从五小时往上,再横过去,读出大约五十五分。
Only predict inside your data.
只在你自己的数据范围内预测。
One full Extended question to finish.
最后来做一整道扩展题。
Twenty students recorded their homework minutes.
二十名学生记录了自己做作业的分钟数。
Grouped data hides the exact values, so we can only estimate the mean.
分组数据把确切数值藏了起来,所以我们只能估算平均数。
Pause here and try it.
先暂停,自己试一试。
Step one: take the midpoint of each group — five, fifteen, twenty-five, thirty-five.
第一步:取每一组的组中值——五、十五、二十五、三十五。
Step two: multiply each midpoint by its frequency — twenty, a hundred and thirty-five, a hundred and twenty-five, seventy.
第二步:把每个组中值乘以它的频数——二十、一百三十五、一百二十五、七十。
Step three: add up both columns — twenty students, three hundred and fifty in total.
第三步:把两列都加起来——二十名学生,总共三百五十。
So the estimated mean is three hundred and fifty divided by twenty, seventeen point five minutes.
所以估算的平均数是三百五十除以二十,等于十七点五分钟。
And the modal class is simply the group with the highest frequency: ten to twenty.
而众数组就是频数最高的那一组:十到二十分钟。
Three marks students give away every year.
三个学生年年都在送掉的分。
First, before any median or quartile, put the data in order — a median read off an unsorted list is simply wrong.
第一,在求中位数或四分位数之前,先把数据排好顺序—— 在没排序的数字上读中位数一定是错的。
Second, from a frequency table, divide by the total frequency, not by how many rows there are.
第二,从频数表求平均数时,要除以总频数,而不是除以有几行。
Third, on a histogram with unequal classes the height is the frequency density, and the area of the bar is the frequency.
第三,在组距不等的直方图上,高度是频数密度,而柱子的面积才是频数。
Four quick questions before you go.
走之前四个小问题。
Which average is the middle value?
哪种平均数是正中间的值?
The median.
中位数。
What is the range of two, four, five, seven, seven?
二、四、五、七、七的极差是多少?
Seven take away two, five.
七减二,等于五。
What is the interquartile range?
四分位距是什么?
Upper quartile take away lower quartile.
上四分位数减下四分位数。
And what does a line of best fit help you do?
最佳拟合线能帮你做什么?
Predict.
预测。