Collecting Data
AP Statistics Topic 3 7:31 English narration · English + 中文 subtitles burned in
Chapters
Transcript
An online poll collects thirty thousand replies.
一个网络投票收到三万份回复。
A careful survey asks only one thousand people.
一次认真的调查只问了一千个人。
Which is closer to the truth?
哪一个更接近真相?
The small one, almost every time.
小的那个,几乎每次都是。
The poll let people choose themselves, so it heard only from those who cared most.
投票让人们自己决定要不要参加,所以它只听到了最在意的那群人。
The survey used chance to choose, so it heard from everyone.
调查用随机来选人,所以它听到了所有人。
How you collect beats how much you collect.
怎么收集,比收集了多少更重要。
Unit three asks where the numbers come from.
第三单元问的是:这些数字从哪里来。
Get this right, and everything you calculate later is worth trusting.
把这一步做对,你后面算出来的一切才值得相信。
Let's begin.
让我们开始吧。
Start with the words.
先把词说清楚。
The population is the whole group you want to know about.
总体是你想了解的那一整群对象。
A sample is the smaller part you actually measure.
样本是你真正去测量的那一小部分。
Measuring everyone is called a census, and it is usually impossible.
把每一个成员都测一遍叫做普查,而普查通常做不到。
So the sample must stand in for the population, and one thing stops it: bias.
所以样本必须代表总体,而有一样东西会挡住它:偏差。
Bias means a consistent error: the method misses the truth in the same direction, over and over.
偏差是指一种一致的误差:这个方法一次又一次地朝同一个方向偏离真相。
Two kinds of study, and they let you say very different things.
研究有两种,而它们允许你说的话很不一样。
In an observational study you watch and measure, but never interfere.
在观察性研究里,你只观察和测量,从不干预。
You may look back at records, or follow people forward in time.
你可以回头去查旧记录,也可以跟踪一群人往后一段时间。 它显示的是关联:两件事一起变化。
It shows association: two things move together. In an experiment you do interfere.
在实验里,你要干预:你给不同的组不同的处理,然后比较。
You give groups different treatments, then compare. Only an experiment can show causation.
只有实验才能显示因果关系,也就是 causation。
So why can watching never prove cause?
那么,为什么只是观察永远证明不了因果?
Look at a real pattern: on days when people buy more ice cream, more people drown.
看一个真实的规律: 在人们买冰淇淋更多的日子里,溺水的人也更多。
Does ice cream cause drowning?
是冰淇淋导致溺水吗?
Of course not.
当然不是。
Hot weather does both. It sends people out for ice cream, and it sends people into the water.
是炎热的天气同时造成了两者:它让人出门去买冰淇淋,也让人下水。
Hot weather is a confounding variable: tied to the explanatory variable, changing the response, so the two effects can never be separated.
炎热的天气就是一个混杂变量:它与解释变量捆在一起,又改变了响应变量, 所以这两种作用永远分不开。
Take a sample by random sampling, and let chance choose.
抽样要用随机抽样,让机会来挑人。
Four designs come up again and again.
有四种设计反复出现。
In a simple random sample, or SRS, every group of the chosen size is equally likely.
简单随机样本,也叫 SRS:任何一组给定大小的个体,被抽中的机会都一样。
In a stratified sample you split the population into strata, groups that are alike inside, and sample randomly within each one.
分层样本:把总体分成若干层,每一层内部彼此相似,然后在每一层里各随机抽一小部分。
In a cluster sample you pick whole clusters at random, and measure everyone inside.
整群样本:随机抽中整群,再测量群里的每一个人。
In a systematic sample you take a random start, then every fifth individual.
系统样本:随机选一个起点,然后每隔五个取一个。
Stratified and cluster sound alike, and students mix them up every year.
分层和整群听起来很像,学生每年都会弄混。
The difference is what sits inside each group.
区别在于每个组里面装的是什么。
A stratum is alike inside: all the tenth graders together, all the eleventh graders together, and you take a few from every stratum.
层的内部是相似的:十年级的放在一起,十一年级的放在一起,你从每一层里各取几个。
A cluster is a small copy of the whole population, like one whole class, and you measure everyone inside it.
群则是整个总体的一个小复制品,比如一个完整的班, 而选中的群,你要把里面的人全部测量。
So: some from every group, or everyone from some groups.
所以:要么每个组各取几个,要么某几个组全取。
The exam often asks you to describe how you would actually take a simple random sample.
考试常常让你说明:你会怎样真正抽出一个简单随机样本。
Three steps.
三步。
One: give every member a different number, from one up to the size of the population.
第一,给每一个成员编一个不同的号码,从一号一直编到总体的人数。
Two: use a random number generator to produce numbers in that range.
第二,用随机数生成器产生这个范围内的号码。
Three: ignore any repeat, and keep drawing until you have enough different people.
第三,重复出现的号码直接跳过,一直抽到凑够足够多的不同的人为止。
Write all three steps.
三步都要写出来。
Even a careful plan can go wrong, and each failure has a name.
就算计划得很小心,也可能出问题,而每一种问题都有名字。
Undercoverage: part of the population had little chance of being chosen.
覆盖不足:一部分总体几乎没有机会被抽到。
A convenience sample or voluntary response: people chose themselves, so only strong feelings reply.
方便样本或自愿回应:人们自己决定要不要参加,于是只有情绪强烈的人回答。
Nonresponse: the people you picked never answered, and they may differ from those who did.
无回应:你抽中的人从来没有回复,而他们可能和回复的人不一样。
Response bias: the answers are wrong, because a question was confusing or too personal.
回应偏差:答案本身就是错的,因为问题让人困惑,或者太私人。
A bigger sample does not fix bias. It only makes a bigger biased sample.
样本变大并不能消除偏差,只会得到一个更大的有偏样本。
Now look inside an experiment.
现在看看实验的内部。
The sixty people being studied are the experimental units, or subjects when they are people.
被研究的这六十个人叫实验单元;如果是人,也叫受试者。
The variable the researcher controls is the explanatory variable, also called a factor, and its levels are the treatments: here, the drug or a placebo.
研究者所控制的那个变量叫解释变量,也叫因子,它的各个水平就是处理:这里是药或安慰剂。
The outcome measured afterwards is the response variable.
之后测量到的结果叫响应变量。
And the middle step is the heart of it: chance, not the researcher, decides who goes where.
而中间那一步才是关键: 由随机、而不是研究者,来决定谁进哪一组。
A completely randomized design needs three things, and the exam wants all three.
完全随机设计需要三样东西,而考试三样都要。
Comparison and control: at least two groups, so there is something to compare against, often a control group with no active treatment.
对比与控制:至少要有两个组,才有可以比较的对象,其中常常有一个不接受有效处理的对照组。
Randomization: chance assigns units to treatments, balancing every other variable, even the ones nobody thought of.
随机化:由随机把单元分配到各个处理,从而把其他每一个变量都平衡掉,包括谁也没想到的那些。
Replication: enough units per group, so a real difference shows through the natural variation.
重复:每一组都要有足够多的单元,真实的差异才能从自然波动中显现出来。
Sometimes you already know that one variable matters.
有时候你已经知道某个变量会有影响。
Suppose men and women respond differently to the drug.
比如男性和女性对这种药的反应不同。
Then do not let chance mix them.
那就别让随机把他们混在一起。
Blocking first splits the subjects into blocks, one of men and one of women, and randomizes inside each block.
区组设计先把受试者分成区组:男性一个区组,女性一个区组, 然后在每个区组内部做随机分配。
That is a randomized block design, and it removes the variation caused by that known variable.
这就是随机区组设计, 它去掉了那个已知变量带来的波动。
A matched pairs design is the smallest version: pair up similar subjects and give both treatments in random order.
配对设计是它最小的版本:把相似的受试者两两配对,再按随机顺序给出两种处理。
People often improve simply because they believe they are being treated.
人们常常只因为相信自己得到了治疗,情况就好转了。
That is the placebo effect, and it is why the control group is usually given a placebo, a dummy treatment with no active ingredient.
这就是安慰剂效应, 也正因为如此,对照组通常会拿到安慰剂:一种不含有效成分的假处理。
To stop belief from bending the result, use blinding: hide who gets what.
为了不让信念扭曲结果,要用盲法:把谁拿到什么藏起来。
In a single-blind experiment one side does not know: usually the subjects, or the people measuring the response.
单盲实验里有一方不知道:通常是受试者,或者是测量结果的那些人。
In a double-blind experiment neither the subjects nor the researchers know.
双盲实验里,受试者和研究者都不知道。
Now the question every exam asks. What may you conclude?
现在是每次考试都会问的问题:你可以下什么结论?
Two separate uses of randomness answer it.
有两种彼此独立的随机性来回答它。
Random assignment, where chance decides who gets which treatment, lets you say the treatment caused the difference.
随机分配,也就是由随机决定谁接受哪种处理, 让你可以说:是这个处理造成了差异。
Random sampling, where chance decides who is in the study at all, lets you generalize to the whole population.
随机抽样,也就是由随机决定谁进入这项研究, 让你可以把结论推广到整个总体。
They are independent, and the four cases give four different conclusions.
这两者互不相干,四种情况给出四种不同的结论。
Let's finish one.
我们来完整做一道。
Researchers ask for volunteers.
研究者招募志愿者。
One hundred people sign up, and chance decides which fifty get the new drug and which fifty get a placebo.
一百人报名, 由随机决定其中哪五十人拿到新药、哪五十人拿到安慰剂。
The drug group improves much more.
用药组的改善明显更大。
What may the researchers conclude?
研究者可以下什么结论?
Ask the two questions.
先问那两个问题。
Was the assignment random?
分配是随机的吗?
Yes, so the improvement can be attributed to the drug. That is causation.
是, 所以这个改善可以归因于这种药,这就是因果关系。
Was the sample random?
样本是随机抽取的吗?
No — they were not randomly sampled; they volunteered.
不是——他们不是随机抽样来的,而是自愿报名的。
So the conclusion covers only these one hundred volunteers.
所以结论只适用于这一百名志愿者。
Three marks students throw away.
学生最常白白丢掉的三分。
First, when a question says describe the method, describe it: number the population, use a random number generator, and say what you do with repeats.
第一,题目说"说明这个方法"时,就要真的说明: 给总体编号、使用随机数生成器、并说清重复的号码怎么处理。
A vague answer earns nothing.
含糊的答案一分也拿不到。
Second, name the bias and say which way it pushes the estimate.
第二,说出偏差的名字,并说清它把估计推高还是推低。
Third, never write the word causation unless the treatments were randomly assigned.
第三,除非处理是随机分配的,否则绝不要写因果关系或 causation。