Sampling and a large data set · 抽样与大规模数据集
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| sampling frame/ˈsæmplɪŋ freɪm/ | 抽样框 | chōu yàng kuāng |
Who is missing from the data?
- A weather database has a missing entry and several stations. Treating each row as identical can distort a comparison.
- This lesson studies sampling frame 抽样框: The list or population definition from which a sample is selected.
Choose the mathematical structure
- Identify the population, sampling unit and frame. Distinguish random, systematic, stratified, quota and opportunity sampling. Missing data is not zero; verify units, dates and variable definitions before comparing samples.
- State the allowed inputs and units before calculating. An equation should express the relationship, not just record a calculator entry.
Which description correctly defines sampling frame? · 以下哪项描述正确定义了抽样框?
The list or population definition from which a sample is selected. · 从中选取样本的列表或总体定义。
Work through a checked case
- Check the result against the starting quantities. Substitute into the original relation, or compare the graph and numerical answer where appropriate.
A population has 120 students in one group and 80 in another. A proportional stratified sample of 30 needs 30×120/200=18 from the first and 12 from the second. Random selection is then needed within each group.
Sampling and a large data set · 抽样与大规模数据集
Identify the population, sampling unit and frame · 识别总体、抽样单元和抽样框
Allocate proportionally, then explain why random selection within each group still matters. · 按比例分配,然后解释为什么在每个组内随机选择仍然很重要。
Find the first-group allocation when sample size is 30 and group sizes are 120 and 80. · 当样本量为30、各组规模分别为120和80时,求第一组的分配数量。
Allocate 30×120/200=18, then select randomly within that group. · 计算分配量 30×120/200=18,然后在该组内随机抽取。
Test a tempting shortcut
- A large biased sample remains biased. Stratification is not the same as selecting whoever is available from each group. AQA large-data-set familiarity requires the actual supplied data and metadata, not invented weather values.
- When a shortcut fails, identify the assumption it breaks. Keep an exact value until the requested final rounding.
Doubling a biased sample automatically removes its selection bias. This claim is false. Explain which definition or assumption it violates.
Find the second-group allocation for that sample. · 求该样本在第二组的分配数量。
Allocate 30×80/200=12. Check that 18+12=30. · 计算分配量 30×80/200=12。验证 18+12=30。
Doubling a biased sample automatically removes its selection bias. · 将存在偏差的样本加倍并不能自动消除其选择偏差。
A large biased sample remains biased. Stratification is not the same as selecting whoever is available from each group. AQA large-data-set familiarity requires the actual supplied data and metadata, not invented weather values. · 一个存在偏差的大样本仍然是有偏差的。分层抽样并不等同于从各组中随意选取可用人员。AQA对大规模数据集的熟悉度要求使用实际提供的数据和元数据,而非虚构的天气数值。
Interpret a new situation
- For AQA, use the official large data set in a supervised spreadsheet task: identify a variable, justify a comparison, inspect missing values, create a display and explain a limitation. Save the decisions with the analysis.
- A complete solution gives the mathematical result and explains what it means. Check that it is possible in the stated context.
A sample includes 45 of 300 people. Find its percentage. · 一个样本包含45人,总人数为300人。求其百分比。
Percentage=45/300×100=15%. · 百分比=45/300×100=15%。
Match each part of a complete solution to its purpose. · 将完整解答的每个部分与其目的相匹配。
An assumption justifies the model; a check tests the result; interpretation connects it to the question. · 假设用于论证模型的合理性;检查用于验证结果;解释用于将其与问题建立联系。
Use this in your course
- edexcel IAL further mathematics; official unit S2. Other-unit enrichment is identified in the scope review; it is not an extra cash-in requirement.
- Give the method before the final answer, and use the paper's calculator and formula rules. Review a wrong answer by locating the first invalid step.
The list or population definition from which a sample is selected. Choose the relationship, show the method, check its assumptions and interpret the result.