Sampling and a large data set
| English | 中文 | Pinyin |
|---|---|---|
| sampling frame/ˈsæmplɪŋ freɪm/ | 抽样框 | chōu yàng kuāng |
Who is missing from the data?
- A weather database has a missing entry and several stations. Treating each row as identical can distort a comparison.
- This lesson studies sampling frame 抽样框: The list or population definition from which a sample is selected.
Choose the mathematical structure
- Identify the population, sampling unit and frame. Distinguish random, systematic, stratified, quota and opportunity sampling. Missing data is not zero; verify units, dates and variable definitions before comparing samples.
- State the allowed inputs and units before calculating. An equation should express the relationship, not just record a calculator entry.
Which description correctly defines sampling frame?
The list or population definition from which a sample is selected.
Work through a checked case
- Check the result against the starting quantities. Substitute into the original relation, or compare the graph and numerical answer where appropriate.
A population has 120 students in one group and 80 in another. A proportional stratified sample of 30 needs 30×120/200=18 from the first and 12 from the second. Random selection is then needed within each group.
Sampling and a large data set
Identify the population, sampling unit and frame
Allocate proportionally, then explain why random selection within each group still matters.
Find the first-group allocation when sample size is 30 and group sizes are 120 and 80.
Allocate 30×120/200=18, then select randomly within that group.
Test a tempting shortcut
- A large biased sample remains biased. Stratification is not the same as selecting whoever is available from each group. AQA large-data-set familiarity requires the actual supplied data and metadata, not invented weather values.
- When a shortcut fails, identify the assumption it breaks. Keep an exact value until the requested final rounding.
Doubling a biased sample automatically removes its selection bias. This claim is false. Explain which definition or assumption it violates.
Find the second-group allocation for that sample.
Allocate 30×80/200=12. Check that 18+12=30.
Doubling a biased sample automatically removes its selection bias.
A large biased sample remains biased. Stratification is not the same as selecting whoever is available from each group. AQA large-data-set familiarity requires the actual supplied data and metadata, not invented weather values.
Interpret a new situation
- For AQA, use the official large data set in a supervised spreadsheet task: identify a variable, justify a comparison, inspect missing values, create a display and explain a limitation. Save the decisions with the analysis.
- A complete solution gives the mathematical result and explains what it means. Check that it is possible in the stated context.
A sample includes 45 of 300 people. Find its percentage.
Percentage=45/300×100=15%.
Match each part of a complete solution to its purpose.
An assumption justifies the model; a check tests the result; interpretation connects it to the question.
Use this in your course
- 7357 · A-level · K. Match the target tier and specification before assigning extensions.
- Give the method before the final answer, and use the paper's calculator and formula rules. Review a wrong answer by locating the first invalid step.
The list or population definition from which a sample is selected. Choose the relationship, show the method, check its assumptions and interpret the result.