Data displays, normal distributions and association
| English | 中文 | Pinyin |
|---|---|---|
| residual/rɪˈsɪdʒuːəl/ | 残差 | cán chà |
| normal distribution/ˈnɔːml ˌdɪstrɪˈbjuːʃn/ | 正态分布 | zhèng tài fēn bù |
A decision before an answer
- A histogram shows frequencies in value intervals. A bar chart of categories answers a different question even if both use rectangles.
- Your goal: Read centre, spread and shape from labelled data displays.
Read the relationship
- The mean weights all observations; the median uses their ordered middle. A histogram groups numeric values into intervals, a box plot displays quartiles and a bar chart compares categories. Inspect units, scale and frequencies before comparing apparent width or height. Use the task’s stated quartile convention where calculation is required.
- Interpret a normal model using mean and standard deviation.
Add 7 to every observation. Standard deviation:
Distances from the shifted mean stay unchanged.
Use the defining rule
- Adding a constant shifts mean and median but leaves standard deviation unchanged; multiplying values by a positive factor scales both centre and spread. For a normal model, the mean is at the symmetric centre; approximately 68% lie within one standard deviation and 95% within two. These approximations concern a stated normal model, not every vaguely symmetric dataset.
- Distinguish two-variable association from causal evidence.
A normal model with mean 100 and SD 10 has about 68% in:
One standard deviation on both sides gives 90–110.
Check the conditions
- A scatterplot pairs quantitative measurements. A fitted line predicts a response; a residual is observed minus predicted. A two-way table instead compares categories and conditional shares. A trend can show association without proving cause, because other variables or selection may explain the relationship.
- Distinguish two-variable association from causal evidence.
A normal-model score distribution with mean 50 and standard deviation 5 places about 68% between 45 and 55 and about 95% between 40 and 60. Adding 10 shifts the mean to 60 but keeps standard deviation 5. For y=2x+4 at x=3, predicted y=10; an observed value 12 gives residual +2. These are hypothetical teaching models.
Observed 18, predicted 15: residual ____.
18-15=3.
Apply the task format
- An informal fitted line should follow the main trend, not join every point. Describe direction, shape and unusual points. Interpolation inside the observed range is better grounded than distant extrapolation. Changes in graphical scale can make an unchanged relationship look stronger or weaker, so read the actual coordinates and axes.
- Distinguish two-variable association from causal evidence.
Do not use normal percentages without a normal-model condition or interpret a scatterplot slope as an automatic causal effect.
Which answer fits this case?
Read centre, spread and shape from labelled data displays
Association alone proves causation.
A third variable or selection can explain association.
Keep the distinctions
- normal distribution 正态分布 — A symmetric bell-shaped model determined by a mean and standard deviation.
- residual 残差 — Observed response minus the model’s predicted response.
- Read centre, spread and shape from labelled data displays.
- Interpret a normal model using mean and standard deviation.
- Distinguish two-variable association from causal evidence.
Match each term with its precise meaning in this lesson.
Keep the distinctions stated in the teaching example.
Put this lesson’s reasoning or event sequence in order.
The order follows the stated process; check each stage before the next.