Should I Worry About Error? · Devo Me Preocupar com o Erro?
Twenty apples leave a population question
- A fictional random sample of twenty apples has mean mass 150 g and sample standard deviation 12 g. The target is the mean mass of apples in the sampled population.
- Another random sample can have a different mean. The sample mean can happen to equal the population mean, but equality is not something this sample establishes.
A sample mean can never happen to equal the population mean.
Equality is possible; it is not guaranteed or established simply by observing the sample mean.
Separate variability from a data error
- Sampling variability is natural variation between valid samples. A mistyped mass or biased sampling method is a different problem.
- An uncertainty formula does not repair a wrong measurement or a sample collected only from the largest apples.
Inference about a mean uses the t-distribution (not z) mainly because...
Estimating σ with s adds uncertainty → use t.
Match the issue to its interpretation.
Inference models sampling variability; it does not repair corrupted or selectively collected data.
Estimate a range or assess a claim
- A confidence interval estimates plausible values for the population mean under its assumptions. A significance test assesses a specified null claim.
- Choose the population and question before the calculation. Random assignment can support a causal comparison without making participants representative of every population.
Sampling variability is a mistake in data collection that must be fixed.
Variability is natural and quantifiable — not a blunder.
Why estimating sigma changes the reference curve
- In these t procedures, the population standard deviation $\sigma$ is unknown. The sample standard deviation $s$ estimates it.
- The standardized statistic $(\bar x-\mu)/(s/\sqrt n)$ has a Student t distribution with $n-1$ degrees of freedom for independent observations from a normal population.
At df = 19 the t curve is already close to normal near its centre, but their tail areas still differ. At df = 3 the heavier tails are more visible.
Compared with the normal, the t-distribution has...
Heavier tails reflect the extra uncertainty from estimating σ.
For n = 16 and sample SD 12 g, calculate the estimated standard error of the mean in grams.
12/sqrt(16) = 3 g.
Read the heavier tails correctly
- With twenty observations, the reference distribution has 19 degrees of freedom. It has more probability in its tails than the standard normal distribution.
- As degrees of freedom increase, the t curve approaches the standard normal curve. A normal curve labelled t does not show this difference.
The reference curve, degrees of freedom and tail rule must match the selected procedure.
Inference for means uses the same two tools as inference for proportions: intervals and tests.
Both quantify uncertainty via intervals and tests.
Use the model with its conditions
- For non-normal populations, t methods can be useful approximations when the sample is sufficiently large and not dominated by extreme observations.
- Check random design, independence and the data shape. A t reference curve accounts for estimating variability; it cannot compensate for every design or distribution problem.
For non-normal populations, t methods can be useful approximations when the sample is sufficiently large and not dominated by extreme observations.
When σ is unknown, mean inference uses the ___-distribution (one letter).
The t-distribution handles unknown σ.