Interval for a Difference of Proportions · Intervalo para Diferença de Proporções
| English | Português |
|---|---|
| two-sample z-interval/tuː ˈsæmpl zed ˈɪntəvl/ | intervalo z de duas amostras |
Compare support in two library populations
- Independent random samples find 60 supporters among 100 users in library 1 and 50 among 100 in library 2. Estimate $p_1-p_2$.
- The point estimate is $\hat p_1-\hat p_2=0.60-0.50=0.10$, or ten percentage points. It is not a 10% relative increase.
Check both groups and their independence
- Check the random design and independence between groups. For sampling without replacement, check $n_1\le0.10N_1$ and · e $n_2\le0.10N_2$.
- The observed counts are 60, 40, 50 and 50; all are at least ten. Surveys of the same people before and after are paired, not independent samples.
p̂1 = 0.60, p̂2 = 0.50. Find the point estimate p̂1 − p̂2.
0.60 − 0.50 = 0.10.
The two-proportion interval estimates...
It's a range of plausible values for the difference.
Combine the two variance estimates
- The two-sample z-interval 双样本z区间 uses · usa $SE=\sqrt{\hat p_1(1-\hat p_1)/n_1+\hat p_2(1-\hat p_2)/n_2}$.
- Here $SE=\sqrt{0.60(0.40)/100+0.50(0.50)/100}=0.07$. The variance terms add; the standard errors do not simply add.
The large counts condition for a difference of proportions requires...
Both groups' successes and failures must each be ≥ 10.
Match each design to the inference issue.
A two-independent-proportion procedure must not treat paired responses as independent samples.
Build the difference interval
- Use $(\hat p_1-\hat p_2)\pm z^*SE$. At 95%, this is $0.10\pm1.96(0.07)$, giving $(-0.0372,0.2372)$.
- The interval uses separate sample proportions, because it does not assume $p_1=p_2$. Do not replace them with a pooled estimate.
The difference is 10 percentage points, with an interval from about -3.72 to 23.72 percentage points. Negative and positive values are both included.
With both n = 100: SE = √(0.0024 + 0.0025). Round to two decimals.
√0.0049 = 0.07.
For a two-proportion confidence interval, each sample uses its own p̂ in the standard error.
No pooling for the interval — that's only for the test.
For 60/100 and 50/100, which belong to this difference interval?
The interval uses sqrt(0.60×0.40/100 + 0.50×0.50/100) = 0.07, then 1.96×0.07 = 0.1372.
Keep the group order throughout
- We are 95% confident that population support in library 1 minus library 2 lies between -3.72 and 23.72 percentage points.
- Reversing the subtraction reverses the endpoints: the interval for $p_2-p_1$ is $(-0.2372,0.0372)$.
Before-and-after responses from the same people are paired. Do not apply an independent-samples formula to them.
If the interval for p1 - p2 is (-0.0372, 0.2372), what is the lower endpoint for p2 - p1?
Negate both endpoints and reverse their order: (-0.2372, 0.0372).
Explain the remaining uncertainty
- The positive estimate favours library 1, but this interval also includes zero and negative differences. It does not establish which population has higher support.
- Both samples contribute uncertainty. A large first sample does not excuse a second sample with too few successes or failures.
The positive estimate favours library 1, but this interval also includes zero and negative differences. It does not establish which population has higher support.