Sampling Distribution of p-hat · 样本比例的抽样分布
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| sample proportion/ˈsæmpl prəˈpɔːʃn/ | 样本比例 | yàng běn bǐ lì |
| large counts condition/lɑːdʒ kaʊnts kənˈdɪʃn/ | 大计数条件 | dà jì shù tiáo jiàn |
| 10% condition/ten pəˈsent kənˈdɪʃn/ | 10%条件 | 10% tiáo jiàn |
The distribution of p-hat
- A sample proportion 样本比例 $\hat{p}$ estimates the true population proportion $p$.
- Its sampling distribution is centered at $p$: $\mu_{\hat{p}} = p$ (so $\hat{p}$ is unbiased).
- Different samples give different $\hat{p}$ — the spread is what we quantify next.
- Under the right conditions, this distribution is approximately normal.
p-hat 的分布
- 样本比例 $\hat{p}$ 估计真实的总体比例 $p$。
- 它的抽样分布以 $p$ 为中心:$\mu_{\hat{p}} = p$(所以 $\hat{p}$ 是无偏的)。
- 不同样本给出不同的 $\hat{p}$——我们接下来量化的就是这种分散。
- 在合适的条件下,这个分布近似正态。
Its standard deviation
- The standard deviation of $\hat{p}$ is:
-
$$\sigma_{\hat{p}} = \sqrt{\frac{p(1-p)}{n}}$$
- Bigger $n$ → smaller spread → a more precise estimate (the $\sqrt{n}$ in the bottom).
- This measures the typical distance of $\hat{p}$ from the true $p$.
它的标准差
- $\hat{p}$ 的标准差是:
-
$$\sigma_{\hat{p}} = \sqrt{\frac{p(1-p)}{n}}$$
- $n$ 越大 → 分散越小 → 估计越精确(分母里的 $\sqrt{n}$)。
- 它衡量 $\hat{p}$ 离真实 $p$ 的典型距离。
The large counts condition
- To use a normal model for $\hat{p}$, check the large counts condition 大计数条件:
-
$$np \ge 10 \quad\text{and}\quad n(1-p) \ge 10$$
- In words: expect at least $10$ successes and at least $10$ failures.
- If either count is under $10$, the sampling distribution is too skewed for the normal model.
大计数条件
- 要对 $\hat{p}$ 使用正态模型,检查大计数条件:
-
$$np \ge 10 \quad\text{and}\quad n(1-p) \ge 10$$
- 用话说:期望至少 $10$ 次成功且至少 $10$ 次失败。
- 如果任一计数小于 $10$,抽样分布就太偏斜,不适合正态模型。
The 10% condition
- When sampling without replacement, check the 10% condition 10%条件.
- The sample must be no more than $10\%$ of the population: $n \le 0.10 N$.
- This keeps the trials "close enough" to independent for the SD formula to hold.
- A small slice of a big population barely changes as you draw.
10% 条件
- 当不放回抽样时,检查 10% 条件。
- 样本必须不超过总体的 $10\%$:$n \le 0.10 N$。
- 这让各次抽取“足够接近”独立,从而使标准差公式成立。
- 大总体的一小片,在抽取时几乎不变。
Two different conditions do two different jobs. The large counts condition ($np\ge10$, $n(1-p)\ge10$) justifies the normal shape; the 10% condition ($n\le0.10N$) justifies the independence behind the SD formula. Check both before modeling $\hat{p}$ as normal — skipping either invalidates the later inference.
两个不同的条件做两件不同的事。****大计数条件($np\ge10$、$n(1-p)\ge10$)为正态形状提供依据;10% 条件($n\le0.10N$)为标准差公式背后的独立性提供依据。把 $\hat{p}$ 建模为正态之前,两者都要检查——漏掉任一个都会使后面的推断失效。
$p = 0.4$ of a population supports a plan; you sample $n = 100$.
- Center: $\mu_{\hat{p}} = 0.4$. SD: $\sqrt{\dfrac{0.4 \times 0.6}{100}} = \sqrt{0.0024} \approx 0.049$.
- Large counts: $np = 40 \ge 10$ and $n(1-p) = 60 \ge 10$. ✓ Normal is OK.
- 10%: fine as long as the population exceeds $1000$.
一个总体中 $p = 0.4$ 支持某方案;你抽 $n = 100$。
- 中心:$\mu_{\hat{p}} = 0.4$。标准差:$\sqrt{\dfrac{0.4 \times 0.6}{100}} = \sqrt{0.0024} \approx 0.049$。
- 大计数:$np = 40 \ge 10$ 且 $n(1-p) = 60 \ge 10$。✓ 可用正态。
- **10%:**只要总体超过 $1000$ 就没问题。
The sample proportion $\hat{p}$ has $\mu_{\hat{p}} = p$ and $\sigma_{\hat{p}} = \sqrt{\frac{p(1-p)}{n}}$. Model it as normal when the large counts condition ($np\ge10$, $n(1-p)\ge10$) holds; when sampling without replacement, also check the 10% condition ($n \le 0.10N$).
样本比例 $\hat{p}$ 有 $\mu_{\hat{p}} = p$ 和 $\sigma_{\hat{p}} = \sqrt{\frac{p(1-p)}{n}}$。当大计数条件($np\ge10$、$n(1-p)\ge10$)成立时把它建模为正态;不放回抽样时还要检查 10% 条件($n \le 0.10N$)。
The sampling distribution of p-hat · p-hat 的抽样分布
Centered at p, spread √(p(1−p)/n); normal under the large counts condition. · 以 p 为中心,分散 √(p(1−p)/n);在大计数条件下正态。
p = 0.4, n = 100. Find σ of p-hat = √(p(1−p)/n). Round to two decimals. · p = 0.4,n = 100。求 p-hat 的 σ = √(p(1−p)/n)。保留两位小数。
√(0.4·0.6/100) = √0.0024 ≈ 0.049 ≈ 0.05. · √(0.4·0.6/100) = √0.0024 ≈ 0.049 ≈ 0.05。
The large counts condition for a normal model of p-hat requires... · p-hat 正态模型的大计数条件要求……
At least 10 expected successes AND 10 expected failures. · 至少 10 次期望成功且 10 次期望失败。
The mean of the sampling distribution of p-hat equals the true proportion p. · p-hat 抽样分布的均值等于真实比例 p。
μ(p-hat) = p, so p-hat is unbiased. · μ(p-hat) = p,所以 p-hat 是无偏的。
When sampling without replacement, the 10% condition requires... · 不放回抽样时,10% 条件要求……
The sample must be at most 10% of the population. · 样本必须至多为总体的 10%。
Increasing n makes the standard deviation of p-hat smaller. · 增大 n 会使 p-hat 的标准差变小。
n is under the square root, so larger n → smaller spread. · n 在平方根之下,所以 n 越大 → 分散越小。