Do Those Points Align? · 那些点排成一线吗?
| English | 中文 | Pinyin · 拼音 |
|---|---|---|
| least-squares regression/liːst skweəz rɪˈɡreʃn/ | 最小二乘回归 | zuì xiǎo èr chéng huí guī |
The slope wobbles too
- In Unit 2 we fit a least-squares line and read its slope $b$.
- But $b$ comes from a sample — a different sample gives a slightly different slope.
- So the slope varies from sample to sample, just like $\bar{x}$ or $\hat{p}$.
- Unit 9 does inference about the true slope $\beta$ behind the data.
斜率也会摆动
- 在第 2 单元我们拟合了一条最小二乘线,并读出它的斜率 $b$。
- 但 $b$ 来自一个样本——不同的样本给出略微不同的斜率。
- 所以斜率在样本之间变动,就像 $\bar{x}$ 或 $\hat{p}$ 一样。
- 第 9 单元对数据背后的真实斜率 $\beta$ 做推断。
The sampling distribution of b
- Collect the slope $b$ from every possible sample → its sampling distribution.
- It's centered at the true population slope $\beta$ (so $b$ is unbiased for $\beta$).
- Its spread is the standard error of the slope, $SE_b$.
- Under the right conditions, this distribution is modeled by a $t$-distribution.
b 的抽样分布
- 从每一个可能的样本收集斜率 $b$ → 它的抽样分布。
- 它以真实总体斜率 $\beta$ 为中心(所以 $b$ 是 $\beta$ 的无偏估计)。
- 它的分散是斜率的标准误 $SE_b$。
- 在合适的条件下,这个分布由一个 $t$ 分布建模。
Inference about the relationship
- Inference for slopes asks: is there really a linear relationship, and how strong?
- The true slope $\beta$ is unknown; $b$ estimates it with uncertainty.
- A confidence interval gives a range for $\beta$; a test judges whether $\beta = 0$.
- Both quantify how sure we are about the line behind the scatter.
关于关系的推断
- 斜率推断问的是:到底有没有线性关系,以及有多强?
- 真实斜率 $\beta$ 未知;$b$ 带着不确定性估计它。
- 置信区间给出 $\beta$ 的一个范围;检验判断 $\beta = 0$ 是否成立。
- 两者都量化我们对散点背后那条线有多确信。
Built on Unit 2
- Everything rests on the least-squares regression 最小二乘回归 model $\hat{y} = a + bx$.
- $b$ is still the predicted change in $y$ per unit $x$ — now with error bars.
- We add a $t$-based interval and test around that slope.
- Regression inference is "Unit 2's slope, plus Unit 7's $t$-machinery."
建立在第 2 单元之上
- 一切都依赖于最小二乘回归模型 $\hat{y} = a + bx$。
- $b$ 仍然是 $x$ 每单位 $y$ 的预测变化量——现在带了误差棒。
- 我们在那个斜率周围加上一个基于 $t$ 的区间和检验。
- 回归推断是“第 2 单元的斜率,加上第 7 单元的 $t$ 机器”。
The slope $b$ from one sample is an estimate, not the truth $\beta$. A different random sample would give a different $b$. Inference for slopes is exactly the same logic as for means — a $t$-based interval and test around a sample statistic — applied to the regression slope, with its own standard error $SE_b$.
**一个样本得到的斜率 $b$ 是一个估计,而非真值 $\beta$。**另一个随机样本会给出不同的 $b$。斜率推断与均值推断的逻辑完全相同——围绕一个样本统计量的、基于 $t$ 的区间和检验——只是应用到回归斜率上,带着它自己的标准误 $SE_b$。
Fitting hours-studied vs. score for one class gives slope $b = 4.2$.
- A different class would give a different $b$ (maybe $3.8$ or $4.6$) — sampling variability.
- The true slope $\beta$ for all students is unknown; $b = 4.2$ estimates it.
- Unit 9 builds an interval and a test for that unknown $\beta$.
对一个班的学习时间对分数拟合,得到斜率 $b = 4.2$。
- 另一个班会给出不同的 $b$(也许 $3.8$ 或 $4.6$)——抽样变异性。
- 所有学生的真实斜率 $\beta$ 未知;$b = 4.2$ 估计它。
- 第 9 单元为那个未知的 $\beta$ 建立区间和检验。
The slope $b$ of a fitted line varies from sample to sample; its sampling distribution is centered at the true slope $\beta$ with spread $SE_b$, modeled by a $t$-distribution. Inference for slopes quantifies the uncertainty about the true linear relationship, building directly on the least-squares regression model.
拟合线的斜率 $b$ 在样本之间变动;它的抽样分布以真实斜率 $\beta$ 为中心、分散为 $SE_b$,由一个 $t$ 分布建模。斜率推断量化关于真实线性关系的不确定性,直接建立在最小二乘回归模型之上。
A fitted slope varies by sample · 拟合斜率随样本变动
Each sample's least-squares line has a slightly different slope b. · 每个样本的最小二乘线都有略微不同的斜率 b。
The slope b from a sample is... · 来自样本的斜率 b 是……
b estimates the unknown population slope β, with uncertainty. · b 带着不确定性估计未知的总体斜率 β。
The sampling distribution of the sample slope b is centered at the true slope β. · 样本斜率 b 的抽样分布以真实斜率 β 为中心。
b is unbiased for β. · b 是 β 的无偏估计。
Inference for slopes is modeled using which distribution? · 斜率推断用哪个分布建模?
Like means, slope inference uses a t-distribution. · 和均值一样,斜率推断用 t 分布。
Regression inference builds on the least-squares model ŷ = a + bx from Unit 2. · 回归推断建立在第 2 单元的最小二乘模型 ŷ = a + bx 之上。
It adds a t-interval and test around the Unit 2 slope. · 它在第 2 单元的斜率周围加上 t 区间和检验。
The spread of the sampling distribution of b is the standard error of the ___ (one word). · b 的抽样分布的分散是 ___ 的标准误(填英文一词 slope)。
SE_b is the standard error of the slope. · SE_b 是斜率的标准误。