跳到主要内容

探索单变量数据

AP 统计学 · 第 1 主题

训练
讲义 词汇表
1.1

统计学导论:我们能从数据中学到什么?

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

VAR-1
Given that variation may be random or not, conclusions are uncertain.

VAR-1.A
Identify questions to be answered, based on variation in one-variable data. [Skill 1.A]

  • VAR-1.A.1 Numbers may convey meaningful information, when placed in context.

来源:美国大学理事会 AP 课程与考试说明

统计学(statistics)是从数据(data)——从真实世界收集的数字或标签——中学习的科学。数据会变化,所以我们描述模式并考虑变异(variation),而不是期望每个值都相同。一个统计问题预期一个基于会变化的数据的答案。

有两个区分贯穿整个课程。参数(parameter)是整个总体的一个数值概括;统计量(statistic)是一个样本的数值概括——我们用统计量去估计无法直接测量的参数。而描述统计(descriptive statistics)只概括手头的数据集,推断统计(inferential statistics)则用一个样本对更大的总体做出并检验论断。

词汇表 训练
英文 中文 拼音
Statistics 统计学 tǒng jì xué
data 数据 shù jù
variation 变异 biàn yì
parameter 参数 cān shù
statistic 统计量 tǒng jì liàng
descriptive statistics 描述统计 miáo shù tǒng jì
inferential statistics 推断统计 tuī duàn tǒng jì
1.2

变异的语言:变量

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

VAR-1
Given that variation may be random or not, conclusions are uncertain.

VAR-1.B
Identify variables in a set of data. [Skill 2.A]

  • VAR-1.B.1 A variable is a characteristic that changes from one individual to another.

VAR-1.C
Classify types of variables. [Skill 2.A]

  • VAR-1.C.1 A categorical variable takes on values that are category names or group labels.
  • VAR-1.C.2 A quantitative variable is one that takes on numerical values for a measured or counted quantity.
    • Illustrative examples for VAR-1.C:
      • Categorical variables:
        • Dominant hand
        • Age group (young or old)
        • Highest degree earned
      • Quantitative variables:
        • Age of a structure
        • Height of a child
        • Concentration of a sample

来源:美国大学理事会 AP 课程与考试说明

一个变量(variable)是一个能在个体之间不同的特征。两种:

  • 分类(categorical)(定性):值是标签/组(眼睛颜色、品牌)。
  • 定量(quantitative):值是你能做算术的数字(身高、年龄)。定量变量是离散(discrete)(可数)或连续(continuous)(测量)的。

选择正确的图和概括取决于你有哪一种。

探索

Categorical or quantitative?

Every variable is either categorical (it labels each unit with a group) or quantitative (a measured number you can average). Which kind it is decides the graphs and summaries you are allowed to use.

词汇表 训练
英文 中文 拼音
variable 变量 biàn liàng
Categorical 分类 fēn lèi
Quantitative 定量 dìng liàng
1.3

用表格表示分类变量

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

UNC-1
Graphical representations and statistics allow us to identify and represent key features of data.

UNC-1.A
Represent categorical data using frequency or relative frequency tables. [Skill 2.B]

  • UNC-1.A.1 A frequency table gives the number of cases falling into each category. A relative frequency table gives the proportion of cases falling into each category.

UNC-1.B
Describe categorical data represented in frequency or relative tables. [Skill 2.A]

  • UNC-1.B.1 Percentages, relative frequencies, and rates all provide the same information as proportions.
  • UNC-1.B.2 Counts and relative frequencies of categorical data reveal information that can be used to justify claims about the data in context.

来源:美国大学理事会 AP 课程与考试说明

一个频数表(frequency table)列出每个类别的计数(count)(频数);一个相对频率(relative frequency)表列出每个类别的比例(proportion)(计数 ÷ 总数)。相对频率让你能公平地比较不同大小的组。

词汇表 训练
英文 中文 拼音
frequency table 频数表 pín shuò biǎo
relative frequency 相对频率 xiāng duì pín lǜ
proportion 比例 bǐ lì
1.4

用图形表示分类变量

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

UNC-1
Graphical representations and statistics allow us to identify and represent key features of data.

UNC-1.C
Represent categorical data graphically. [Skill 2.B]

  • UNC-1.C.1 Bar charts (or bar graphs) are used to display frequencies (counts) or relative frequencies (proportions) for categorical data.
  • UNC-1.C.2 The height or length of each bar in a bar graph corresponds to either the number or proportion of observations falling within each category.
  • UNC-1.C.3 There are many additional ways to represent frequencies (counts) or relative frequencies (proportions) for categorical data.

UNC-1.D
Describe categorical data represented graphically. [Skill 2.A]

  • UNC-1.D.1 Graphical representations of a categorical variable reveal information that can be used to justify claims about the data in context.

UNC-1.E
Compare multiple sets of categorical data. [Skill 2.D]

  • UNC-1.E.1 Frequency tables, bar graphs, or other representations can be used to compare two or more data sets in terms of the same categorical variable.

来源:美国大学理事会 AP 课程与考试说明

条形图(bar charts)把每个类别的计数或比例显示为分开的条;一个饼图(pie chart)显示每个类别在整体里的份额。条的高度(或扇区)让你能一眼比较类别。条可以按大小或按自然的类别顺序排列。

探索

Show a categorical variable as a pie chart

A pie chart turns each category's share of the whole into a slice: a bigger share is a bigger slice, and every slice together makes 100%. It is a picture of a relative-frequency table.

词汇表 训练
英文 中文 拼音
Bar charts 条形图 tiáo xíng tú
1.5

用图形表示定量变量

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

UNC-1
Graphical representations and statistics allow us to identify and represent key features of data.

UNC-1.F
Classify types of quantitative variables. [Skill 2.A]

  • UNC-1.F.1 A discrete variable can take on a countable number of values. The number of values may be finite or countably infinite, as with the counting numbers.
  • UNC-1.F.2 A continuous variable can take on infinitely many values, but those values cannot be counted. No matter how small the interval between two values of a continuous variable, it is always possible to determine another value between them.
    • Illustrative examples for UNC-1.F:
      • A discrete variable:
        • Number of students in a class
      • A continuous variable:
        • Height of a child

UNC-1.G
Represent quantitative data graphically. [Skill 2.B]

  • UNC-1.G.1 In a histogram, the height of each bar shows the number or proportion of observations that fall within the interval corresponding to that bar. Altering the interval widths can change the appearance of the histogram.
  • UNC-1.G.2 In a stem and leaf plot, each data value is split into a "stem" (the first digit or digits) and a "leaf" (usually the last digit).
  • UNC-1.G.3 A dotplot represents each observation by a dot, with the position on the horizontal axis corresponding to the data value of that observation, with nearly identical values stacked on top of each other.
  • UNC-1.G.4 A cumulative graph represents the number or proportion of a data set less than or equal to a given number.
  • UNC-1.G.5 There are many additional ways to graphically represent distributions of quantitative data.

来源:美国大学理事会 AP 课程与考试说明

对于数字,用一个点图(dotplot)、茎叶图(stem-and-leaf plot),或直方图(histogram)(在称为区间的值区间上的条)。这些显示分布(distribution)——值如何散开。一个直方图的区间宽度改变图景,所以选择它以揭示形状。

On a histogram with unequal class widths the bar area is the frequency
在一个不等组宽的直方图上条的面积是频数
探索

Explore how bin width shapes a histogram

A histogram groups data into equal-width bins and draws a bar over each. Change the bins and notice how the same data can look jagged (too narrow) or smooth (too wide) — the shape is a choice.

词汇表 训练
英文 中文 拼音
dotplot 点图 diǎn tú
stem-and-leaf plot 茎叶图 jīng yè tú
histogram 直方图 zhí fāng tú
distribution 分布 fēn bù
1.6

描述定量变量的分布

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

UNC-1
Graphical representations and statistics allow us to identify and represent key features of data.

UNC-1.H
Describe the characteristics of quantitative data distributions. [Skill 2.A]

  • UNC-1.H.1 Descriptions of the distribution of quantitative data include shape, center, and variability (spread), as well as any unusual features such as outliers, gaps, clusters, or multiple peaks.
  • UNC-1.H.2 Outliers for one-variable data are data points that are unusually small or large relative to the rest of the data.
  • UNC-1.H.3 A distribution is skewed to the right (positive skew) if the right tail is longer than the left. A distribution is skewed to the left (negative skew) if the left tail is longer than the right. A distribution is symmetric if the left half is the mirror image of the right half.
  • UNC-1.H.4 Univariate graphs with one main peak are known as unimodal. Graphs with two prominent peaks are bimodal. A graph where each bar height is approximately the same (no prominent peaks) is approximately uniform.
  • UNC-1.H.5 A gap is a region of a distribution between two data values where there are no observed data.
  • UNC-1.H.6 Clusters are concentrations of data usually separated by gaps.
  • UNC-1.H.7 Descriptive statistics does not attribute properties of a data set to a larger population, but may provide the basis for conjectures for subsequent testing.

来源:美国大学理事会 AP 课程与考试说明

描述四件事(记住 SOCS):

  • 形状(shape):对称,或偏斜(skewed)左/右(那一侧一条长尾),以及有多少个峰——一个主峰是单峰(unimodal),两个明显的峰是双峰(bimodal),各柱大致相等是均匀(uniform)。
  • 离群值(outliers):远离其余的不寻常的值。
  • 中心(center):一个典型的值(均值或中位数)。
  • 散布(spread):值变化多少(范围、IQR、标准差)。

总是在上下文里、带单位地描述形状/中心/散布。

The shape of a distribution: symmetric, skewed right (long right tail), or skewed left
一个分布的形状:对称、右偏(长右尾),或左偏
词汇表 训练
英文 中文 拼音
Shape 形状 xíng zhuàng
skewed 偏斜 piān xié
unimodal 单峰 dān fēng
bimodal 双峰 shuāng fēng
uniform 均匀 jūn yún
Outliers 离群值 lí qún zhí
1.7

定量变量的汇总统计量

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

UNC-1
Graphical representations and statistics allow us to identify and represent key features of data.

UNC-1.I
Calculate measures of center and position for quantitative data. [Skill 2.C]

  • UNC-1.I.1 A statistic is a numerical summary of sample data.
  • UNC-1.I.2 The mean is the sum of all the data values divided by the number of values. For a sample, the mean is denoted by $x$-bar: $\bar{x} = \dfrac{1}{n}\sum_{i=1}^{n} x_i$, where $x_i$ represents the $i^{\text{th}}$ data point in the sample and $n$ represents the number of data values in the sample.
  • UNC-1.I.3 The median of a data set is the middle value when data are ordered. When the number of data points is even, the median can take on any value between the two middle values. In AP Statistics, the most commonly used value for the median of a data set with an even number of values is the average of the two middle values.
  • UNC-1.I.4 The first quartile, Q1, is the median of the half of the ordered data set from the minimum to the position of the median. The third quartile, Q3, is the median of the half of the ordered data set from the position of the median to the maximum. Q1 and Q3 form the boundaries for the middle 50% of values in an ordered data set.
  • UNC-1.I.5 The $p^{\text{th}}$ percentile is interpreted as the value that has $p\%$ of the data less than or equal to it.

UNC-1.J
Calculate measures of variability for quantitative data. [Skill 2.C]

  • UNC-1.J.1 Three commonly used measures of variability (or spread) in a distribution are the range, interquartile range, and standard deviation.
  • UNC-1.J.2 The range is defined as the difference between the maximum data value and the minimum data value. The interquartile range (IQR) is defined as the difference between the third and first quartiles: $Q3 - Q1$. Both the range and the interquartile range are possible ways of measuring variability of the distribution of a quantitative variable.
  • UNC-1.J.3 Standard deviation is a way to measure variability of the distribution of a quantitative variable. For a sample, the standard deviation is denoted by $s$: $s_x = \sqrt{\dfrac{1}{n-1}\sum(x_i - \bar{x})^2}$. The square of the sample standard deviation, $s^2$, is called the sample variance.
  • UNC-1.J.4 Changing units of measurement affects the values of the calculated statistics.

UNC-1.K
Explain the selection of a particular measure of center and/or variability for describing a set of quantitative data. [Skill 4.B]

  • UNC-1.K.1 There are many methods for determining outliers. Two methods frequently used in this course are:
    • UNC-1.K.1.i An outlier is a value greater than $1.5 \times \text{IQR}$ above the third quartile or more than $1.5 \times \text{IQR}$ below the first quartile.
    • UNC-1.K.1.ii An outlier is a value located 2 or more standard deviations above, or below, the mean.
  • UNC-1.K.2 The mean, standard deviation, and range are considered nonresistant (or non-robust) because they are influenced by outliers. The median and IQR are considered resistant (or robust), because outliers do not greatly (if at all) affect their value.

来源:美国大学理事会 AP 课程与考试说明

标准差:数据相对平均值的离散
  • 中心: 均值(mean)$\bar{x}=\dfrac{\sum x_i}{n}$(平均)和中位数(median)(中间的值)。中位数抵抗离群值;均值被拉向偏斜。
  • 散布: 范围(range)、四分位距(interquartile range)$\text{IQR}=Q_3-Q_1$(中间 50%),以及标准差(standard deviation)$s_x=\sqrt{\dfrac{\sum(x_i-\bar{x})^2}{n-1}}$(离均值的典型距离;它的平方是方差(variance))。
  • 五数概括(five-number summary):min、$Q_1$、中位数、$Q_3$、max。

对偏斜数据用抵抗性(resistant)测度(中位数、IQR);对大致对称的数据用均值和标准差。

一个值的百分位数(percentile)是数据中在它或以下的百分比——所以中位数是第 50 百分位数,$Q_1$ 是第 25。一个累积相对频率图(cumulative relative frequency graph)让百分位数容易读出:对每个值它画出数据中在它或以下的比例,从 0 上升到 1。从一个值向上到曲线再横过去到它的百分位数,或反过来找一个给定百分位数处的值(同样的读法对一张累积频率表也有效)。

Worked example. 对于数据 $4, 8, 6, 10, 7$:均值是 $\bar{x}=\dfrac{4+8+6+10+7}{5}=\dfrac{35}{5}=7$。排序到 $4,6,7,8,10$,中位数是中间的值,$7$。这里均值和中位数一致,因为数据大致对称。

词汇表 训练
英文 中文 拼音
mean 均值 jūn zhí
median 中位数 zhōng wèi shù
interquartile range 四分位距 sì fēn wèi jù
standard deviation 标准差 biāo zhǔn chà
variance 方差 fāng chà
five-number summary 五数概括 wǔ shù gài kuò
percentile 百分位数 bǎi fēn wèi shù
cumulative relative frequency graph 累积相对频率图 lěi jī xiāng duì pín lǜ tú
练习卷
1.8

汇总统计量的图形表示

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

UNC-1
Graphical representations and statistics allow us to identify and represent key features of data.

UNC-1.L
Represent summary statistics for quantitative data graphically. [Skill 2.B]

  • UNC-1.L.1 Taken together, the minimum data value, the first quartile (Q1), the median, the third quartile (Q3), and the maximum data value make up the five-number summary.
  • UNC-1.L.2 A boxplot is a graphical representation of the five-number summary (minimum, first quartile, median, third quartile, maximum). The box represents the middle 50% of data, with a line at the median and the ends of the box corresponding to the quartiles. Lines ("whiskers") extend from the quartiles to the most extreme point that is not an outlier, and outliers are indicated by their own symbol beyond this.

UNC-1.M
Describe summary statistics of quantitative data represented graphically. [Skill 2.A]

  • UNC-1.M.1 Summary statistics of quantitative data, or of sets of quantitative data, can be used to justify claims about the data in context.
  • UNC-1.M.2 If a distribution is relatively symmetric, then the mean and median are relatively close to one another. If a distribution is skewed right, then the mean is usually to the right of the median. If the distribution is skewed left, then the mean is usually to the left of the median.

来源:美国大学理事会 AP 课程与考试说明

一个箱线图(boxplot)画五数概括:一个从 $Q_1$$Q_3$、中位数在里面的箱,以及到最极端的非离群值的须。一个点是一个离群值,若它落在一个四分位数之外超过 $1.5\times\text{IQR}$ ——一个你可能被要求应用的规则。箱线图对并排比较几个组很理想。

Worked example. 一个数据集有 $Q_1=20$$Q_3=32$,所以 $\text{IQR}=12$。离群值围栏是 $Q_1-1.5(12)=2$$Q_3+1.5(12)=50$。任何低于 $2$ 或高于 $50$ 的值被标记为一个离群值。

A box-and-whisker plot shows the quartiles and the range
一个箱须图显示四分位数和范围
A boxplot draws the five-number summary; the box spans the IQR
一个箱线图画五数概括;箱跨越 IQR
探索

Explore the five-number summary as a boxplot

Drag $Q_1$, the median, and $Q_3$ to see the box (its length is the IQR) and how the median's position inside the box reveals skew — a median close to $Q_1$ signals a right-skewed distribution.

词汇表 训练
英文 中文 拼音
boxplot 箱线图 xiāng xiàn tú
1.9

比较定量变量的分布

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

UNC-1
Graphical representations and statistics allow us to identify and represent key features of data.

UNC-1.N
Compare graphical representations for multiple sets of quantitative data. [Skill 2.D]

  • UNC-1.N.1 Any of the graphical representations, e.g., histograms, side-by-side boxplots, etc., can be used to compare two or more independent samples on center, variability, clusters, gaps, outliers, and other features.

UNC-1.O
Compare summary statistics for multiple sets of quantitative data. [Skill 2.D]

  • UNC-1.O.1 Any of the numerical summaries (e.g., mean, standard deviation, relative frequency, etc.) can be used to compare two or more independent samples.

来源:美国大学理事会 AP 课程与考试说明

要比较两个或更多组,比较形状、中心和散布,并提及离群值——总是用比较性词语("A 组有一个更高的中位数比 B 组")并在上下文里。不要只是分别描述每个组;把比较显式化。

探索

Compare distributions with box plots

A box plot draws the five-number summary. Placing two box plots on the same scale compares their centre (median), spread (IQR = box width) and skew at a glance — the fair way to compare groups.

1.10

正态分布

大纲
Enduring UnderstandingLearning ObjectiveEssential Knowledge

VAR-2
The normal distribution can be used to represent some population distributions.

VAR-2.A
Compare a data distribution to the normal distribution model. [Skill 2.D]

  • VAR-2.A.1 A parameter is a numerical summary of a population.
  • VAR-2.A.2 Some sets of data may be described as approximately normally distributed. A normal curve is mound-shaped and symmetric. The parameters of a normal distribution are the population mean, $\mu$, and the population standard deviation, $\sigma$.
  • VAR-2.A.3 For a normal distribution, approximately 68% of the observations are within 1 standard deviation of the mean, approximately 95% of observations are within 2 standard deviations of the mean, and approximately 99.7% of observations are within 3 standard deviations of the mean. This is called the empirical rule.
  • VAR-2.A.4 Many variables can be modeled by a normal distribution.
    • Illustrative examples for VAR-2.A:
      • Variables that can be modeled by a normal distribution:
        • Body temperature
        • Weight of a loaf of bread

VAR-2.B
Determine proportions and percentiles from a normal distribution. [Skill 3.A]

  • VAR-2.B.1 A standardized score for a particular data value is calculated as (data value − mean)/(standard deviation), and measures the number of standard deviations a data value falls above or below the mean.
  • VAR-2.B.2 One example of a standardized score is a $z$-score, which is calculated as $z\text{-score} = \left(\dfrac{x_i - \mu}{\sigma}\right)$. A $z$-score measures how many standard deviations a data value is from the mean.
  • VAR-2.B.3 Technology, such as a calculator, a standard normal table, or computer-generated output, can be used to find the proportion of data values located on a given interval of a normally distributed random variable.
  • VAR-2.B.4 Given the area of a region under the graph of the normal distribution curve, it is possible to use technology, such as a calculator, a standard normal table, or computer-generated output, to estimate parameters for some populations.

VAR-2.C
Compare measures of relative position in data sets. [Skill 2.D]

  • VAR-2.C.1 Percentiles and $z$-scores may be used to compare relative positions of points within a data set or between data sets.

来源:美国大学理事会 AP 课程与考试说明

一个正态分布(normal distribution)是一个由它的均值 $\mu$ 和标准差 $\sigma$ 描述的对称、钟形模型。经验法则(empirical rule)(68–95–99.7):约 68% 的值落在离均值 $1\sigma$ 内、95% 在 $2\sigma$ 内,而 99.7% 在 $3\sigma$ 内。

The normal curve: a probability is the area under it, centred on the mean
正态曲线:一个概率是它下面的面积,以均值为中心

一个**$z$ 分数**(z-score)测量一个值离均值多少个标准差:

$$z=\frac{x-\mu}{\sigma}.$$
转换成一个 $z$ 分数,然后用正态表或技术求以下的、以上的,或之间的比例(面积)——并反转这个过程以从一个给定百分位数求一个值。

Worked example. 测验分数是正态的,$\mu=500$$\sigma=100$。一个 $700$ 的分数有 $z=\dfrac{700-500}{100}=2$。由经验法则,$95\%$ 的分数落在 $2\sigma$ 内,所以 $2.5\%$ 落在 $700$ 以上——意味着一个 $700$ 大约在第 $97.5$ 百分位。

The normal curve and the 68-95-99.7 empirical rule
正态曲线和 68-95-99.7 经验法则
探索

Explore area under the normal curve

The proportion of data below a value equals the area under the curve to its left. Shade a tail or a central band to see the 68–95–99.7 empirical rule and read a $z$-score as an area.

词汇表 训练
英文 中文 拼音
normal distribution 正态分布 zhèng tài fēn bù
empirical rule 经验法则 jīng yàn fǎ zé
$z$-score 标准分数 biāo zhǔn fēn shù
1.10

考试技巧

  • 形状、中心、散布和离群值(SOCS)描述一个分布——总是在上下文里。
  • 均值被离群值拉动;中位数抵抗它们,所以对偏斜数据首选中位数。
  • 对一个正态分布用 68–95–99.7 法则和 z 分数 $z=\tfrac{x-\mu}{\sigma}$
  • 用并排箱线图比较分布并评论中心、散布和形状。
  • 标准差测量离均值的一个典型距离;IQR 与中位数配对。

本主题的互动课程

逐步学习,并即时检测练习。

AP 统计学历年真题

AP 统计学的更多主题

登录或创建账号

IGCSE, A-Level & AP