English narration · English + 中文 subtitles burned in · การบรรยายภาษาอังกฤษ · คำบรรยายภาษาอังกฤษ + 中文 ลอยตัวบนภาพ
English
This handout covers Topic 4: Further Probability & Statistics 进阶概率统计. It adds continuous distributions, small-sample inference, the chi-squared and non-parametric tests, and probability generating functions.
Continuous random variables · ตัวแปรสุ่มต่อเนื่อง
Syllabus · หลักสูตร
English
Candidates should be able to:
Notes and examples
use a probability density function which may be defined piecewise
use the general result $\text{E}(g(X)) = \int f(x)g(x) \, \mathrm{d}x$ where $f(x)$ is the probability density function of the continuous random variable $X$ and $g(X)$ is a function of $X$
understand and use the relationship between the probability density function (PDF) and the cumulative distribution function (CDF), and use either to evaluate probabilities or percentiles
use cumulative distribution functions (CDFs) of related variables in simple cases.
e.g. given the CDF of a variable $X$, find the CDF of a related variable $Y$, and hence its PDF, e.g. where $Y = X^3$.
Source: Cambridge International syllabus · แหล่งที่มา: หลักสูตร Cambridge International
English
A continuous variable $X$ is described by a probability density function 概率密度函数$f(x)$, which may be defined piecewise. The probability over a range is the area under $f$, and the mean of any function of $X$ is
$$E(g(X)) = \int g(x)\,f(x)\,dx.$$
The cumulative distribution function 累积分布函数$F(x) = P(X \leqslant x)$ is the running total: $F(x) = \displaystyle\int_{-\infty}^{x} f(t)\,dt$, and $f(x) = F'(x)$. Use $F$ to find probabilities and percentiles 百分位数 (for example, the median 中位数 solves $F(x) = 0.5$).
Worked example. A variable has $f(x) = \tfrac12 x$ for $0 \leqslant x \leqslant 2$. Find the median.
The cumulative distribution function is $F(x) = \displaystyle\int_0^x \tfrac12 t\,dt = \tfrac14 x^2$. Set $F(m) = 0.5$:
The distribution of a related variable. If $Y=g(X)$, find $Y$'s distribution through its cumulative function. For $Y=X^2$: $F_Y(y)=P(X^2\leqslant y)=P(-\sqrt y\leqslant X\leqslant\sqrt y)=F_X(\sqrt y)-F_X(-\sqrt y)$, then differentiate for $f_Y=F_Y'$. When $g$ is monotonic increasing there is a shortcut, $F_Y(y)=F_X\big(g^{-1}(y)\big)$.
Worked example. With $f_X(x)=\tfrac12 x$ on $[0,2]$ (so $F_X(x)=\tfrac14 x^2$), let $Y=X^2$. For $0\leqslant y\leqslant 4$, $F_Y(y)=F_X(\sqrt y)=\tfrac14 y$, so $f_Y(y)=F_Y'(y)=\tfrac14$ – that is, $Y$ is uniform on $[0,4]$.
probability density function/ˌprɒbəˈbɪlɪti ˈdensɪti ˈfʌŋkʃn/
ฟังก์ชันความหนาแน่นความน่าจะเป็น
cumulative distribution function/ˈkjuːmjʊlətɪv ˌdɪstrɪˈbjuːʃn ˈfʌŋkʃn/
ฟังก์ชันการแจกแจงสะสม
percentiles/pəˈsentaɪlz/
percentiles
median/ˈmiːdiːən/
มัธยฐาน
hypothesis test/haɪˈpɒθəsɪs test/
การทดสอบสมมติฐาน
confidence interval/ˈkɒnfɪdəns ˈɪntəvl/
confidence interval
4.2
Inference using normal and t-distributions · การอนุมานโดยใช้การแจกแจงnormalและt
Syllabus · หลักสูตร
English
Candidates should be able to:
Notes and examples
formulate hypotheses and apply a hypothesis test concerning the population mean using a small sample drawn from a normal population of unknown variance, using a t-test
calculate a pooled estimate of a population variance from two samples
Calculations based on either raw or summarised data may be required.
formulate hypotheses concerning the difference of population means, and apply, as appropriate: - a 2-sample t-test - a paired sample t-test - a test using a normal distribution
The ability to select the test appropriate to the circumstances of a problem is expected.
determine a confidence interval for a population mean, based on a small sample from a normal population with unknown variance, using a t-distribution
determine a confidence interval for a difference of population means, using a t-distribution or a normal distribution, as appropriate.
กำหนด ช่วงความเชื่อมั่น สำหรับค่าเฉลี่ยประชากร โดยอ้างอิงจากตัวอย่างเล็กจากประชากรที่มีการแจกแจงปกติที่มีความแปรปรวนไม่ทราบค่า โดยใช้ การแจกแจง t
กำหนด ช่วงความเชื่อมั่น สำหรับผลต่างของค่าเฉลี่ยประชากร โดยใช้ การแจกแจง t หรือ การแจกแจงปกติ ตามความเหมาะสม
Source: Cambridge International syllabus · แหล่งที่มา: หลักสูตร Cambridge International
English
When a sample is small and the population variance is unknown, base your hypothesis test 假设检验 on the $t$-distribution instead of the normal. The same idea gives a confidence interval 置信区间 for the mean:
$$\bar{x} \pm t\,\frac{s}{\sqrt{n}},$$
where $t$ comes from the $t$-tables with $n - 1$ degrees of freedom. To compare two populations, use a two-sample (2-sample) or paired-sample$t$-test, after finding a pooled estimate 合并估计 of the shared variance when appropriate.
Worked example. A sample of $n = 10$ has mean $\bar{x} = 50$ and standard deviation $s = 4$. Find a $95\%$ confidence interval for the mean (use $t = 2.262$ for $9$ degrees of freedom).
Why small samples need t instead of z · ทำไมตัวอย่างขนาดเล็กจึงต้องใช้ t แทน z
With $\sigma$ unknown you use $t$, and $t$ has heavier tails than the normal (drawn dashed behind it) — so its critical values are larger and the interval is wider. At the worked example's $9$ degrees of freedom the widget reads $t^* = 2.262$, exactly the table value used above. Sweep df up and $t$ collapses onto the normal. · เมื่อ $\sigma$ ไม่ทราบใช้ $t$ และ $t$ มีหางหนากว่าปกติ (วาดด้วยเส้นประอยู่เบื้องหลัง) — ดังนั้นค่าวิกฤตจึง มากกว่า และช่วงมีความกว้างกว่า ที่ $9$ Degrees of freedom ของตัวอย่างที่คำนวณ Widget อ่านค่าได้ $t^* = 2.262$ ตรงกับค่าในตารางที่ใช้อ้างอิงข้างต้น ลาก df ขึ้นไปและ $t$ จะซ้อนทับกับกราฟปกติ
Explore · สำรวจ
The normal distribution · การแจกแจงปกติ
Shade a tail to find a probability — the basis of confidence intervals and hypothesis tests. · แรเงาหางเพื่อหาค่าความน่าจะเป็น — เป็นพื้นฐานของช่วงความเชื่อมั่นและการทดสอบสมมติฐาน
fit a theoretical distribution, as prescribed by a given hypothesis, to given data
Questions will not involve lengthy calculations.
use a $\chi^2$-test, with the appropriate number of degrees of freedom, to carry out the corresponding goodness of fit analysis
Classes should be combined so that each expected frequency is at least 5.
use a $\chi^2$-test, with the appropriate number of degrees of freedom, for independence in a contingency table.
Yates’ correction is not required. Where appropriate, either rows or columns should be combined so that the expected frequency in each cell is at least 5.
Source: Cambridge International syllabus · แหล่งที่มา: หลักสูตร Cambridge International
English
A $\chi^2$-test (chi-squared test 卡方检验) compares observed counts $O$ with expected counts $E$ from a theoretical distribution 理论分布:
$$\chi^2 = \sum \frac{(O - E)^2}{E}.$$
Compare this with a table value for the right number of degrees of freedom 自由度. Two uses: a goodness of fit 拟合优度 test (does the data follow the proposed model?), and a test for independence 独立性 of two variables in a contingency table 列联表.
Worked example. Four equally likely categories give observed counts $20, 30, 25, 25$ (so each expected count is $25$). Test the fit at the $5\%$ level.
The chi-squared distribution and its 5% tail · การแจกแจง chi-squared และหาง 5%
The worked example on this page gives $\chi^2 = 2$ with $3$ degrees of freedom against a table value of $7.815$ — the widget reproduces both. Drag df to see why the critical value changes with the number of categories. · ตัวอย่างที่คำนวณหน้านี้เป็น $\chi^2 = 2$ กับ $3$ Degrees of freedom เทียบกับค่าในตาราง $7.815$ — Widget นี้ทำซ้ำทั้งสองอย่าง ลาก df เพื่อดูว่าค่าวิกฤตเปลี่ยนตามจำนวนหมวดหมู่อย่างไร
Explore · สำรวจ
Chi-squared test route · เส้นทางทดสอบ Chi-squared
Follow observed and expected counts to a test decision. · ติดตามจำนวนสังเกตและจำนวนคาดหวังไปสู่การตัดสินใจทดสอบ
understand the idea of a non-parametric test and appreciate situations in which such a test might be useful
e.g. when sampling from a population which cannot be assumed to be normally distributed.
understand the basis of the sign test, the Wilcoxon signed-rank test and the Wilcoxon rank-sum test
Including knowledge that Wilcoxon tests are valid only for symmetrical distributions.
use a single-sample sign test and a single-sample Wilcoxon signed-rank test to test a hypothesis concerning a population median
Including the use of normal approximations where appropriate. Questions will not involve tied ranks or observations equal to the population median value being tested.
use a paired-sample sign test, a Wilcoxon matched-pairs signed-rank test and a Wilcoxon rank-sum test, as appropriate, to test for identity of populations.
Including the use of normal approximations where appropriate. Questions will not involve tied ranks or zero‑difference pairs.
Source: Cambridge International syllabus · แหล่งที่มา: หลักสูตร Cambridge International
English
A non-parametric test 非参数检验 makes no assumption that the data is normal, so it is useful when that assumption fails. The basic ones are:
the sign test 符号检验: count how many values fall above and below a proposed median, and test those counts with a binomial model;
the Wilcoxon signed-rank test 威尔科克森符号秩检验 (the matched-pairs test for paired data), which also uses the sizes of the differences, not just their signs;
the Wilcoxon rank-sum test 威尔科克森秩和检验, for comparing two separate samples.
Worked example. Test whether a median is $5$. In a sample of $10$ values (none equal to $5$), $9$ lie above$5$ and $1$ lies below. Test at the $5\%$ level (two-tailed).
Under $H_0$ (median $= 5$) the number above follows $B(10, 0.5)$. The observed result ($9$ above) is extreme, so find $P(X \geq 9) = \binom{10}{9}(0.5)^{10} + (0.5)^{10} = \dfrac{11}{1024} = 0.0107$. For a two-tailed test compare with $\tfrac{1}{2}(5\%) = 0.025$. Since $0.0107 < 0.025$, reject $H_0$: there is evidence the median is not $5$.
probability generating function/ˌprɒbəˈbɪlɪti ˈdʒenəreɪtɪŋ ˈfʌŋkʃn/
ฟังก์ชันสร้างความน่าจะเป็น
Further Probability & Statistics/ˈfɜːðə ˌprɒbəˈbɪlɪti ænd stəˈtɪstɪks/
ความน่าจะเป็นและสถิติเพิ่มเติม
4.5
Probability generating functions · ฟังก์ชันสร้างโอกาส (probability generating functions)
Syllabus · หลักสูตร
English
Candidates should be able to:
Notes and examples
understand the concept of a probability generating function (PGF) and construct and use the PGF for given distributions
Including the discrete uniform, binomial, geometric and Poisson distributions.
use formulae for the mean and variance of a discrete random variable in terms of its PGF, and use these formulae to calculate the mean and variance of a given probability distribution
use the result that the PGF of the sum of independent random variables is the product of the PGFs of those random variables.
Source: Cambridge International syllabus · แหล่งที่มา: หลักสูตร Cambridge International
English
The probability generating function 概率母函数 of a discrete variable $X$ is
$$G(t) = E(t^X) = \sum_x P(X = x)\,t^x.$$
It packs the whole distribution into one function. The mean and variance come from its derivatives at $t = 1$: $E(X) = G'(1)$ and $\mathrm{Var}(X) = G''(1) + G'(1) - \big(G'(1)\big)^2$. Also, the PGF of a sum of independent variables is the product of their PGFs.
Worked example.$X$ has $P(X=0) = 0.5$, $P(X=1) = 0.3$, $P(X=2) = 0.2$. Find $E(X)$ using the PGF.
Here $G(t) = 0.5 + 0.3t + 0.2t^2$, so $G'(t) = 0.3 + 0.4t$ and
More topics in A-Level Further Mathematics · Further Mathematics A-Level · หัวข้อเพิ่มเติมใน A-Level Further Mathematics · Further Mathematics A-Level
Pick one and the site follows you — notes, papers, videos and practice all open on it. · เลือกหนึ่งตัว และเว็บจะติดตามคุณ — หมายเหตุ, ใบงาน, วิดีโอ และการฝึกฝนจะเปิดอยู่ที่นั้น
Type to search notes, lessons, code, vocabulary and past-paper questions across every subject. · พิมพ์เพื่อค้นหาบันทึก, บทเรียน, โค้ด, คำศัพท์ และคำถามข้อสอบเก่าในทุกวิชา