Skip to content · ⁨コンテンツへスキップ⁩

Inference for Categorical Data: Chi-Square · ⁨カテゴリーデータのための推論:カイ二乗⁩

AP Statistics · Topic 8 · ⁨トピック 8⁩

View Slides · ⁨查看幻灯片⁩ Train · ⁨練習する⁩
Video lesson for this topic · ⁨このトピックのビデオレッスン⁩ Open the video page · ⁨動画ページを開く⁩
7:49

カテゴリーデータのための推論:カイ二乗

公平なサイコロを六十回振ります。各面は十回ずつ出るはずです。しかし、そうはなりません—ここ八回、あそこ十三回。ギャップは常に存在します。もう一度振って…

English narration · English + 中文 subtitles burned in · ⁨英語ナレーション・英語+中文字幕 burning-in⁩

8.1

Are My Results Unexpected? · ⁨結果が予期せぬものですか?⁩

Syllabus · ⁨シラバス⁩
English

Enduring Understanding (VAR-1): Given that variation may be random or not, conclusions are uncertain.

Learning Objective VAR-1.J: Identify questions suggested by variation between observed and expected counts in categorical data. [Skill 1.A]

  • VAR-1.J.1 Variation between what we find and what we expect to find may be random or not.
日本語

長期的理解(VAR-1): ばらつきがランダムである場合でもそうでない場合でも、結論には不確実性が伴う。

学習目標 VAR-1.J: 分類データにおいて観測値と期待値の間のばらつきによって提起される質問を特定せよ。[スキル 1.A]

  • VAR-1.J.1 私たちが見つけるものと期待するものの間のばらつきは、偶然によるものかもしれないし、そうでないかもしれない。

Source: College Board AP Course and Exam Description · ⁨出典: College Board AP コースおよび試験説明書⁩

English

When data are counts spread across several categories, we test whether the observed counts differ from what a claim predicts. The tool is the chi-square 卡方 ($\chi^2$) statistic, which adds up the standardized gaps between observed and expected counts:

$$\chi^2=\sum \frac{(\text{observed}-\text{expected})^2}{\text{expected}}.$$
A large $\chi^2$ means the observed counts are far from expected – evidence against the claim. The chi-square distribution is right-skewed and depends on its degrees of freedom 自由度.

日本語

データが複数のカテゴリにわたってカウントとして分布する場合、観測されたカウントが主張で予測される値と異なるかどうかを検定します。ツールはカイ二乗 ($\chi^2$) 統計量であり、観測カウントと期待カウントの標準化された乖離の総和です:

$$\chi^2=\sum \frac{(\text{observed}-\text{expected})^2}{\text{expected}}.$$
大きな$\chi^2$は、観測カウントが期待値から大きく外れていることを意味し、主張に対する証拠となります。カイ二乗分布は右に歪んでおり、自由度に依存します。

8.2

Setting Up a Goodness-of-Fit Test · ⁨適合度検定の設定⁩

Syllabus · ⁨シラバス⁩
English

Enduring Understanding (VAR-8): The chi-square distribution may be used to model variation.

Learning Objective VAR-8.A: Describe chi-square distributions. [Skill 3.C]

  • VAR-8.A.1 Expected counts of categorical data are counts consistent with the null hypothesis. In general, an expected count is a sample size times a probability.

    The chi-square statistic measures the distance between observed and expected counts relative to expected counts.

    Chi-square distributions have positive values and are skewed right. Within a family of density curves, the skew becomes less pronounced with increasing degrees of freedom.

Learning Objective VAR-8.B: Identify the null and alternative hypotheses in a test for a distribution of proportions in a set of categorical data. [Skill 1.F]

  • VAR-8.B.1 For a chi-square goodness-of-fit test, the null hypothesis specifies null proportions for each category, and the alternative hypothesis is that at least one of these proportions is not as specified in the null hypothesis.

Learning Objective VAR-8.C: Identify an appropriate testing method for a distribution of proportions in a set of categorical data. [Skill 1.E]

  • VAR-8.C.1 When considering a distribution of proportions for one categorical variable, the appropriate test is the chi-square test for goodness of fit.

Learning Objective VAR-8.D: Calculate expected counts for the chi-square test for goodness of fit. [Skill 3.A]

  • VAR-8.D.1 Expected counts for a chi-square goodness-of-fit test are (sample size)(null proportion).

Learning Objective VAR-8.E: Verify the conditions for making statistical inferences when testing goodness of fit for a chi-square distribution. [Skill 4.C]

  • VAR-8.E.1 In order to make statistical inferences for a chi-square test for goodness of fit we must check the following:
    • a. To check for independence:
      • i. Data should be collected using a random sample or randomized experiment.
      • ii. When sampling without replacement, check that $n \leq 10\%N$.
    • b. The chi-square test for goodness of fit becomes more accurate with more observations, so large counts should be used (shape).
      • i. A conservative check for large counts is that all expected counts should be greater than 5.
日本語

持続的理解 (VAR-8): カイ二乗分布は、ばらつきのモデル化に使用できる。

学習目標 VAR-8.A: カイ二乗分布を説明せよ。[スキル 3.C]

  • VAR-8.A.1 分類データの期待度数とは、帰無仮説と整合する度数である。一般的に、期待度数はサンプルサイズに確率を掛けたものである。

    カイ二乗統計量は、観測度数と期待度数の間の距離を、期待度数に対して相対的に測定する。

    カイ二乗分布は正の値を取り、右に歪む。密度曲線のファミリー内では、自由度が増加するにつれて歪みが小さくなる。

学習目標 VAR-8.B: 分類データの比率分布に対する検定における帰無仮説と対立仮説を特定せよ。[スキル 1.F]

  • VAR-8.B.1 カイ二乗適合度検定において、帰無仮説は各カテゴリーに対する帰無比率を指定し、対立仮説はこれらの比率のうち少なくとも1つが帰無仮説で指定された通りでないというものである。

学習目標 VAR-8.C: 分類データの比率分布に対して適切な検定手法を特定せよ。[スキル 1.E]

  • VAR-8.C.1 1つの分類変数に対する比率分布を検討する場合、適切な検定はカイ二乗適合度検定である。

学習目標 VAR-8.D: カイ二乗適合度検定のための期待度数を計算せよ。[スキル 3.A]

  • VAR-8.D.1 カイ二乗適合度検定のための期待度数は(サンプルサイズ)×(帰無比率)である。

学習目標 VAR-8.E: カイ二乗分布の適合度検定を行う際の推論の条件を確認せよ。[スキル 4.C]

  • VAR-8.E.1 カイ二乗適合度検定に対する統計的推論を行うためには、以下の点を確認しなければならない:
    • a. 独立性を確認するために:
      • i. データはランダムサンプリングまたは無作為割付け実験を用いて収集されなければならない。
      • ii. 非復元抽出を行う場合は、$n \leq 10\%N$ を確認する。
    • b. カイ二乗適合度検定は観察数が多いほど正確になるため、大きな度数を使用する(形状)。
      • i. 大きな度数に対する保守的なチェックとしては、すべての期待度数が5より大きいことが挙げられる。

Source: College Board AP Course and Exam Description · ⁨出典: College Board AP コースおよび試験説明書⁩

English
The chi-square (χ²) test

A goodness-of-fit (GOF) 拟合优度 test checks whether one categorical variable follows a claimed distribution (e.g. "the die is fair"). Hypotheses:

$$H_0:\text{the distribution is as claimed}\qquad H_a:\text{at least one proportion differs}.$$
Expected count for each category $=n\times(\text{claimed proportion})$. Conditions: random sample, all expected counts $\ge 5$, and the 10% condition.

日本語
カイ二乗 (χ²) 検定

適合度 (GOF) 検定は、単一のカテゴリ変数が主張された分布に従っているか確認します(例:「サイコロは公平である」)。仮説:

$$H_0:\text{the distribution is as claimed}\qquad H_a:\text{at least one proportion differs}.$$
各カテゴリの期待カウント $=n\times(\text{claimed proportion})$。条件:無作為標本、すべての期待カウント $\ge 5$、および10%条件。

カイ二乗分布とその右側拒絶領域
カイ二乗分布は右に歪んでいます。大きい統計量は有意水準を超えた影付きの右側に位置し、ここでモデルを棄却します。
Vocabulary · ⁨語彙⁩ Train · ⁨練習する⁩
English 日本語
chi-square/kaɪ skweə/ カイ二乗
degrees of freedom/dɪˈɡriːz ɒv ˈfriːdəm/ 自由度
goodness-of-fit (GOF)/ˈɡʊdnəs ɒv fɪt/ 適合度検定(GOF)
Test for homogeneity/test fɔː ˈhɒməʊdʒneɪti/ 均質性検定
Test for independence/test fɔː ˌɪndɪˈpendəns/ 独立性検定
8.3

Carrying Out a Goodness-of-Fit Test · ⁨適合度検定の実施⁩

Syllabus · ⁨シラバス⁩
English

Enduring Understanding (VAR-8): The chi-square distribution may be used to model variation.

Learning Objective VAR-8.F: Calculate the appropriate statistic for the chi-square test for goodness of fit. [Skill 3.E]

  • VAR-8.F.1 The test statistic for the chi-square test for goodness of fit is
    • Equation: $\chi^2 = \sum \dfrac{(Observed\ count - Expected\ count)^2}{Expected\ count}$, with $degrees\ of\ freedom = number\ of\ categories - 1$.
  • VAR-8.F.2 The distribution of the test statistic assuming the null hypothesis is true (null distribution) can be either a randomization distribution or, when a probability model is assumed to be true, a theoretical distribution (chi-square).

Learning Objective VAR-8.G: Determine the $p$-value for chi-square test for goodness of fit significance test. [Skill 3.E]

  • VAR-8.G.1 The $p$-value for a chi-square test for goodness of fit for a number of degrees of freedom is found using the appropriate table or computer generated output.

Enduring Understanding (DAT-3): Significance testing allows us to make decisions about hypotheses within a particular context.

Learning Objective DAT-3.I: Interpret the $p$-value for the chi-square test for goodness of fit. [Skill 4.B]

  • DAT-3.I.1 An interpretation of the $p$-value for the chi-square test for goodness of fit is the probability, given the null hypothesis and probability model are true, of obtaining a test statistic as, or more, extreme than the observed value.

Learning Objective DAT-3.J: Justify a claim about the population based on the results of a chi-square test for goodness of fit. [Skill 4.E]

  • DAT-3.J.1 A decision to either reject or fail to reject the null hypothesis is based on comparison of the $p$-value to the significance level, $\alpha$.
  • DAT-3.J.2 The results of a chi-square test for goodness of fit can serve as the statistical reasoning to support the answer to a research question about the population that was sampled.
日本語

持続的理解 (VAR-8): カイ二乗分布は、ばらつきのモデル化に使用できる。

学習目標 VAR-8.F: カイ二乗適合度検定に必要な統計量を計算せよ。[スキル 3.E]

  • VAR-8.F.1 カイ二乗適合度検定のための検定統計量は
    • 式: $\chi^2 = \sum \dfrac{(Observed\ count - Expected\ count)^2}{Expected\ count}$、ここで$degrees\ of\ freedom = number\ of\ categories - 1$。
  • VAR-8.F.2 帰無仮説が真であるとしたときの検定統計量の分布(帰無分布)は、無作為化分布であるか、確率モデルが真であると仮定される場合は理論分布(カイ二乗分布)である。

学習目標 VAR-8.G: カイ二乗適合度検定における$p$値を求める。[スキル 3.E]

  • VAR-8.G.1 自由度のあるカイ二乗適合度検定における$p$値は、適切な表やコンピュータ出力を用いて求める。

持続的理解 (DAT-3): 有意性検定により、特定の文脈内で仮説に関する判断を下すことができます。

学習目標 DAT-3.I: カイ二乗適合度検定における$p$値を解釈せよ。[スキル 4.B]

  • DAT-3.I.1 カイ二乗適合度検定における$p$値の解釈は、帰無仮説と確率モデルが真である場合に、観測値と同じか、それよりも極端な値の検定統計量を得る確率である。

学習目標 DAT-3.J: カイ二乗適合度検定の結果に基づき、母に関する主張を正当化せよ。[スキル 4.E]

  • DAT-3.J.1 帰無仮説を棄却するか否かの決定は、$p$値と有意水準$\alpha$を比較することに基づく。
  • DAT-3.J.2 カイ二乗適合度検定の結果は、サンプルされた母に関する研究問題への回答を裏付ける統計的根拠として機能しうる。

Source: College Board AP Course and Exam Description · ⁨出典: College Board AP コースおよび試験説明書⁩

English

Compute $\chi^2=\sum\dfrac{(O-E)^2}{E}$ with $df=(\text{number of categories})-1$. Find the $p$-value from the chi-square distribution (upper tail), compare to $\alpha$, and conclude in context. A large component of the sum points to the category that deviates most.

Worked example. A die rolled $60$ times gives counts $8,10,12,9,11,10$. If it is fair, each expected count is $60/6=10$, so

$$\chi^2=\frac{(8-10)^2}{10}+\frac{(10-10)^2}{10}+\frac{(12-10)^2}{10}+\frac{(9-10)^2}{10}+\frac{(11-10)^2}{10}+\frac{(10-10)^2}{10}=0.4+0+0.4+0.1+0.1+0=1.0,$$
with $df=6-1=5$. Write out every category, including the two that match their expected count exactly and so add $0$ – the sum runs over all six categories, and $df$ counts categories, not just the ones that differ. This $\chi^2$ is small (a large $p$-value), so we fail to reject $H_0$ – no evidence the die is unfair.

日本語

$\chi^2=\sum\dfrac{(O-E)^2}{E}$ を $df=(\text{number of categories})-1$ で計算する。カイ二乗分布(上端尾部)から $p$ 値を求め、 $\alpha$ と比較して文脈に基づいて結論を出す。和の大きな成分は、最も逸脱したカテゴリを指す。

カイ二乗は観測カウントと帰無仮説に基づく期待カウントを比較します
カイ二乗は観測カウントと帰無仮説に基づく期待カウントを比較します

** worked example.** サイコロを$60$回振り、カウントが$8,10,12,9,11,10$となりました。公平であれば、各期待カウントは$60/6=10$なので、

$$\chi^2=\frac{(8-10)^2}{10}+\frac{(10-10)^2}{10}+\frac{(12-10)^2}{10}+\frac{(9-10)^2}{10}+\frac{(11-10)^2}{10}+\frac{(10-10)^2}{10}=0.4+0+0.4+0.1+0.1+0=1.0,$$
は$df=6-1=5$です。すべてのカテゴリを書き出し、期待カウントと完全に一致する2つを含む必要があります。これらは$0$を加えるため、合計は6つのカテゴリすべてにわたるものであり、$df$は単に異なるカテゴリだけでなくカテゴリ数を数えます。この$\chi^2$は小さい($p$-値が大きい)ため、$H_0$を棄却できない—— サイコロが不公平である証拠はありません。

Explore · ⁨探索⁩

Explore the chi-square distribution and its p-value · ⁨カイ二乗分布とそのp値を調べる⁩

The p-value is the area in the right tail beyond your test statistic, so a larger $\chi^2$ means a smaller p-value. Drag $\chi^2$ to watch that area shrink, and drag df to see the whole family change shape — strongly right-skewed at small df, more symmetric as df grows. · ⁨p値は検定統計量より右側にある尾の面積であり、大きな $\chi^2$ は 小さい p値を意味する。$\chi^2$ をドラッグしてその面積が縮小する様子を確認し、自由度 (df) を動かして分布族全体の形状変化を見る(自由度が小さいときは強い右偏、自由度が増えると対称性を持つ)。⁩

8.4

Expected Counts in Two-Way Tables · ⁨2次元表における期待カウント⁩

Syllabus · ⁨シラバス⁩
English

Enduring Understanding (VAR-8): The chi-square distribution may be used to model variation.

Learning Objective VAR-8.H: Calculate expected counts for two-way tables of categorical data. [Skill 3.A]

  • VAR-8.H.1 The expected count in a particular cell of a two-way table of categorical data can be calculated using the formula:
    • Equation: $expected\ count = \dfrac{(row\ total)(column\ total)}{table\ total}$.
日本語

持続的理解 (VAR-8): カイ二乗分布は、ばらつきのモデル化に使用できる。

学習目標 VAR-8.H: 分類データの2次元表における期待度数を計算せよ。[スキル 3.A]

  • VAR-8.H.1 分類データの2次元表の特定のセルにおける期待度数は、以下の式を用いて計算できる:
    • 式: $expected\ count = \dfrac{(row\ total)(column\ total)}{table\ total}$。

Source: College Board AP Course and Exam Description · ⁨出典: College Board AP コースおよび試験説明書⁩

English

For a two-way table, the expected count in a cell (under "no association") is

$$E=\frac{(\text{row total})\times(\text{column total})}{\text{grand total}}.$$
This is the count you would see if the row and column variables were unrelated.

Worked example. In a two-way table a cell's row total is $40$, its column total is $50$, and the grand total is $200$. Its expected count is $E=\dfrac{40\times50}{200}=10$. Repeating for every cell gives the expected table to compare against the observed one.

日本語

2次元表において、「関連なし」の下でのセルの期待カウントは

$$E=\frac{(\text{row total})\times(\text{column total})}{\text{grand total}}.$$
これは行変数と列変数が無関係であった場合に観察されるカウントです。

** worked example.** 2次元表であるセルの行合計が$40$、列合計が$50$、全体合計が$200$の場合、その期待カウントは$E=\dfrac{40\times50}{200}=10$です。すべてのセルについて繰り返すことで、観測表と比較するための期待表が得られます。

スプレッドシートはカイ二乗検定前にカテゴリカルなカウントを整理します
スプレッドシートはカイ二乗検定前にカテゴリカルなカウントを整理します
8.5

Homogeneity or Independence? · ⁨均質性か独立性か?⁩

Syllabus · ⁨シラバス⁩
English

Enduring Understanding (VAR-8): The chi-square distribution may be used to model variation.

Learning Objective VAR-8.I: Identify the null and alternative hypotheses for a chi-square test for homogeneity or independence. [Skill 1.F]

  • VAR-8.I.1 The appropriate hypotheses for a chi-square test for homogeneity are:

    $H_0$: There is no difference in distributions of a categorical variable across populations or treatments.

    $H_a$: There is a difference in distributions of a categorical variable across populations or treatments.

  • VAR-8.I.2 The appropriate hypotheses for a chi-square test for independence are:

    $H_0$: There is no association between two categorical variables in a given population or the two categorical variables are independent.

    $H_a$: Two categorical variables in a population are associated or dependent.

Learning Objective VAR-8.J: Identify an appropriate testing method for comparing distributions in two-way tables of categorical data. [Skill 1.E]

  • VAR-8.J.1 When comparing distributions to determine whether proportions in each category for categorical data collected from different populations are the same, the appropriate test is the chi-square test for homogeneity.
  • VAR-8.J.2 To determine whether row and column variables in a two-way table of categorical data might be associated in the population from which the data were sampled, the appropriate test is the chi-square test for independence.

Learning Objective VAR-8.K: Verify the conditions for making statistical inferences when testing a chi-square distribution for independence or homogeneity. [Skill 4.C]

  • VAR-8.K.1 In order to make statistical inferences for a chi-square test for two-way tables (homogeneity or independence), we must verify the following:
    • a. To check for independence:
      • i. For a test for independence: Data should be collected using a simple random sample.
      • ii. For a test for homogeneity: Data should be collected using a stratified random sample or randomized experiment.
      • iii. When sampling without replacement, check that $n \leq 10\%N$.
    • b. The chi-square tests for independence and homogeneity become more accurate with more observations, so large counts should be used (shape).
      • i. A conservative check for large counts is that all expected counts should be greater than 5.
日本語

持続的理解 (VAR-8): カイ二乗分布は、ばらつきのモデル化に使用できる。

学習目標 VAR-8.I: カイ二乗の同質性検定または独立性検定における帰無仮説と対立仮説を特定する。[スキル 1.F]

  • VAR-8.I.1 カイ二乗の同質性検定に適切な仮説は以下の通りである:

    $H_0$: 人口群や処理間でのカテゴリー変数の分布に違いがない。

    $H_a$: 人口群や処理間でのカテゴリー変数の分布に違いがある。

  • VAR-8.I.2 カイ二乗の独立性検定に適切な仮説は以下の通りである:

    $H_0$: 特定の人口群における2つのカテゴリー変数間に相関がなく、あるいは2つのカテゴリー変数が独立している。

    $H_a$: 人口群内の2つのカテゴリー変数間に相関があり、あるいは依存関係にある。

学習目標 VAR-8.J: 2次元表形式のカテゴリカルデータの分布比較に適した検定手法を特定する。[スキル 1.E]

  • VAR-8.J.1 異なる人口群から収集されたカテゴリカルデータのカテゴリごとの比率が同一かどうかを確認して分布を比較する場合、適切な検定はカイ二乗の同質性検定である。
  • VAR-8.J.2 データがサンプリングされた人口群において、2次元表の行変数と列変数が相関関係にある可能性を確認する場合、適切な検定はカイ二乗の独立性検定である。

学習目標 VAR-8.K: カイ二乗の独立性検定または同質性検定を行う際の統計的推論のための条件を確認する。[スキル 4.C]

  • VAR-8.K.1 2次元表(同質性または独立性)のカイ二乗検定に対する統計的推論を行うには、以下の条件を確認する必要がある:
    • a. 独立性を確認するために:
      • i. 独立性検定の場合:データは単純無作為抽出によって収集されるべきである。
      • ii. 同質性検定の場合:データは層別無作為抽出または無作為化実験によって収集されるべきである。
      • iii. 無放回抽出を行う場合は、 $n \leq 10\%N$ を確認する。
    • b. 独立性および同質性のカイ二乗検定は観測値が増えるほど精度が高まるため、大きな期待度数を使用する(形状)。
      • i. 大きな度数に対する保守的なチェックとしては、すべての期待度数が5より大きいことが挙げられる。

Source: College Board AP Course and Exam Description · ⁨出典: College Board AP コースおよび試験説明書⁩

English

Two tests use the same $\chi^2$ math but answer different questions:

  • Test for homogeneity 同质性: are the distributions of one categorical variable the same across several populations or groups (separate samples/treatments)?
  • Test for independence 独立性: are two categorical variables associated within a single population (one sample, two variables measured)?

The design (several samples vs one sample) decides which name and hypotheses to use.

日本語

2つの検定は同じ$\chi^2$数学を使いますが、異なる質問に答えます:

  • 均質性の検定: 1つのカテゴリ変数の分布が複数の集団やグループで同一か(分離された標本/処理)?
  • 独立性の検定: 2つのカテゴリ変数が単一の集団内で関連しているか(1つの標本、2つの変数を測定)?

設計(複数標本 vs 1つ)がどの名前と仮説を使うかを決定します。

8.6

Carrying Out a Test for Homogeneity or Independence · ⁨均質性または独立性の検定の実施⁩

Syllabus · ⁨シラバス⁩
English

Enduring Understanding (VAR-8): The chi-square distribution may be used to model variation.

Learning Objective VAR-8.L: Calculate the appropriate statistic for a chi-square test for homogeneity or independence. [Skill 3.E]

  • VAR-8.L.1 The appropriate test statistic for a chi-square test for homogeneity or independence is the chi-square statistic:
    • Equation: $\chi^2 = \sum \dfrac{(Observed\ count - Expected\ count)^2}{Expected\ count}$, with degrees of freedom equal to: $(number\ of\ rows - 1)(number\ of\ columns - 1)$.

Learning Objective VAR-8.M: Determine the $p$-value for a chi-square significance test for independence or homogeneity. [Skill 3.E]

  • VAR-8.M.1 The $p$-value for a chi-square test for independence or homogeneity for a number of degrees of freedom is found using the appropriate table or technology.
  • VAR-8.M.2 For a test of independence or homogeneity for a two-way table, the $p$-value is the proportion of values in a chi-square distribution with appropriate degrees of freedom that are equal to or larger than the test statistic.

Enduring Understanding (DAT-3): Significance testing allows us to make decisions about hypotheses within a particular context.

Learning Objective DAT-3.K: Interpret the $p$-value for the chi-square test for homogeneity or independence. [Skill 4.B]

  • DAT-3.K.1 An interpretation of the $p$-value for the chi-square test for homogeneity or independence is the probability, given the null hypothesis and probability model are true, of obtaining a test statistic as, or more, extreme than the observed value.

Learning Objective DAT-3.L: Justify a claim about the population based on the results of a chi-square test for homogeneity or independence. [Skill 4.E]

  • DAT-3.L.1 A decision to either reject or fail to reject the null hypothesis for a chi-square test for homogeneity or independence is based on comparison of the $p$-value to the significance level, $\alpha$.
  • DAT-3.L.2 The results of a chi-square test for homogeneity or independence can serve as the statistical reasoning to support the answer to a research question about the population that was sampled (independence) or the populations that were sampled (homogeneity).
日本語

持続的理解 (VAR-8): カイ二乗分布は、ばらつきのモデル化に使用できる。

学習目標 VAR-8.L: カイ二乗の同質性検定または独立性検定に適切な検定 statistic を計算する。[スキル 3.E]

  • VAR-8.L.1 カイ二乗の同質性検定または独立性検定に適切な検定 statistic はカイ二乗 statistic である:
    • 数式: $\chi^2 = \sum \dfrac{(Observed\ count - Expected\ count)^2}{Expected\ count}$、自由度は次の通り: $(number\ of\ rows - 1)(number\ of\ columns - 1)$。

学習目標 VAR-8.M: カイ二乗の独立性検定または同質性検定に対する $p$ 値を求める。[スキル 3.E]

  • VAR-8.M.1 自由度の数に対するカイ二乗の独立性検定または同質性検定の $p$ 値は、適切な表または技術工具を用いて求める。
  • VAR-8.M.2 2次元表に対する独立性検定または同質性検定では、 $p$ 値は適切な自由度を持つカイ二乗分布における値のうち、検定 statistic に等しいかそれより大きいものの割合である。

持続的理解 (DAT-3): 有意性検定により、特定の文脈内で仮説に関する判断を下すことができます。

学習目標 DAT-3.K: カイ二乗の同質性検定または独立性検定に対する $p$ 値を解釈する。[スキル 4.B]

  • DAT-3.K.1 カイ二乗の同質性検定または独立性検定に対する $p$ 値の解釈は、帰無仮説と確率モデルが正しいと仮定した場合、観測値と同程度以上極端な検定 statistic が得られる確率である。

学習目標 DAT-3.L: カイ二乗の同質性検定または独立性検定の結果に基づき、人口群に関する主張を正当化する。[スキル 4.E]

  • DAT-3.L.1 カイ二乗の同質性検定または独立性検定における帰無仮説の棄却または棄却不能の決定は、 $p$ 値と有意水準 $\alpha$ の比較に基づいて行われる。
  • DAT-3.L.2 カイ二乗の同質性検定または独立性検定の結果は、サンプリングされた人口群(独立性)またはサンプリングされた各人口群(同質性)に関する研究質問への回答を裏付けるための統計的推論として機能しうる。

Source: College Board AP Course and Exam Description · ⁨出典: College Board AP コースおよび試験説明書⁩

English

Compute expected counts, then $\chi^2=\sum\dfrac{(O-E)^2}{E}$ over all cells, with

$$df=(\text{rows}-1)(\text{columns}-1).$$
Conditions: random data, all expected counts $\ge 5$, 10% condition. Find the $p$-value, compare to $\alpha$, and conclude in context – evidence of a difference between groups (homogeneity) or of an association (independence).

日本語

期待カウントを計算し、すべてのセルについて$\chi^2=\sum\dfrac{(O-E)^2}{E}$を合計し、

$$df=(\text{rows}-1)(\text{columns}-1).$$
条件:ランダムデータ、すべての期待個数が$\ge 5$以上、10%条件。$p$-値を見つけ、$\alpha$と比較して文脈の中で結論付けます – グループ間の違い(同質性)または相関(独立)の証拠。

8.7

Choosing the Right Categorical Procedure · ⁨適切なカテゴリカル手続の選択⁩

Syllabus · ⁨シラバス⁩
English

This topic is intended to focus on the skill of selecting an appropriate inference procedure now that students have a range of options. Students should be given opportunities to practice when and how to apply all learning objectives relating to inference for categorical data.

日本語

このトピックは、学生が多様な選択肢を有するようになった今、適切な推論手順を選択するスキルに焦点を当てることを意図している。学生は、カテゴリカルデータに関するすべての学習目標(推論)に対する適用時期と方法を練習する機会を得るべきである。

Source: College Board AP Course and Exam Description · ⁨出典: College Board AP コースおよび試験説明書⁩

English

Decide by the setup: one categorical variable against a claimed distribution $\Rightarrow$ goodness-of-fit; one sample cross-classified by two variables $\Rightarrow$ independence; several samples/groups compared $\Rightarrow$ homogeneity. Comparing just two proportions can use either a two-proportion $z$-test or a chi-square test, but only for a two-tailed alternative, where they agree exactly ($\chi^2=z^2$). A chi-square test is always two-tailed, so it cannot give a directional conclusion: if $H_a$ is one-tailed (say $p_1>p_2$), use the $z$-test.

日本語

設定によって判断します:1つのカテゴリ変数と主張された分布の比較 $\Rightarrow$ 適合度検定;1つのサンプルが2つの変数で交差分類される場合 $\Rightarrow$ 独立性検定;複数のサンプル/グループを比較する場合 $\Rightarrow$ 同質性検定。単に2つの比率を比較する際は、2比率 $z$-検定またはカイ二乗検定のどちらかを使用できますが、それらは完全に一致するのは両側対立仮説の場合のみです($\chi^2=z^2$)。カイ二乗検定は常に両側であるため、方向性の結論を出すことはできません。もし$H_a$が片側(例えば$p_1>p_2$)である場合は、$z$-検定を使用してください。

Explore · ⁨探索⁩

Which chi-square test is this? · ⁨これはどのカイ二乗検定か?⁩

All three tests use the same $\chi^2$ arithmetic, so the marks are won by naming the right one. The design decides — how many samples were taken, and how many variables were measured on each unit. · ⁨3つの検定すべて同じ $\chi^2$ の計算式を使用するため、正解を名指すことで得点されます。決めるのは設計です。いくつのサンプルを取り、各単位についていくつの変数を測定したかです。⁩

8.7

Exam tips · ⁨試験対策⁩

English
  • Use $\chi^2=\sum\tfrac{(O-E)^2}{E}$ for categorical data; always divide by the expected count.
  • Pick the right test: goodness-of-fit (one variable), independence, or homogeneity (two-way table).
  • Compute expected counts as $\tfrac{\text{row total}\times\text{column total}}{\text{grand total}}$ and check each is $\ge5$.
  • A large $\chi^2$ (small p-value) means observed counts differ from expected by more than chance.
  • State degrees of freedom correctly (categories $-1$, or $(r-1)(c-1)$).
日本語
  • カテゴリカルデータには$\chi^2=\sum\tfrac{(O-E)^2}{E}$を使用し、常に期待度数で割ります。
  • 適切な検定を選択します:適合度検定(1変数)、独立性、または同質性(2次表)。
  • 期待度数を$\tfrac{\text{row total}\times\text{column total}}{\text{grand total}}$として計算し、それぞれが$\ge5$以上であることを確認します。
  • 大きな$\chi^2$(小さなp値)は、観測度数と期待度数の差が偶然によるものより大きいことを示します。
  • 自由度を正しく述べる(カテゴリ$-1$、または$(r-1)(c-1)$)。

Interactive lessons on this topic · ⁨このトピックのインタラクティブ授業⁩

Work through it step by step, with instant-check exercises. · ⁨一歩ずつ進め、即時チェック付きの問題で学習します。⁩

Past Papers · ⁨過去問⁩

More topics in AP Statistics · ⁨AP Statistics の他のトピック⁩

Log in or create account · ⁨ログインまたはアカウント作成⁩

IGCSE, A-Level & AP