Carrying Out a Two-Mean Test · Realización de una Prueba para Dos Medias
A rounded p-value cannot settle a boundary
- A display reports p=0.05 to two decimal places. The underlying result might be 0.047 or 0.053, giving different decisions at alpha=0.05.
- Use enough precision before comparing with alpha. A rounded boundary value is not necessarily exact equality; t=2 alone is also insufficient without degrees of freedom and the alternative.
Find the p-value
- Convert the two-sample $t$ statistic to a p-value from the $t$-distribution. One-sided $H_a$: the tail area in that direction; two-sided: double the smaller tail area for the symmetric null distribution.
- The · El $df$ come from technology (a calculator's two-sample $t$ routine). Under the null model, it is the probability of a test statistic at least as extreme in the direction or directions specified by the alternative.
The paired t-test
- If the design is matched pairs, run a paired $t$-test on the differences. Compute each $d = x_1 - x_2$, then test $H_0: \mu_d = 0$ with a one-sample $t$ on the $d$'s.
- Test statistic: $t = \dfrac{\bar{d} - 0}{s_d/\sqrt{n}}$, with $df = n - 1$ ($n$ = number of pairs). It's just the one-sample $t$-test applied to the differences.
A paired t-test has twelve complete pairs. What are its degrees of freedom?
The analysis has twelve differences, so df = 12 - 1 = 11, not 24 - 1.
Make a decision
- Compare the p-value to $\alpha$: p $\le \alpha$ → reject $H_0$: convincing evidence the means differ.
- p $> \alpha$ → fail to reject: not convincing evidence. Same rule for both the two-sample and paired tests.
For a matched-pairs design, the correct test is...
Collapse pairs to differences, then a one-sample (paired) t-test.
With an unrounded two-sided p-value exactly 0.05 and alpha exactly 0.05, which decision follows the stated p ≤ alpha rule?
Under the stated exact boundary convention, reject. A rounded display of 0.05 could instead represent a value above or below alpha.
The paired t-test statistic is t = d-bar / (s_d/√n) with df = n − 1.
It's a one-sample t-test on the paired differences.
Running a two-sample test on paired data gives the correct p-value if the arithmetic is right.
Wrong procedure → wrong p-value, even with correct arithmetic.
A paired-test conclusion is phrased about...
Paired tests are about the mean of the differences.
Conclude in context
- State the conclusion about the two population means, in context. "Convincing evidence the two group means differ" (if rejected), or "not convincing" (if not).
- Name the groups and what the mean measures, with units. For a paired test, phrase it about the mean difference.
Pick the test from the design, then compute. Matched pairs → paired $t$-test on differences ($df = n-1$, one SE from $s_d$). Independent groups → two-sample $t$-test (added variances, technology $df$). Running a two-sample test on paired data — or vice versa — gives the wrong p-value even when every number is entered correctly.
A valid two-sided test reports p = 0.053 at alpha = 0.05. Which should a report include?
Non-rejection is a threshold decision, not proof of no difference. Report the estimate and how the evidence was obtained.
Suppose technology reports a two-sided p-value of $0.0473$ for a valid two-sample test; $\alpha=0.05$.
- Compare: $0.0473<0.05$ → reject $H_0$. Do not decide from a two-decimal display of $0.05$.
- Conclude: convincing evidence the two group means differ.
- A paired version would instead test $H_0: \mu_d = 0$ on the differences.
Carry the reasoning to a new case
- For independent samples use the two-sample procedure; for matched measurements analyse differences.
- Keep the design and alternative fixed before the significance decision.
Match each unrounded p-value to the decision at alpha=0.05.
Use the unrounded value or sufficient precision. Failure to reject does not prove equality.
Find the p-value from the $t$ statistic (twice the smaller tail area for a two-sided test) and compare to $\alpha$: reject if p $\le \alpha$, else fail to reject. Use a two-sample $t$-test for independent groups and a paired $t$-test ($t = \frac{\bar{d}}{s_d/\sqrt{n}}$) for matched pairs. State the conclusion about the two population means in context.