What a t-test answers
It asks one narrow question: if the two groups really had the same mean, how often would random sampling produce a difference at least as big as the one you measured? That is the p-value. It is not the probability that your hypothesis is true.
Picking the right one
- Paired: the same subjects measured twice, before and after. Test the differences, which removes the variation between subjects and makes the test much more sensitive. Use this whenever the data really is paired.
- Welch's two-sample: two independent groups, without assuming the variances match. This is the sensible default for independent groups.
- Student's two-sample: the same, pooling the variances. Only worth it when the groups are similar in size and spread, and it gains very little over Welch.
- One-sample: one group against a fixed reference value.
A worked example
Group A: 5.1, 5.3, 4.9, 5.0, 5.2. Group B: 4.5, 4.7, 4.3, 4.4, 4.6.
- Means 5.10 and 4.50, each with a standard deviation of 0.158.
- Standard error of the difference: √(0.025/5 + 0.025/5) = 0.1.
- t = 0.60 ÷ 0.1 = 6.00, with 8 degrees of freedom.
- Two-tailed p = 0.00032.
A difference that large between five-point groups this tight would turn up by chance about three times in ten thousand.
One tail or two
Use two-tailed unless a difference in one direction would be meaningless to you, and decide before seeing the data. Switching to one-tailed afterwards to get under 0.05 halves the p-value for free and is the most common way these tests are abused.
What the test assumes
- Independent observations. This is the assumption that gets broken most often, and no test can rescue it. Repeated measurements of the same subject are not independent.
- Roughly normal data, or enough of it. Above about 30 per group the test is robust to non-normality; below that, strong skew or outliers matter.
- No equal-variance assumption, if you use Welch.
Clearly skewed small samples are better served by Mann-Whitney or Wilcoxon.
Report the effect size too
A p-value says nothing about how big the difference is. With large samples a trivial difference is significant; with small ones a large difference may not be. Give the difference in means with a confidence interval, and Cohen's d (difference ÷ pooled standard deviation, where 0.2 is small, 0.5 medium and 0.8 large). In the example above d is 3.8, which is enormous.
Tools
The T-Test Calculator runs paired, two-sample and Welch tests, and gives t, the degrees of freedom and the p-value. The Standard Deviation Calculator gives the mean and spread of each group on its own.