Sampling Distributions for Differences in Sample Means (College Board AP® Statistics): Study Guide

Mark Curtis

Written by: Mark Curtis

Reviewed by: Dan Finlay

Updated on

Sampling distributions for differences in sample means

What is a one-sample problem?

  • When one random sample of size n has been taken from one population

    • with population mean μ and population standard deviation σ

      • The sample mean is x¯

      • This is a one-sample problem

What is a two-sample problem?

  • If one random sample of size n1 is taken from one population with population mean μ1 and population standard deviation σ1

    • then a different random sample of size n2is taken from a different population (that is independent to the first population) with population mean μ2 and population standard deviation σ2

      • then this is a two-sample problem

      • The sample means are x¯1 and x¯2

What is the difference in sample means?

  • In a two-sample problem you can compare the sample means from separate samples of two independent populations

    • You can look at the difference in sample means, x¯1x¯2

      • e.g. if x¯1x¯2>0 then the mean of the first sample is greater than the mean of the second sample

What is the sampling distribution for differences in sample means?

  • You can find the differences in sample means, if

    • you take all possible samples of size n1 from the first population and calculate their sample means, x¯1

    • then take all possible samples of size n2 from the second population and calculate their sample means, x¯2

    • then work out all the possible values that the difference x¯1x¯2 can take

      • The collection of all these values is called the sampling distribution for differences in sample means

What are the mean and standard deviation of the sampling distribution for differences in sample means?

  • If the first population has a population mean of μ1 and a population standard deviation of σ1

    • and the second independent population has a population mean of μ2 and population standard deviation of σ2

  • Then the sampling distribution for differences in sample means, x¯1x¯2

    • has a mean of μ1μ2

    • and a standard deviation of σ12n1+σ22n2

    • where n1 is the size of the first sample

    • and n2 is the size of the second sample

  • The standard deviation of σ12n1+σ22n2 assumes sampling was done with replacement

  • If the sampling is carried out without replacement then the sample must satisy:

    • The randomization conditon

      • a random sampling method should be used for both samples

    • The 10% condition

      • each sample size is less than 10% of the population size

      • n110%N1 and n210%N2

Examiner Tips and Tricks

The mean, μ1μ2, and the standard deviation, σ12n1+σ22n2, are given in the exam under 'Sampling distributions for means', in the row called 'For two populations'.

What conditions are needed for normality?

  • If in addition to the above, the two independent populations are also known to be normally distributed

    • then the sampling distribution for differences in sample means is also normally distributed

      • with mean μ1μ2 and standard deviation σ12n1+σ22n2

  • You can use these properties to calculate probabilities involving differences in sample means, x¯1x¯2, as they follow a normal distribution

    • Its standardized z-statistic is (x1x2)(μ1μ2)σ12n1+σ22n2

      • μ1, μ2, σ1 and σ2 will be given in the question

What do I do if the populations are not normally distributed?

  • If the populations are not normally distributed, then the sampling distribution for differences in sample means is not guaranteed to be normally distributed

  • However, despite not knowing its shape, the sampling distribution for differences in sample means still has a

    • mean of μ1μ2 and a standard deviation of σ12n1+σ22n2

      • i.e. you can always write these down, even though the distribution is unknown

Can I use the Central Limit theorem if the populations are not normally distributed?

  • If the populations are not normally distributed, but both sample sizes are greater than or equal to 30 (n130 and n230)

    • then the Central Limit theorem can be applied

    • meaning the sampling distribution for differences in sample means is approximately normally distributed with the parameters above

      • i.e. mean μ1μ2 and standard deviation σ12n1+σ22n2

  • You can use these properties to estimate probabilities involving differences in sample means, x¯1x¯2, as they follow an approximate normal distribution

    • Its standardized z-statistic is (x1x2)(μ1μ2)σ12n1+σ22n2

      • μ1, μ2, σ1 and σ2 will be given in the question

Worked Example

The average lifetime of bulbs from a company called Brite have a mean of 900 hours and a standard deviation of 25 hours. The average lifetime of bulbs from a company called Shine have a mean of 800 hours and a standard deviation of 15 hours.

Estimate the probability that the mean of a sample of 40 bulbs from Brite is at least 108 hours more than the mean of a sample of 50 bulbs from Shine.

Answer:

This a probability question about the difference in means of two samples, so requires the sampling distribution for differences in sample means

Start by labeling each population

Population 1 is the average lifetime of bulbs from Brite

Population 2 is the average lifetime of bulbs from Shine

You are not told the lifetimes of the bulbs are normally distributed but both sample sizes are greater than 30 so the Central Limit theorem can be applied

n1=4030 and n2=5030 so use the Central Limit theorem

The difference in sample means follows an approximate normal distribution with mean μ1μ2 and standard deviation σ12n1+σ22n2

Substitute μ1=900 and μ2=800 into μ1μ2

μ1μ2=900800=100

Substitute σ1=25, n1=40, σ2=15 and n2=50 into σ12n1+σ22n2

σ12n1+σ22n2=25240+15250=4.4860896...

The wording in the question asks for the probability that x¯1>x¯2+108

Rearrange this to form the difference of sample means, x¯1x¯2

P(X¯1>X¯2+108)=P(X¯1X¯2>108)

The difference in sample means follows an approximate normal distribution with mean 100 and standard deviation 4.4860896... from above

To find the probability that the difference in sample means is greater than 108, first calculate the z-score for 108

1081004.4860896...=1.783...

Then find P(Z>1.783...), e.g. using the normal tables

P(Z>1.783...)=1P(Z<1.783...)=10.9625=0.0375

The probability that the mean of a sample of 40 bulbs from Brite is at least 108 hours more than the mean of a sample of 50 bulbs from Shine is approximately 0.0375

Unlock more, it's free!

Join the 100,000+ Students that ❤️ Save My Exams

the (exam) results speak for themselves:

Build on this topic

Mark Curtis

Author: Mark Curtis

Expertise: Maths Content Creator

Mark graduated twice from the University of Oxford: once in 2009 with a First in Mathematics, then again in 2013 with a PhD (DPhil) in Mathematics. He has had nine successful years as a secondary school teacher, specialising in A-Level Further Maths and running extension classes for Oxbridge Maths applicants. Alongside his teaching, he has written five internal textbooks, introduced new spiralling school curriculums and trained other Maths teachers through outreach programmes.

Dan Finlay

Reviewer: Dan Finlay

Expertise: Portfolio Lead

Dan graduated from the University of Oxford with a First class degree in mathematics. As well as teaching maths for over 8 years, Dan has marked a range of exams for Edexcel, tutored students and taught A Level Accounting. Dan has a keen interest in statistics and probability and their real-life applications.