Skip to main content

The sampling distribution of the mean and the central limit theorem: the mean and variance of a sample mean, and probabilities that it lies within given bounds (new in the 2024 Extension 1 syllabus)

Syllabus dot point

“Apply the central limit theorem to estimate the probability that the sample mean lies within given bounds”

HSCMaths Extension 1Statistical Analysis (ME-S1)13 min read

Quick answer

Sample means vary from sample to sample, but E(Xˉ)=μE(\bar{X}) = \mu and Var⁡(Xˉ)=σ2n\operatorname{Var}(\bar{X}) = \frac{\sigma^2}{n}. By the central limit theorem, for n≥30n \geq 30 the sample mean is approximately N ⁣(μ,σ2n)N\!\left(\mu, \frac{\sigma^2}{n}\right) whatever the population's shape, so probabilities about Xˉ\bar{X} come from z=xˉ−μσ/nz = \frac{\bar{x} - \mu}{\sigma/\sqrt{n}} and the standard normal table. New in the 2024 Extension 1 syllabus.

Jump to a section
  1. What this dot point is asking
  2. The answer
  3. Exam-style questions
  4. Practice questions

What this dot point is asking

This is new content in the Mathematics Extension 1 11-12 Syllabus (2024), first examined in the 2027 HSC. It replaces the 2017 course's normal approximation for the sample proportion. NESA's Year 12 focus area "The binomial distribution and the sampling distribution of the mean" asks you to:

  • recognise that sample means from repeated samples differ, even for the same sample size, and that Xˉ\bar{X} is a random variable that estimates μ\mu
  • use E(Xˉ)=μE(\bar{X}) = \mu and Var⁡(Xˉ)=σ2n\operatorname{Var}(\bar{X}) = \frac{\sigma^2}{n}
  • state the central limit theorem: for a population with mean μ\mu and variance σ2\sigma^2, provided nn is large enough (n≥30n \geq 30), Xˉ\bar{X} is approximately N ⁣(μ,σ2n)N\!\left( \mu, \frac{\sigma^2}{n} \right), whatever the shape of the population
  • apply the central limit theorem to estimate the probability that the sample mean lies within given bounds, which is the exam skill.
Note

Imagine asking one random student how long they spent on homework last night: the answer could be anything from zero to four hours. Now ask fifty random students and average their answers. That average is far more predictable, because the very long and very short answers cancel out. Do it again with another fifty students and you get a slightly different average, but close to the first. The central limit theorem says those averages always pile up in a bell shape around the true average, and the bigger the group, the narrower the bell.

The answer

Key fact

For random samples of size nn from a population with mean μ\mu and standard deviation σ\sigma:

E(Xˉ)=μ,Var⁡(Xˉ)=σ2n,standard deviation of Xˉ=σn.E(\bar{X}) = \mu, \qquad \operatorname{Var}(\bar{X}) = \frac{\sigma^2}{n}, \qquad \text{standard deviation of } \bar{X} = \frac{\sigma}{\sqrt{n}}.

Central limit theorem: if n≥30n \geq 30, then Xˉ\bar{X} is approximately N ⁣(μ,σ2n)N\!\left( \mu, \frac{\sigma^2}{n} \right), even when the population is not normal. So

z=xˉ−μσ/n.z = \frac{\bar{x} - \mu}{\sigma / \sqrt{n}}.

Why sample means behave this way

A single observation XX can land anywhere in the population's spread. The sample mean averages nn independent observations, so extreme values tend to cancel. Its centre stays at μ\mu (on average a sample neither overestimates nor underestimates), but its spread shrinks: the variance is divided by nn, so the standard deviation is divided by n\sqrt{n}. Quadrupling the sample size halves the spread.

The remarkable part is the shape. Even if the population is skewed, the distribution of Xˉ\bar{X} becomes approximately normal as nn grows. NESA takes n≥30n \geq 30 as large enough. That is what lets you use the standard normal table for questions about averages of non-normal quantities such as waiting times, incomes or dice scores.

Population distribution and the sampling distribution of the meanA strongly right-skewed population of waiting times with mean 12 minutes and standard deviation 12 minutes, drawn as a decreasing curve starting high at 0. Overlaid is the sampling distribution of the mean for samples of size 36: a tall, narrow, symmetric bell curve centred on the same mean of 12 with standard deviation 12 divided by the square root of 36, which is 2. The population is skewed but the sample means are approximately normal and far less spread out. 010203040 waiting time (minutes) μ = 12 sample means, n = 36 approx. normal, SD 12/√36 = 2 population (skewed) mean 12, SD 12 Same centre, far less spread: X̄ is approximately N(μ, σ²/n) for n ≥ 30.

The method for probability questions

  1. Check the conditions. Say that n≥30n \geq 30, so the central limit theorem applies.
  2. Find the parameters of Xˉ\bar{X}. Mean μ\mu, standard deviation σn\frac{\sigma}{\sqrt{n}}. Use the population's standard deviation, not the variance, in the numerator.
  3. Standardise each bound. z=xˉ−μσ/nz = \frac{\bar{x} - \mu}{\sigma/\sqrt{n}}.
  4. Read the standard normal table and use symmetry: P(Z>z)=1−P(Z≤z)P(Z > z) = 1 - P(Z \leq z) and P(Z<−z)=1−P(Z≤z)P(Z < -z) = 1 - P(Z \leq z).
  5. Answer in context, usually as a decimal to four places or a percentage.

Sample mean versus a single value

The most common error in these questions is using σ\sigma instead of σn\frac{\sigma}{\sqrt{n}}. Ask yourself whether the question is about one individual (use σ\sigma, and only if the population is normal) or about the average of a sample (use σn\frac{\sigma}{\sqrt{n}} and the central limit theorem). The same value is far less unusual for one person than for the average of fifty.

How the new content connects to the old

The 2017 Extension 1 course used the normal approximation for the sample proportion p^\hat{p}. The 2024 course keeps Bernoulli and binomial distributions (explicitly excluding the normal approximation to the binomial) and moves to the sample mean. The standardising step is the same idea you met with zz-scores in Mathematics Advanced; our pages on sample proportions and the normal approximation of the binomial show the older, related method.

Worked examples

Probability below a bound

Battery lifetimes have mean 400400 hours and standard deviation 6060 hours. For a random sample of 3636 batteries, find P(Xˉ<385)P(\bar{X} < 385).

Conditions and parameters
n=36≥30n = 36 \geq 30, so Xˉ≈N(400,102)\bar{X} \approx N(400, 10^2) since 6036=10\frac{60}{\sqrt{36}} = 10.
Standardise
z=385−40010=−1.5z = \frac{385 - 400}{10} = -1.5.
Table and symmetry
P(Z<−1.5)=1−P(Z≤1.5)=1−0.9332=0.0668P(Z < -1.5) = 1 - P(Z \leq 1.5) = 1 - 0.9332 = 0.0668.

Marker's note: one mark each for the standard deviation of Xˉ\bar{X}, the zz-score and the probability.

Probability between two bounds

Bags of rice have mean mass 1.021.02 kg and standard deviation 0.050.05 kg. For a random sample of 100100 bags, find the probability that the mean mass is between 1.011.01 kg and 1.031.03 kg.

Parameters
0.05100=0.005\frac{0.05}{\sqrt{100}} = 0.005.
Standardise
z=1.01−1.020.005=−2z = \frac{1.01 - 1.02}{0.005} = -2 and z=1.03−1.020.005=2z = \frac{1.03 - 1.02}{0.005} = 2.
Probability
P(−2<Z<2)=2P(Z≤2)−1=2(0.9772)−1=0.9544P(-2 < Z < 2) = 2P(Z \leq 2) - 1 = 2(0.9772) - 1 = 0.9544, close to the empirical rule's 95%95\%.

Marker's note: the symmetric interval makes 2P(Z≤2)−12P(Z \leq 2) - 1 the quickest route; showing both zz-scores earns the method mark.

Choosing a sample size

How large must a sample be for the standard deviation of Xˉ\bar{X} to be at most 22, when σ=15\sigma = 15?

Set up. 15n≤2\frac{15}{\sqrt{n}} \leq 2 gives n≥7.5\sqrt{n} \geq 7.5, so n≥56.25n \geq 56.25.

Round up. The smallest sample size is n=57n = 57 (which also satisfies n≥30n \geq 30).

Marker's note: sample sizes are whole numbers and must be rounded up to meet the condition.

Common traps
Using σ\sigma instead of σn\frac{\sigma}{\sqrt{n}}
A question about an average needs the standard deviation of Xˉ\bar{X}.
Dividing by nn instead of n\sqrt{n}
Var⁡(Xˉ)=σ2n\operatorname{Var}(\bar{X}) = \frac{\sigma^2}{n}, so the standard deviation is σn\frac{\sigma}{\sqrt{n}}, not σn\frac{\sigma}{n}.
Forgetting to justify the central limit theorem
State n≥30n \geq 30; markers award a mark for it.
Assuming the population must be normal
The whole point of the theorem is that it need not be, once nn is large enough.
Rounding a sample size down
If n≥61.47n \geq 61.47, the answer is 6262.
Exam technique

Write one line of justification ("since n=50≥30n = 50 \geq 30, by the central limit theorem Xˉ\bar{X} is approximately normal") before any calculation: it is often worth a mark on its own. Keep σn\frac{\sigma}{\sqrt{n}} exact (for example 6.550\frac{6.5}{\sqrt{50}}) until you compute zz, round zz to two decimal places to match the table, and sketch a normal curve with the region shaded to check whether you need P(Z≤z)P(Z \leq z) or 1−P(Z≤z)1 - P(Z \leq z).

Exam-style questions

Questions in the style of NESA exam questions on this dot point, each with a worked answer. They are written by ExamExplained unless tagged "Past paper"; the year shows the paper a question is modelled on.

NESA 2024-syllabus samplePast paper3 marks
In a large school, the average amount of money spent per student per day at the canteen is &#36;8 with a standard deviation of 6.5. At the end of each day, 50 randomly chosen students are asked how much they spent at the canteen on that day. Use the standard normal distribution to find the probability that the sample mean on a particular day is greater than &#36;10. You may use the information provided on page 16. [Page 16 of the sample paper is a table of standard normal probabilities.]
Show worked answer →

Because the sample size is 50≥3050 \geq 30, the central limit theorem applies: the mean Xˉ\bar{X} of the 5050 amounts is approximately normal with mean μ=8\mu = 8 and standard deviation 6.550≈0.919\frac{6.5}{\sqrt{50}} \approx 0.919.

Standardise: z=10−86.5/50≈2.18z = \frac{10 - 8}{6.5/\sqrt{50}} \approx 2.18 (to two decimal places).

From the table, P(Z≤2.18)=0.9854P(Z \leq 2.18) = 0.9854, so P(Xˉ>10)≈1−0.9854=0.0146P(\bar{X} > 10) \approx 1 - 0.9854 = 0.0146, about 1.46%1.46\%.

NESA's marking guidelines award 3 marks for the correct solution; 2 marks for recognising the standard normal distribution and finding the zz-score (or equivalent merit); 1 mark for recognising that n≥30n \geq 30 allows use of the central limit theorem, or for finding the standard deviation of Xˉ\bar{X} in terms of nn (or equivalent merit).

Source: NESA HSC Mathematics Extension 1 annotated sample examination materials (Mathematics Extension 1 11-12 Syllabus (2024)), Question 14(c), and marking guidelines.

HSC-style3 marks
Battery lifetimes have mean 400400 hours and standard deviation 6060 hours. A random sample of 3636 batteries is tested. Find the probability that the sample mean lifetime is less than 385385 hours. (Use P(Z≤1.5)=0.9332P(Z \leq 1.5) = 0.9332.)
Show worked answer →

With n=36≥30n = 36 \geq 30, the central limit theorem gives Xˉ\bar{X} approximately normal with mean 400400 and standard deviation 6036=10\frac{60}{\sqrt{36}} = 10.

z=385−40010=−1.5z = \frac{385 - 400}{10} = -1.5, so P(Xˉ<385)≈P(Z<−1.5)=1−0.9332=0.0668P(\bar{X} < 385) \approx P(Z < -1.5) = 1 - 0.9332 = 0.0668.

Markers expect the central limit theorem to be named or justified (n≥30n \geq 30), the standard deviation of Xˉ\bar{X} (not of one battery), and a correct use of symmetry to handle the negative zz-score.

Practice questions

Original practice questions graded from foundation to exam level, each with a full worked solution. Try them before revealing the solution.

foundation2 marks
A population has mean μ=50\mu = 50 and standard deviation σ=12\sigma = 12. Random samples of size n=36n = 36 are taken. State the mean and the standard deviation of the sample mean Xˉ\bar{X}.
Show worked solution →

Mean of the sample mean. E(Xˉ)=μ=50E(\bar{X}) = \mu = 50.

Standard deviation of the sample mean. Var⁡(Xˉ)=σ2n\operatorname{Var}(\bar{X}) = \frac{\sigma^2}{n}, so

σXˉ=σn=1236=126=2.\sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}} = \frac{12}{\sqrt{36}} = \frac{12}{6} = 2.

Marker's note: one mark for each value. Dividing by nn instead of n\sqrt{n} (giving 13\frac{1}{3}) is the usual error.

foundation2 marks
A population has variance σ2=81\sigma^2 = 81. For samples of size 4949, find Var⁡(Xˉ)\operatorname{Var}(\bar{X}) and the standard deviation of Xˉ\bar{X}.
Show worked solution →

Variance. Var⁡(Xˉ)=σ2n=8149≈1.65\operatorname{Var}(\bar{X}) = \frac{\sigma^2}{n} = \frac{81}{49} \approx 1.65.

Standard deviation. Take the square root: 8149=97≈1.29\sqrt{\frac{81}{49}} = \frac{9}{7} \approx 1.29.

Marker's note: one mark for the variance, one for the standard deviation. Keep the exact 97\frac{9}{7} if you will use it later.

core3 marks
The heights of adults in a large population have mean 170170 cm and standard deviation 88 cm. A random sample of 6464 adults is taken. Use the central limit theorem to find the probability that the sample mean height is greater than 172172 cm. (Use P(Z≤2)=0.9772P(Z \leq 2) = 0.9772.)
Show worked solution →
Apply the central limit theorem
Since n=64≥30n = 64 \geq 30, Xˉ\bar{X} is approximately normal with mean 170170 and standard deviation 864=1\frac{8}{\sqrt{64}} = 1.
Standardise
z=172−1701=2z = \frac{172 - 170}{1} = 2.
Find the probability

P(Xˉ>172)≈P(Z>2)=1−0.9772=0.0228.P(\bar{X} > 172) \approx P(Z > 2) = 1 - 0.9772 = 0.0228.

Marker's note: one mark for the standard deviation 11 (with the CLT justified by n≥30n \geq 30), one for z=2z = 2, one for 0.02280.0228.

core3 marks
Waiting times at a call centre are strongly right-skewed, with mean 1212 minutes and standard deviation 1212 minutes. For a random sample of 3636 calls, find the approximate probability that the mean waiting time is between 1010 and 1515 minutes, and explain why a normal model is reasonable. (Use P(Z≤1)=0.8413P(Z \leq 1) = 0.8413 and P(Z≤1.5)=0.9332P(Z \leq 1.5) = 0.9332.)
Show worked solution →
Why normal
The population is not normal, but the central limit theorem says that for n≥30n \geq 30 the sampling distribution of the mean is approximately normal whatever the population's shape. Here n=36n = 36.
Parameters
E(Xˉ)=12E(\bar{X}) = 12 and σXˉ=1236=2\sigma_{\bar{X}} = \frac{12}{\sqrt{36}} = 2.
Standardise both bounds
z1=10−122=−1z_1 = \frac{10 - 12}{2} = -1 and z2=15−122=1.5z_2 = \frac{15 - 12}{2} = 1.5.

P(10<Xˉ<15)≈P(−1<Z<1.5)=0.9332−(1−0.8413)=0.9332−0.1587=0.7745.P(10 < \bar{X} < 15) \approx P(-1 < Z < 1.5) = 0.9332 - (1 - 0.8413) = 0.9332 - 0.1587 = 0.7745.

Marker's note: one mark for the CLT explanation, one for both zz-scores, one for 0.77450.7745. Using P(Z<−1)=0.8413P(Z < -1) = 0.8413 instead of 0.15870.1587 is the common slip.

core2 marks
A population has standard deviation 2020. Compare the standard deviation of the sample mean for samples of size 2525 and 100100, and describe the effect of increasing the sample size.
Show worked solution →

Compute both. For n=25n = 25: 2025=4\frac{20}{\sqrt{25}} = 4. For n=100n = 100: 20100=2\frac{20}{\sqrt{100}} = 2.

Interpret. Multiplying the sample size by 44 halves the standard deviation of Xˉ\bar{X}, because it depends on 1n\frac{1}{\sqrt{n}}. Larger samples give sample means that cluster more tightly around μ\mu.

Marker's note: one mark for both values, one for the 1n\frac{1}{\sqrt{n}} explanation (quadruple nn, halve the spread).

exam4 marks
A machine fills bottles with a mean of 500500 mL and a standard deviation of 44 mL. Quality control takes a random sample of nn bottles, where n≥30n \geq 30. Find the smallest nn such that the probability that the sample mean is within 11 mL of 500500 mL is at least 0.950.95. (Use P(Z≤1.96)=0.975P(Z \leq 1.96) = 0.975.)
Show worked solution →

Set up the condition. By the central limit theorem Xˉ\bar{X} is approximately N(500,16n)N\left( 500, \frac{16}{n} \right), with standard deviation 4n\frac{4}{\sqrt{n}}. We need

P(499<Xˉ<501)≥0.95.P(499 < \bar{X} < 501) \geq 0.95.

Symmetric bounds. P(−z<Z<z)=0.95P(-z < Z < z) = 0.95 when z=1.96z = 1.96, since P(Z≤1.96)=0.975P(Z \leq 1.96) = 0.975 leaves 0.0250.025 in each tail. So we need

14/n≥1.96⇒n≥7.84⇒n≥61.4656.\frac{1}{4/\sqrt{n}} \geq 1.96 \quad\Rightarrow\quad \sqrt{n} \geq 7.84 \quad\Rightarrow\quad n \geq 61.4656.

Smallest whole number. n=62n = 62.

Marker's note: one mark for σXˉ=4n\sigma_{\bar{X}} = \frac{4}{\sqrt{n}}, one for linking 0.950.95 to z=1.96z = 1.96, one for the inequality in nn, one for rounding up to 6262 (not down to 6161).

exam4 marks
A population has standard deviation 1010 and unknown mean μ\mu. For random samples of size 4040, the probability that the sample mean exceeds 5252 is 0.02280.0228. Find μ\mu, correct to two decimal places. (Use P(Z≤2)=0.9772P(Z \leq 2) = 0.9772.)
Show worked solution →
Find the zz-score
P(Z>z)=0.0228P(Z > z) = 0.0228 means P(Z≤z)=0.9772P(Z \leq z) = 0.9772, so z=2z = 2.
Standard deviation of Xˉ\bar{X}
1040≈1.5811\frac{10}{\sqrt{40}} \approx 1.5811.
Solve for μ\mu

52−μ10/40=2⇒μ=52−2×1040≈52−3.1623=48.84.\frac{52 - \mu}{10/\sqrt{40}} = 2 \quad\Rightarrow\quad \mu = 52 - 2 \times \frac{10}{\sqrt{40}} \approx 52 - 3.1623 = 48.84.

Marker's note: one mark for z=2z = 2, one for 1040\frac{10}{\sqrt{40}}, one for the equation, one for μ≈48.84\mu \approx 48.84.

exam5 marks
A fair six-sided die is rolled and XX is the number shown. (a) Show that E(X)=3.5E(X) = 3.5 and Var⁡(X)=3512\operatorname{Var}(X) = \frac{35}{12}. (b) The die is rolled 5050 times. Use the central limit theorem to estimate the probability that the mean score is greater than 44. (Use P(Z≤2.07)=0.9808P(Z \leq 2.07) = 0.9808.)
Show worked solution →

(a) Mean and variance of one roll. E(X)=1+2+3+4+5+66=216=3.5E(X) = \frac{1 + 2 + 3 + 4 + 5 + 6}{6} = \frac{21}{6} = 3.5.

E(X2)=1+4+9+16+25+366=916E(X^2) = \frac{1 + 4 + 9 + 16 + 25 + 36}{6} = \frac{91}{6}, so

Var⁡(X)=E(X2)−μ2=916−494=182−14712=3512.\operatorname{Var}(X) = E(X^2) - \mu^2 = \frac{91}{6} - \frac{49}{4} = \frac{182 - 147}{12} = \frac{35}{12}.

(b) Sample mean of 50 rolls. n=50≥30n = 50 \geq 30, so Xˉ\bar{X} is approximately normal with mean 3.53.5 and standard deviation

35/1250=35600≈0.2415.\sqrt{\frac{35/12}{50}} = \sqrt{\frac{35}{600}} \approx 0.2415.

z=4−3.50.2415≈2.07z = \frac{4 - 3.5}{0.2415} \approx 2.07, so

P(Xˉ>4)≈1−0.9808=0.0192.P(\bar{X} > 4) \approx 1 - 0.9808 = 0.0192.

Marker's note: two marks for (a) (the mean, then the variance via E(X2)−μ2E(X^2) - \mu^2); in (b), one for the standard deviation of Xˉ\bar{X}, one for z≈2.07z \approx 2.07, one for 0.01920.0192.

Practise this

Sources & how we know this

ExamExplained