Exam code: H240
1/230Still learning
Know0
Define continuous random variable.
A continuous random variable is a random variable that can take any value within a range, rather than only certain separate values.
Continuous random variables usually measure something, so height, weight and time are all continuous.
That is the contrast with a discrete random variable, which counts, and the difference decides which distributions can model the variable.

Join for free to unlock a full flashcard set, track what you know,
and turn revision into real progress.
For a continuous distribution, what is , and what follows from it?
for every value of
, because probability is the area under the graph and a single value is a line with no width.
What has a probability is a range of values: the area between and
is
, and the total area under the graph is 1.
What follows is convenient: since the endpoints contribute nothing, strict and weak inequalities give the same answer.
True or False?
In , the second number in the bracket is the standard deviation.
False.
The second number is the variance, : the notation says so, but it is easy to read past.
So has variance 36 and standard deviation
.
This matters because calculators ask for the standard deviation, so entering the second number straight from the bracket is a common error; square root it first, unless the variance is already written as a square, as in .
Was this flashcard helpful?
Define continuous random variable.
A continuous random variable is a random variable that can take any value within a range, rather than only certain separate values.
Continuous random variables usually measure something, so height, weight and time are all continuous.
That is the contrast with a discrete random variable, which counts, and the difference decides which distributions can model the variable.
For a continuous distribution, what is , and what follows from it?
for every value of
, because probability is the area under the graph and a single value is a line with no width.
What has a probability is a range of values: the area between and
is
, and the total area under the graph is 1.
What follows is convenient: since the endpoints contribute nothing, strict and weak inequalities give the same answer.
True or False?
In , the second number in the bracket is the standard deviation.
False.
The second number is the variance, : the notation says so, but it is easy to read past.
So has variance 36 and standard deviation
.
This matters because calculators ask for the standard deviation, so entering the second number straight from the bracket is a common error; square root it first, unless the variance is already written as a square, as in .
A normal distribution is symmetrical about . What two things follow from that?
The three averages all coincide:
and half the area lies on each side of the mean:
The second is worth having ready. It gives you a probability with no calculation, and it is often the quickest way to check that an answer is on the right side of the mean.
Complete the proportions of a normal distribution that lie within one, two and three standard deviations of the mean:
within , about
% of the data
within , about
% of the data
within , about
% of the data
The completed proportions are:
within , about 68% of the data, which is roughly two thirds
within , about 95% of the data
within , about 99.7% of the data, which is nearly all of it
These are worth knowing by heart as a sense check. If a calculation says that 40% of the data lies within one standard deviation of the mean, something has gone wrong.
Where are the points of inflection on a normal distribution curve?
At , exactly one standard deviation either side of the mean.
Those are the two places where the curve stops bending one way and starts bending the other, as it changes from falling ever more steeply to falling ever less steeply.
This gives you a way to read off a sketch: it is the horizontal distance from the mean to a point of inflection.
How does a normal curve change when changes, and how when
changes?
Changing translates the curve horizontally, moving the whole shape along without altering it.
Changing stretches it horizontally: a small variance gives a tall curve with a narrow centre, and a large variance a short curve with a wide centre.
The reason a narrower curve has to be taller is that the total area is always 1, so squeezing the curve inwards must push it upwards.
What has to be true of a real-life variable before a normal distribution is a sensible model for it?
It must be continuous, so it measures something, its distribution must be symmetrical and bell-shaped with a single mode, and the population needs to be large enough.
So, for example, a variable produced by a random number generator cannot be modelled this way, because every value is equally likely and it has no mode.
Nor can how long a human lives, because that distribution is not symmetrical.
True or False?
A normal distribution cannot model height, because it allows any real value and a height cannot be negative.
False.
It is true that a normal distribution is defined for every real number, but values more than about four standard deviations from the mean have a probability density of practically zero.
So a normal model of human height puts a negligible probability on the impossible values and describes the realistic ones well, which is what makes it usable for quantities like height and weight that have a natural floor.
A model does not have to be perfect to be useful, only good enough over the range that matters.
How do you find for a normal distribution?
The probability is the area under the normal curve between and
, and since that curve is too complicated to integrate the area is found numerically, using the Normal Cumulative Distribution function on a calculator.
Often shortened to NCD, Normal CD or Normal Cdf, it needs four inputs: the lower bound , the upper bound
, the mean, and the standard deviation.
Sketch the curve and shade the area you want before you start: it costs a few seconds and tells you whether the answer you get back is plausible.
True or False?
A calculator's Normal Probability Density function gives you .
False.
It gives the probability density at that point, which is the height of the curve, not a probability.
For a normal distribution for every value of
, so a function returning a non-zero answer cannot be giving you one.
The function you want is always the Normal Cumulative Distribution; the density function has no use in this course, so if your calculator offers you Normal PD or Normal Pdf, that is not it.
Your calculator wants both a lower and an upper bound, but you need , which has no upper bound. What do you enter?
A value far enough above the mean that everything beyond it is negligible: more than four standard deviations above is accurate enough, and the easiest thing to type is a string of 9s, or .
It works because the probability of being more than three standard deviations above the mean is already less than 0.0015, and beyond four it is less than 0.000032, so the tail you are chopping off cannot affect the answer at the accuracy you are working to.
For , do the same at the other end with a large negative lower bound.
A calculator will work out any normal probability directly, so when are results like still needed?
Whenever you cannot simply type the numbers in.
That happens when the mean or the standard deviation is unknown, when you have been given only a diagram rather than values, and when you are working with the inverse distribution and have a probability rather than a value.
The ones worth having ready are
and they hold because the probability of a single value is zero, so nothing is lost or double counted at the boundaries.
You know that and you need the value of
. What do you use?
The Inverse Normal Distribution function, sometimes called InvN.
It runs the cumulative function backwards: the cumulative function takes a value and returns a probability, while the inverse takes a probability and returns a value.
Enter the area , the mean, and the standard deviation; if your calculator asks which tail, this is the left tail, because
is the area to the left of
.
For you are told
and want
. Complete the two things needed before the inverse normal function can be used:
the area to the left,
the standard deviation,
The completed values are:
the area to the left,
the standard deviation,
The area is converted with , because the function expects the area to the left; if your calculator offers a right tail, you can enter 0.175 directly instead.
The standard deviation is , since the 36 in the bracket is the variance. Those two conversions give
to 3 significant figures.
What check should you always run on an answer from the inverse normal distribution?
Compare it with the mean, which tells you at once whether it is on the right side.
If is less than 0.5 then
must be below the mean, and if it is more than 0.5 then
must be above it.
This catches the commonest mistake in inverse normal questions, which is working with the wrong tail; a quick sketch with the area shaded makes the comparison obvious.
Define standard normal distribution.
The standard normal distribution is the normal distribution with mean 0 and standard deviation 1, written as in the usual notation.
The letter is reserved for it, so seeing
in a question tells you its mean and standard deviation without their having to be stated.
Complete the formula that turns a value of into a value of the standard normal variable:
The completed formula is:
Subtracting the mean moves the centre of the distribution to 0, and dividing by the standard deviation rescales the spread so that the standard deviation becomes 1, which is what makes it standard.
For a normal variable with mean
and standard deviation
, how is
written in terms of the standard normal variable?
Standardise the boundary value in exactly the same way as you would standardise any other value:
The two probabilities are equal because standardising only relabels the horizontal axis, and relabelling does not change the area under the curve.
True or False?
A negative value always comes from a value of
that lies below the mean.
True.
Standardising subtracts the mean, so the numerator is negative exactly when
lies below the mean, and dividing by the positive standard deviation cannot change that sign.
It is a quick check worth making: if a question describes a value below the mean and your comes out positive, you have worked from the wrong tail.
What does the table of percentage points of the standard normal distribution give you?
For each of nine commonly used probabilities, it gives to three decimal places the value of for which
holds.
The four met most often are 1.645 when is 0.95, then 1.960 for 0.975, 2.326 for 0.99 and 2.576 for 0.995.
The same values come from the inverse normal function on a calculator, so the table is a convenience rather than the only way to reach them.
The standard deviation of a normal variable is known but its mean is not, and you are told that for that variable.
How do you find the mean?
Find the value for which
holds, which comes out negative because 12 lies below the mean.
Then put that value, the value 12 and the known standard deviation into
or its rearrangement
and solve the resulting equation to find
itself.
Carry more decimal places in than you need in the answer, because it gets multiplied by the standard deviation and any rounding is magnified.
Both the mean and the standard deviation of a normal variable are unknown. What must the question give you, and what do you do with it?
It must give you two probabilities, each one attached to a specific value of .
Each probability gives a value, each
value with its own
gives one equation of the form
, and the two together are simultaneous equations in the two unknowns.
Keep careful track of which value belongs to which value of
in the two equations, because pairing them the wrong way round gives a plausible-looking answer that is wrong.
By signing up you agree to our Terms and Privacy Policy