Working with Distributions (Cambridge (CIE) A Level Maths: Probability & Statistics 1): Flashcards

Exam code: 9709

1/13

0Still learning

Know0

  • When choosing a distribution to model something, what is the first question to ask?

Cards in this collection (13)

  • When choosing a distribution to model something, what is the first question to ask?

    Whether the variable counts or measures.

    Counting points to a discrete distribution and measuring to a continuous one.

    That single question rules out two of the three distributions straight away, which is why it is worth asking before anything else.

  • Fill in which distribution each description points to:

    counts successes in a fixed number of trials: \_\_\_\_\_\_

    counts trials up to and including the first success: \_\_\_\_\_\_

    measures something with a symmetrical, bell-shaped spread: \_\_\_\_\_\_

    The completed matches are:

    counts successes in a fixed number of trials: binomial

    counts trials up to and including the first success: geometric

    measures something with a symmetrical, bell-shaped spread: normal

    The two discrete ones differ only in what is fixed and what varies, so the phrase to look for is whether the number of trials is decided in advance.

  • How can real data tell you whether a normal model is reasonable?

    Draw a histogram of the data and look at its shape: roughly symmetrical and bell-shaped supports a normal model.

    As more data is collected the outline of the histogram should smooth out and come to resemble the normal curve itself.

    A histogram that is lopsided, or that has two peaks, is telling you to look elsewhere.

  • True or False?

    A variable that counts something can never be modelled by a normal distribution.

    False.

    A count is discrete, so strictly a continuous distribution does not fit it.

    But where the count ranges over a great many values, a normal distribution can describe it closely enough to be worth using, which is exactly what a normal approximation does.

  • Cow masses follow \text{N}(550, 80^{2}) and a cow is "beefy" if it exceeds 700 kg. For 10 cows, how do you find the probability that at most one is beefy?

    Use the normal distribution first to find the probability that one cow is beefy: p = \text{P}(M > 700) = 0.0303.

    Then use that as the probability of success in a binomial, so the number of beefy cows is X \sim \text{B}(10, 0.0303), and find \text{P}(X \le 1).

    The normal supplies the parameter and the binomial answers the question.

  • In a question using two distributions, why must you state what each variable represents?

    Because the two variables measure completely different things: one is a mass in kilograms, the other a count of cows out of 10.

    Mixing them up means putting a mass into a binomial or a count into a normal, which is the usual way these questions go wrong.

    Writing "let M be the mass of a cow" and "let X be the number of beefy cows" before calculating anything prevents it.

  • Why approximate a binomial by a normal distribution rather than using the binomial formula?

    A cumulative binomial probability needs one term for every value in the range, so a range spanning dozens of values means dozens of separate calculations.

    A normal calculation replaces all of them with a single standardisation and a couple of table lookups.

    The pay-off grows with n, which is exactly when the binomial becomes unmanageable.

  • For a normal approximation, n must be large enough that two quantities both exceed 5. Fill them in:

    \_\_\_\_\_\_ > 5

    \_\_\_\_\_\_ > 5

    The completed conditions are:

    np > 5

    nq > 5

    where q = 1 - p.

    Both are needed because each guards one end: a small np crowds the distribution against 0 and a small nq crowds it against n, and either way it is too lopsided for a symmetrical curve to fit.

  • Which normal distribution approximates X \sim \text{B}(n, p)?

    The one with the same mean and the same variance as the binomial, so

    \mu = np \text{ and } \sigma^{2} = np(1 - p)

    Matching those two quantities is the whole of what makes the approximation work.

    Remember that the second of them is the variance, so it needs square-rooting before it can be used in a standardisation.

  • Define continuity correction.

    A continuity correction is the adjustment made when a discrete variable is approximated by a continuous one.

    Each whole number is widened to the interval reaching half a unit either side of it, since every number in that interval rounds to it.

    So X = 4 becomes 3.5 < X_{N} < 4.5.

  • Fill in the continuity correction each inequality needs:

    \text{P}(X \le k) becomes \text{P}(X_{N} < \_\_\_\_\_\_)

    \text{P}(X < k) becomes \text{P}(X_{N} < \_\_\_\_\_\_)

    The completed corrections are:

    \text{P}(X \le k) becomes \text{P}(X_{N} < k + 0.5)

    \text{P}(X < k) becomes \text{P}(X_{N} < k - 0.5)

    The correction always reaches to the upper edge of the largest whole number you want to include.

    For X \le k that largest value is k itself, whose edge is at k + 0.5; for X < k it is k - 1, whose edge is at k - 0.5.

  • For X \sim \text{B}(1250, 0.4), how do you approximate \text{P}(485 \le X \le 530)?

    First find the approximating distribution: \mu = 500 and \sigma^{2} = 300, so use \text{N}(500, 300).

    Both endpoints are included, so the correction widens the range outwards to 484.5 \le X_{N} \le 530.5.

    Standardising both ends and using the table then gives 0.776.

  • True or False?

    The normal approximation gives a probability to values the binomial says are impossible.

    True.

    The approximating normal distribution is continuous and defined for every real number, so it assigns probability to non-integer values, and to values below 0 or above n.

    Those probabilities are negligible when np and nq both exceed 5, which is precisely what those conditions are protecting against.

Sign up to unlock flashcards

or