Representation of Data (Cambridge (CIE) A Level Maths: Probability & Statistics 1): Flashcards

Exam code: 9709

1/26

0Still learning

Know0

  • Complete the kind of data these two diagrams are drawn from:

    A histogram and a cumulative frequency graph both need \_\_\_\_\_\_ data that has been \_\_\_\_\_\_ into classes.

Cards in this collection (26)

  • Complete the kind of data these two diagrams are drawn from:

    A histogram and a cumulative frequency graph both need \_\_\_\_\_\_ data that has been \_\_\_\_\_\_ into classes.

    The completed sentence is:

    A histogram and a cumulative frequency graph both need continuous data that has been grouped into classes.

    Both diagrams are built from class intervals rather than from individual values, so the data has to be capable of being divided into classes in the first place.

  • What can a stem-and-leaf diagram show you that a box plot cannot?

    Every individual data value, because a stem-and-leaf diagram lists them all while a box plot reduces the data to five numbers.

    That also means further statistics such as the mean, the mode and the standard deviation can be worked out from a stem-and-leaf diagram but not from a box plot, which only gives the median, the quartiles and the range.

  • True or False?

    A grouped frequency table contains enough information to draw a stem-and-leaf diagram.

    False.

    A stem-and-leaf diagram displays every individual data value, and a grouped table records only how many values fell in each class, so the values it would need have already been thrown away.

    The reverse is possible: a grouped table can always be built from a stem-and-leaf diagram.

  • What three things about the data should you establish before choosing how to present it?

    Whether the data is raw and ungrouped or has already been grouped, whether it is discrete or continuous, and whether you need to display one set or compare two.

    Each of the three rules some diagrams out: a histogram cannot be drawn from discrete data, and a comparison needs a diagram that can carry two data sets at once.

  • A grouped frequency table has classes of different widths. Which diagram is designed to cope with that, and what would go wrong without it?

    A histogram, which is built to handle unequal class intervals.

    A diagram that simply plotted the frequency as the height of each bar would misrepresent the data, because a wide class would appear more important than a narrow one purely for being wide.

  • Define a stem-and-leaf diagram.

    A stem-and-leaf diagram displays numerical data by splitting each value into a stem and a leaf, with all the values sharing a stem written along a single row.

    It is always read together with a key, and it suits two-digit data best, although three-digit data can be shown as well.

  • Why must a stem-and-leaf diagram always be given a key?

    Because the split between the stem and the leaf carries no place value of its own, so a row reading 2 \mid 3 could mean 23, or 2.3, or 230.

    The key fixes which, by stating what one particular entry represents.

  • How would the three-digit value 122 be entered on a stem-and-leaf diagram?

    As 12 \mid 2, with the leading digits 12 forming the stem and the final digit 2 forming the leaf.

    A leaf is always a single digit, so any extra digits are absorbed into the stem.

  • What do the numbers in brackets down the side of a stem-and-leaf diagram tell you?

    How many data values are on that row, which is the frequency of that class interval.

    They are not always included, but they are worth having when there is a lot of data, because they let the total be checked without counting the leaves again.

  • True or False?

    The leaves in a finished stem-and-leaf diagram must be written in order of size.

    True.

    An ordered diagram is what makes it possible to count to a particular position in the data, so an unordered one is only ever a working draft.

    Starting from unordered data it is usual to draw the diagram twice, once to collect the leaves onto the right rows and once to put them in order.

  • With a large amount of data each stem may be listed twice. Complete the split:

    The first row holds the leaves 0 to \_\_\_\_\_\_, and the second row holds the leaves \_\_\_\_\_\_ to 9.

    The completed split is:

    The first row holds the leaves 0 to 4, and the second row holds the leaves 5 to 9.

    Splitting the stems this way halves the length of each row, which spreads the data out and makes its shape easier to see.

  • In a back-to-back stem-and-leaf diagram, in which direction do the left-hand leaves increase?

    Outwards from the stem, so the smallest leaf sits next to the stem in the middle and the values get larger as you read leftwards.

    The right-hand side increases outwards too, which keeps both sides reading from the centre and lets the two distributions be compared shape against shape.

  • What five values is a box plot built from, and what does the box itself contain?

    The minimum, the lower quartile, the median, the upper quartile and the maximum.

    The box runs from the lower quartile to the upper quartile, so it holds the middle 50% of the data, and each whisker covers a further quarter.

    Those are exactly the five values an ordered stem-and-leaf diagram gives you, which is why a box plot can be drawn straight from one.

  • True or False?

    The depth of the box on a box plot tells you how many data values there are.

    False.

    The box may be drawn to any depth at all, and its depth carries no information: only the positions of its edges along the scale mean anything.

    A box plot uses a single axis, so everything it has to tell you is read off horizontally.

  • How must two box plots be drawn if the two data sets are to be compared?

    One above the other against a single shared scale, so that the medians, quartiles and extremes line up and can be compared directly by eye.

    On two different scales the comparison would be meaningless, because a wider box would no longer mean a more varied data set.

  • A box plot has its minimum at 8 cm, its lower quartile at 14 cm, its upper quartile at 30 cm and its maximum at 58 cm. Find the range and the interquartile range.

    Subtracting the two extremes gives the range, and subtracting the two quartiles gives the interquartile range:

    \text{range} = 58 - 8 = 50 \text{ cm}

    \text{IQR} = 30 - 14 = 16 \text{ cm}

    The median is not needed for either, which is why a box plot gives both at a glance.

  • Why is cumulative frequency plotted against the upper boundary of each class?

    Because the cumulative frequency counts every value up to and including that class, and the largest value the class can hold is its upper boundary.

    Plotting it against the midpoint or the lower boundary would claim the running total had been reached earlier than it really was.

  • On a cumulative frequency graph of n values, complete the cumulative frequencies at which the quartiles are read:

    The lower quartile is read at \_\_\_\_\_\_, the median at \frac{n}{2}, and the upper quartile at \_\_\_\_\_\_.

    The completed readings are:

    The lower quartile is read at \frac{n}{4}, the median at \frac{n}{2}, and the upper quartile at \frac{3 n}{4}.

    Read across from that value on the vertical axis to the curve, then straight down to the horizontal axis to get the quartile itself, and find any percentile the same way from the matching percentage of the total frequency.

  • A cumulative frequency graph of 30 puppies shows that 25 of them are shorter than 53.5 cm. What percentage are longer than 53.5 cm?

    The graph gives the number below a value, so subtract from the total: 30 - 25 = 5 puppies are longer than 53.5 cm.

    As a percentage that is \frac{5}{30} = 16.7\% to three significant figures.

  • How do you find the frequency of a single class from a cumulative frequency graph?

    Read the cumulative frequency at the class's upper boundary, then subtract the cumulative frequency at its lower boundary.

    For a class running from 40 cm to 45 cm, readings of 16 and 8 give 16 - 8 = 8 values in that class, because the running total has grown by 8 across it.

  • True or False?

    A cumulative frequency graph can never fall as you read it from left to right.

    True.

    A frequency can never be negative, so adding the next class's frequency to the running total can only leave it unchanged or make it larger.

    A graph that dipped would be claiming a class with a negative number of values in it.

  • Complete the formula for the height of a bar on a histogram:

    \text{frequency density} = \frac{\_\_\_\_\_\_}{\_\_\_\_\_\_}

    The completed formula is:

    \text{frequency density} = \frac{\text{frequency}}{\text{class width}}

    Dividing by the class width is what lets classes of different widths share one diagram fairly, because it measures how tightly the values are packed rather than simply how many of them there are.

  • A bar on a histogram has height 3 and spans from 13 to 15. How many data values does it represent?

    Its area, which is 3 \times 2 = 6 values.

    The area of a bar, or of any part of a bar, gives the frequency it stands for, so multiply the height by the width measured along the horizontal axis.

  • True or False?

    There are never gaps between the bars of a histogram.

    True.

    A histogram displays continuous data, whose classes run edge to edge with no values able to fall between them, so the bars meet.

    Where a grouped table appears to leave a gap, its boundaries are adjusted to close it before the histogram is drawn.

  • Why must a histogram's horizontal axis have an even scale?

    Because the width of each bar has to be read off that scale to get its class width, and the bars are deliberately allowed to differ in width.

    On an uneven scale two bars drawn the same width would stand for different class widths, and nothing on the diagram could be trusted.

  • A class running from 15 to 30 contains 6 values. What is its frequency density, and what does that number mean?

    Its frequency density is \frac{6}{15} = 0.4.

    That says the values are spread thinly, at 0.4 of a value per unit along the horizontal axis, which is why a wide class draws a much lower bar than a narrow class holding the same 6 values.

Sign up to unlock flashcards

or