Exam code: YMA01
1/480Still learning
Know0
Define bivariate data.
Bivariate data records two variables for each item, such as the height and the arm span of every person in a group.
It is displayed on a scatter diagram, where each item contributes a single point plotted from its two values.
Every other diagram in this topic handles one variable at a time.

Join for free to unlock a full flashcard set, track what you know,
and turn revision into real progress.
Before choosing a diagram to display a set of data, what three things about the data or the task do you need to establish?
Three things decide it:
whether the data is raw and ungrouped or has been grouped into classes
whether it is discrete or continuous
whether you are displaying a single set of data or comparing two
The third of these is the one most often forgotten, because a diagram that shows one set of data well is not always the one that lets two sets be read against each other.
Two diagrams are drawn from the same data. One shows every individual value and the other shows only a handful of summary values. What does the second one gain by throwing information away?
It becomes far quicker to read at a glance, and far easier to set beside a second data set for comparison.
A diagram showing every value carries more information, but it takes longer to interpret and two of them side by side are awkward to compare.
Choosing a diagram is always this trade-off: how much detail it keeps against how quickly it can be read.
Was this flashcard helpful?
Define bivariate data.
Bivariate data records two variables for each item, such as the height and the arm span of every person in a group.
It is displayed on a scatter diagram, where each item contributes a single point plotted from its two values.
Every other diagram in this topic handles one variable at a time.
Before choosing a diagram to display a set of data, what three things about the data or the task do you need to establish?
Three things decide it:
whether the data is raw and ungrouped or has been grouped into classes
whether it is discrete or continuous
whether you are displaying a single set of data or comparing two
The third of these is the one most often forgotten, because a diagram that shows one set of data well is not always the one that lets two sets be read against each other.
Two diagrams are drawn from the same data. One shows every individual value and the other shows only a handful of summary values. What does the second one gain by throwing information away?
It becomes far quicker to read at a glance, and far easier to set beside a second data set for comparison.
A diagram showing every value carries more information, but it takes longer to interpret and two of them side by side are awkward to compare.
Choosing a diagram is always this trade-off: how much detail it keeps against how quickly it can be read.
True or False?
Two bar charts drawn from exactly the same figures can give very different impressions of the data.
True.
The scale on the vertical axis is a choice: starting it well above zero, or stretching it, makes small differences between the bars look dramatic.
Nothing about the figures has changed, which is why the axes, their units and their labels have to be read before any conclusion is drawn from the picture.
A graph's vertical axis is labelled Population (millions), and one of its bars reaches 60. What is the population?
60 000 000, because the axis label says the numbers along it have been given in millions, so each unit on the scale stands for 1 000 000.
Abbreviating a scale like this is how very large numbers are made to fit on an axis, and the label is the only thing that tells you it has been done.
Define a stem and leaf diagram.
A stem and leaf diagram shows every raw data value, sorted into class intervals.
The stem is the leading digit or digits of a value and the leaf is its final digit, so a stem of 2 with a leaf of 3 might represent 23.
The leaves are written in order along each row, so the individual values and the shape of the distribution are both visible at once.
Why must a stem and leaf diagram always carry a key?
Because the stems and the leaves are only digits, and nothing in the diagram itself says what place value they stand for.
A key saying that a stem of 2 with a leaf of 3 represents 2.3 settles it, since the same row would otherwise be read as 23, or as 230.
Leaves are always single digits, so the key also shows how each value is split between its stem and its leaf.
Starting from an unordered list of data, why is a stem and leaf diagram usually drawn twice?
The first version simply gets every value into the right row, without worrying about order, and that is the stage where a value is most easily missed or written down twice.
The second version is then rewritten with the leaves in order within each row, and given its key.
Splitting the job in two makes it much less likely that a value is lost, and only the ordered version can be used to find the median and the quartiles.
True or False?
The numbers in brackets down the side of a stem and leaf diagram are the values in that row added together.
False.
They are the frequency of the row: how many leaves it contains.
They are not always included, but with a large amount of data they are worth having, both as a check that no value has been lost and as a quick way of counting along to the position of the median.
On a back-to-back stem and leaf diagram the two sets of leaves share one central column of stems. How do you read the values on the left-hand side?
Outwards from the stem, so the left-hand leaves increase from the centre towards the edge, which is the opposite direction to ordinary reading.
The leaf nearest the stem is therefore the smallest value in that row, not the largest.
A single key covers both sides, saying what a stem and leaf mean for each of the two groups.
A stem and leaf diagram holds so much data that its rows have become very long. What is usually done, and what does each row then hold?
Every stem is listed twice, so its data is spread over two rows instead of one.
The first row for a stem holds the leaves 0 to 4 and the second holds the leaves 5 to 9.
That halves the width of each class interval, which brings the shape of the distribution out more clearly.
A stem and leaf diagram has a row labelled LOW and a row labelled HIGH, each holding a value written out in full rather than as a leaf. What are those rows for?
They hold the outliers, kept outside the main body of the diagram.
Writing an outlier out in full avoids having to add a long run of empty stems just to reach one extreme value.
The stems in between then stay close together, so the shape of the rest of the data can still be seen.
Why can the median and the quartiles be found exactly from a stem and leaf diagram, when a grouped frequency table only gives estimates of them?
Because a stem and leaf diagram keeps every raw value, and keeps them in order, so you can count along to the position you need and read off the value actually sitting there.
Grouping into a frequency table throws the individual values away and records only how many fell into each class, so anything worked out from it has to be estimated.
That is why a stem and leaf diagram is such a convenient starting point when a box plot is wanted.
Which values can be read from a box plot, and which cannot?
You can read the minimum, the lower quartile, the median, the upper quartile and the maximum, together with any outliers.
No other individual data value can be read from a box plot: it shows the shape of the distribution, not the data itself.
What proportion of the data lies inside the box of a box plot, and what proportion lies along each whisker?
The box holds the middle 50% of the data, between the lower and upper quartiles.
Each whisker holds 25%: the lowest quarter of the data on one side and the highest quarter on the other.
True or False?
A box plot drawn with a deeper box represents a larger set of data.
False.
The box can be drawn to any depth, and the depth carries no information at all.
Only positions along the scale mean anything on a box plot.
How is an outlier shown on a box plot, and where does the whisker end when there is one?
The outlier is marked with a cross beyond the end of the whisker.
The whisker then stops at the most extreme value that is not an outlier, so it no longer reaches the smallest or largest value in the data set.
Two sets of data are to be compared using box plots. How should the two box plots be presented?
One above the other, drawn against the same scale on the horizontal axis, with each one clearly labelled.
Reading across from one plot to the other is only valid if both use the same scale.
Define cumulative frequency.
The running total of the frequencies: the frequency of a class added to the frequencies of all the classes below it.
When a cumulative frequency graph is plotted, which value of each class is the cumulative frequency plotted against, and why?
The upper boundary of the class.
The cumulative frequency counts every value in that class as well as every value in the classes below it, so that total is only reached at the top of the class.
Complete the positions used to estimate the quartiles of data values from a cumulative frequency graph:
The completed positions are:
Find each position on the cumulative frequency axis, read across to the curve, then down to the horizontal axis to get the data value.
A question asks how many data values are greater than a particular value. How is that found from a cumulative frequency graph?
Read up from that value on the horizontal axis to the curve, then across to the cumulative frequency axis.
That gives the number of values less than it, so subtract the result from the total frequency to get the number greater than it.
True or False?
The median read from a cumulative frequency graph is the exact median of the data.
False.
The data has been grouped, so the individual values are no longer known.
The median, the quartiles and any percentile read from the graph are all estimates.
What is plotted on the vertical axis of a histogram, and why is it not the frequency?
Frequency density.
The classes of a histogram may have different widths, so it is the area of each bar that represents the frequency, not its height. Plotting frequency density is what makes the areas come out in proportion to the frequencies.
Complete the formula used to find the height of each bar of a histogram:
The completed formula is:
In some questions the area of a bar is only proportional to the frequency, and a scale factor has to be found first.
True or False?
In a histogram the bars for adjacent classes always touch.
True.
A histogram displays grouped continuous data, so each class ends where the next begins and there is no gap between the bars.
A bar chart, which displays discrete or qualitative data, does have gaps between its bars.
A grouped frequency table gives the classes as 1 to 5, 6 to 10 and 11 to 15. What must you do before you can find the class widths for a histogram?
Close the gaps between the classes first, so that the upper boundary of one class is the lower boundary of the next.
Check whether the values were rounded or truncated before deciding where the boundaries lie.
For data rounded to the nearest whole number the boundaries become 0.5 to 5.5, 5.5 to 10.5 and 10.5 to 15.5, each of class width 5.
A histogram question tells you that the area of each bar is proportional to the frequency rather than equal to it. How do you find the frequencies?
Use a class whose frequency you are given to find the scale factor: work out the area of that bar and compare it with the frequency it represents.
That tells you how many data values one unit of area stands for, and every other bar can then be converted in the same way.
A question of this kind will always give you enough information to find the scale factor.
How do you estimate the frequency for part of a class from a histogram, for example the values above 13 in the class ?
Find the area of that part of the bar, from 13 to 15, and convert it to a frequency exactly as you would for a whole bar.
This assumes the values are spread evenly across the class, so the answer is an estimate.
True or False?
The two ends of a frequency polygon are joined down to the horizontal axis.
False.
The polygon is left open, unless a frequency at the end happens to be zero. The first and last points are not joined to the axis, and the last point is not joined back to the first.
The points themselves are plotted at the midpoint of each class and joined with straight lines.
Define an outlier.
An outlier is an extreme data value that does not fit the general pattern of the rest of the data.
It may come from a genuine one-off event, or from a mistake made when the data was collected or recorded.
Which of those two it is decides whether the value should be kept.
Complete the two boundaries beyond which a value counts as an outlier, by filling in the missing subscripts:
The completed boundaries are:
Each boundary sits interquartile ranges out from the near end of the middle 50%, so the lower one is measured from
and the upper one from
.
The value of is always given to you, and is commonly 1.5.
Outliers are not always defined using the quartiles. What is the other definition you may be given, and what is typically in it?
A value counts as an outlier if it lies more than standard deviations from the mean, so below
or above
.
Here is commonly 2, where for the interquartile-range definition it is commonly 1.5.
The question tells you which definition to use and what is, and the two need not flag the same values in the same data set.
The ages of children at a party include a value of 29, which the outlier test flags. Should it be removed from the data set?
Yes, because 29 cannot be the age of a child, so it is almost certainly an error in the data.
The general rule is that an outlier is kept if it is a valid piece of data and removed if it is likely to be a mistake, so the decision comes from the context rather than from the test.
A flagged value of 13 at the same party would be investigated rather than deleted, because a 13-year-old there is unusual rather than impossible.
A single outlier is added to a data set. What happens to the range, and what happens to the interquartile range?
The range is changed completely, because it is the largest value minus the smallest and the outlier immediately becomes one of those two.
The interquartile range is essentially unaffected, because it depends only on the two quartiles, which sit well inside the data.
That is what makes the interquartile range the more reliable measure of spread whenever a data set may contain extreme values.
You are drawing a box plot from summary statistics alone, with no individual data values, and the minimum turns out to be an outlier. Where does the lower whisker end?
At the lower outlier boundary itself, .
Without the individual values there is no way to tell which value is the smallest one that is not an outlier, so the boundary is drawn to instead.
Define a measure of spread.
A measure of spread gives information about how varied a set of data is, rather than about where it sits.
The range, the interquartile range, the standard deviation and the variance are all measures of spread.
A smaller value means the data is more consistent, and a larger one means it is more varied.
When comparing two data sets, what two kinds of statistic must your comparison mention, and which ones pair together?
A measure of location and a measure of spread, one of each.
The two pairings are the mean with the standard deviation or the variance, and the median with the range or the interquartile range.
A data set contains extreme values. Which pairing of location and spread should be used to describe it, and why?
The median with the interquartile range, because both are resistant to extreme values while the mean and the standard deviation are not.
The mean and the standard deviation are each worked out from every value in the set, so a single extreme one pulls both of them away from where the bulk of the data actually sits.
Use the mean and the standard deviation instead when the data is roughly symmetrical and free of extreme values.
Two groups are timed completing a puzzle, and group A has the lower median. Why is it not enough to write only that group A's median is lower?
Because a comparison has to be written in the context of the data: here a lower median time means group A was faster, which is the thing actually being asked about.
Whether a lower value is the better one depends entirely on what is being measured, and it would be the other way round for a set of test scores.
A value is added to a data set. What decides whether the mean goes up or down?
Whether the new value is above or below the current mean: one above it pulls the mean up, and one below it pulls the mean down.
Removing a value does the opposite, so taking out a value from above the mean pulls the mean down.
A value exactly equal to the mean leaves it unchanged.
True or False?
Adding a very extreme value to a data set moves the median by about as much as adding an ordinary value would.
True.
The median is decided by position, so a new value counts once wherever it lands, and a value of 500 shifts the middle position by exactly as much as a value of 5 landing on the same side of it.
It is the mean that behaves differently, because it is built from the sizes of the values rather than their positions, so an extreme one moves it a long way.
A value is added to a data set. Why can you not say in advance whether the median will change?
Because the median is decided by position, and adding a value shifts every position above where it lands while leaving the ones below it alone.
Whether the middle position ends up on a different value depends on where in the ordered data the new one falls, so each case has to be checked individually.
The same is true of the quartiles.
Define skewness.
Skewness describes the way a non-symmetrical distribution leans.
A distribution whose tail stretches out to the right has positive skew, and one whose tail stretches out to the left has negative skew.
Any distribution is either symmetrical or it has skewness of one of these two kinds.
True or False?
A positively skewed distribution has most of its data concentrated at the high end.
False.
The skew is named after the direction of the tail, not of the bulk of the data.
In a positively skewed distribution most of the data is bunched at the low end, and it is the long thin tail that stretches away to the right.
How do the gaps between the three quartiles tell you which way a distribution is skewed?
Compare with
, which on a box plot are the widths of the two halves of the box.
If the median sits closer to the lower quartile, so that , the skew is positive.
If it sits closer to the upper quartile, so that , the skew is negative.
In a positively skewed distribution, how do the mean, the median and the mode compare in size, and why do they come out in that order?
They come out as .
The mode sits at the peak, which is down at the low end, while the mean is dragged out towards the long tail because it is worked out from every value including the extreme ones lying in that tail.
The median settles between the two, since it moves with position rather than with size; for negative skew the whole order reverses, giving .
A distribution is symmetrical. What does that tell you about its quartiles, and about its three averages?
The median sits exactly midway between the quartiles, so and the two halves of the box on a box plot come out equal.
The mean, the median and the mode are all equal, or very nearly so, because there is no tail on either side to drag the mean away from the peak.
Symmetry is the case that skewness is measured against, which is why the quartile test and the ordering of the averages are both written as comparisons rather than as formulae.
By signing up you agree to our Terms and Privacy Policy