Tabulation, Diagrams & Representation (Edexcel GCSE Statistics: Foundation): Flashcards

Exam code: 1ST0

1/75

0Still learning

Know0

Cards in this collection (75)

  • Complete the sentence about bar charts.

    On a bar chart the horizontal axis shows the different \_\_\_\_\_\_ being counted, and the height of each bar shows its \_\_\_\_\_\_ in the data.

    The completed sentence is:

    On a bar chart the horizontal axis shows the different outcomes being counted, and the height of each bar shows its frequency in the data.

  • Why are the bars on a bar chart drawn with gaps between them?

    Because a bar chart shows discrete data, where each bar is a separate category rather than part of a continuous scale.

    The gaps make it clear that there are no values in between.

  • True or False?

    All the bars on a bar chart must be drawn with the same width.

    True.

    Bars on a bar chart are always drawn with equal widths, so that the height on its own represents the frequency.

    If the widths varied, a taller bar would no longer reliably mean a larger frequency.

  • When is a vertical line graph a better choice than a bar chart?

    When the data is numerical and there are a lot of different values to show, such as test scores out of 20.

    Thin lines fit where wide bars would be crowded, so the spread of the data and any unusual values stand out clearly.

  • What is the difference between a dual bar chart and a bar-line chart?

    A dual bar chart shows two data sets measuring the same variable, so the bars sit side by side on one shared scale.

    A bar-line chart shows two different variables, one as bars and one as a line, so each can have its own scale: monthly rainfall as bars with monthly temperature as a line.

  • Define pictogram.

    A pictogram represents frequencies using symbols instead of bars, with no axes at all.

    A key states what one symbol is worth, and part-symbols stand for part-values, so if one symbol means 2 people then half a symbol means 1 person.

  • What is a composite bar chart, and what does the height of one coloured section show?

    A composite bar chart stacks the categories within each bar, so the overall height of a bar gives the total frequency.

    A coloured section shows the frequency for that category alone, read as the difference between the two ends of the section rather than as the value at its top.

  • Why can a composite bar chart tell you something a plain bar chart of the same totals cannot?

    Because it shows the breakdown within each bar as well as the total.

    A time slot can have the smallest total traffic of the day and still contain the largest number of lorries, and only the composite chart reveals that.

  • What does a pie chart show about the categories in a data set?

    A pie chart shows the proportions of the whole that each category takes up.

    That makes it easy to see at a glance which categories are the largest and smallest shares of the data, and how they compare with one another.

  • Complete the formula for the angle of a sector in a pie chart.

    \text{angle} = \frac{\_\_\_\_\_\_}{\text{total frequency}} \times \_\_\_\_\_\_

    The completed formula is:

    \text{angle} = \frac{\text{frequency}}{\text{total frequency}} \times 360^{\circ}

    The angles for all the categories should add up to 360 degrees, which is a useful check on the arithmetic.

  • In a survey of 40 people, 8 chose walking. What angle represents walking on a pie chart?

    Take the fraction of the total and multiply it by 360 degrees:

    \frac{8}{40} \times 360^{\circ} = 72^{\circ}

    The fraction \frac{8}{40} simplifies to \frac{1}{5}, so walking takes up a fifth of the circle.

  • True or False?

    A pie chart on its own tells you how many people chose each option.

    False.

    A pie chart shows only what share of the total each option has, so a sector covering a quarter could be 50 people or 5000.

    You need to be told the total frequency before any sector can be turned into a number.

  • A pie chart is drawn for 30 students. How many degrees represent one student?

    The 360 degrees of the circle are shared between 30 students, so \frac{360}{30} = 12 degrees each.

    Every category's angle can then be found by multiplying its frequency by 12, which is quicker than working out a separate fraction for each one.

  • A sector of a pie chart measures 30 degrees and represents 15 people. How many people are in the whole data set?

    The whole circle is 360 degrees, which is 12 times the 30 degree sector.

    So the whole data set holds 15 \times 12 = 180 people.

  • The total value of a shop's stock is £48 000, and the tennis sector of a pie chart measures 72 degrees. What is the tennis stock worth?

    A sector's share of the circle is the same as its share of the value, so multiply the total by that fraction of 360 degrees.

    That gives 48000 \times \frac{72}{360} = 9600, so the tennis stock is worth £9 600.

  • Why must a stem-and-leaf diagram always have a key?

    Because the same diagram could stand for very different numbers.

    A stem of 3 with a leaf of 6 could mean 36 or 3.6, and only the key says which.

  • Define stem-and-leaf diagram.

    A stem-and-leaf diagram splits each value into a stem, its leading digits, and a leaf, its last digit, putting all the values that share a stem on one row.

    It orders the data and groups it into classes while still showing every original value.

  • Complete the sentence about stem-and-leaf diagrams.

    The stems act as the \_\_\_\_\_\_ intervals, and the number of leaves on a stem gives the \_\_\_\_\_\_ for that interval.

    The completed sentence is:

    The stems act as the class intervals, and the number of leaves on a stem gives the frequency for that interval.

  • In a stem-and-leaf diagram the median turns out to be a leaf of 1 on the stem of 2. What is the median?

    The median is 21, not 1.

    A leaf gives only one digit, so you must put the value back together with its stem before writing it down, which is a common place to lose marks.

  • How do you find the modal class from a stem-and-leaf diagram?

    Look for the stem with the most leaves.

    Each stem is a class interval, so the longest row of leaves is the class holding the most data.

  • What is a back-to-back stem-and-leaf diagram used for?

    Comparing two related sets of data, such as boys against girls, on one diagram instead of two.

    The two sets share a single central column of stems, with one set's leaves running out to the left and the other's to the right.

  • True or False?

    On a back-to-back stem-and-leaf diagram, the leaves on the left-hand side increase from left to right.

    False.

    Leaves always increase as you move away from the stems.

    On the left-hand side that means they get larger from right to left, which is the mirror image of the right-hand side.

  • Why is a rough version of a stem-and-leaf diagram drawn first, and when can you skip it?

    Because the data usually arrives out of order, so a rough version gets every value into the right stem before anything is arranged neatly.

    You can skip it when the data has already been given to you in order.

  • What check tells you a completed two-way table has no mistakes in it?

    The row totals and the column totals must both add up to the same overall figure in the corner.

    That is because the corner figure is the whole data set counted twice over, once across the rows and once down the columns, so a disagreement means something is wrong.

  • Define two-way table.

    A two-way table compares two different characteristics of the same group, with one set of categories across the columns and the other down the rows.

    Each cell gives the number having one particular combination of the two, such as Year 12 students who study Spanish.

  • In a two-way table, 15 children in total chose clay modelling and 13 of them were in class B. How many class A children chose clay modelling?

    2, found by taking the class B figure away from the column total, 15 - 13 = 2 children.

    Every empty cell can be reached this way, by subtracting the known entries in a row or column from that row or column's total.

  • Complete the sentence about reading a Venn diagram.

    The region where two circles overlap counts the items in set A \_\_\_\_\_\_ set B, while the two circles taken together count the items in set A \_\_\_\_\_\_ set B.

    The completed sentence is:

    The region where two circles overlap counts the items in set A and set B, while the two circles taken together count the items in set A or set B.

    Note that 'A or B' includes everything in the overlap as well.

  • On a Venn diagram of two sets, what do the numbers outside both circles but inside the rectangle represent?

    The items that are in neither set.

    The rectangle stands for the entire data set, so everything is somewhere inside it, and whatever sits outside both circles has neither characteristic.

  • A Venn diagram shows 12 in A only, 4 in the overlap, 21 in B only and 8 outside both circles. How many items are in set B?

    25, which is the 4 in the overlap plus the 21 in B only.

    The whole of circle B counts as set B, including the part it shares with A, and the complete data set here is 12 + 4 + 21 + 8 = 45 items.

  • True or False?

    The numbers written in the regions of a Venn diagram are always frequencies.

    False.

    The regions can hold percentages instead, in which case every number in the diagram has to add up to 100.

    Whatever they hold, the numbers across the whole diagram must account for the complete data set.

  • Define population pyramid.

    A population pyramid is a diagram showing how a population is spread across age groups, with one horizontal bar for each group.

    It is usually split down the middle by gender, with females on one side and males on the other.

  • What can the horizontal scale of a population pyramid show?

    Either the actual number of people in each age group, or, more often, a percentage.

    Where it is a percentage it may be of the total population or of that gender's population alone, so the labelling has to be read carefully before anything is calculated.

  • A population pyramid gives 5.4% of the population as female aged 0 to 9, and 7.4% as female aged 10 to 19. What percentage is female and under 20?

    Add the two percentages together, giving 5.4 + 7.4 = 12.8 percent of the population.

    Age groups can be combined like this because every person falls into exactly one of them.

  • True or False?

    An age group showing 0.0% on a population pyramid must contain nobody at all.

    False.

    When percentages are rounded to one decimal place, any group holding less than 0.05% of the population is shown as 0.0%.

    There may still be people in it.

  • On a population pyramid, how can you tell at a glance whether more males or more females are aged 40 and over?

    Compare the two sides bar by bar, from the 40 to 49 group upwards.

    If every male bar in that range is longer than the female bar beside it, there are more males aged 40 and over, with no need to add anything up.

  • Define choropleth map.

    A choropleth map shows data for the different regions of a geographical area by shading or patterning each region according to its value.

    Regions with similar values are given the same shade, so patterns across the whole area stand out at a glance.

  • How is a choropleth map shaded when colour is not available?

    With different depths of grey, running from black for the densest values through to white for the least dense.

    Patterns such as cross-hatching can be used in the same way.

  • True or False?

    A choropleth map must be an accurate geographical map of the area it represents.

    False.

    An idealised or abstract diagram can be used instead, such as a field divided into a grid of equal squares.

    What matters is that each region carries a shade matching its own value.

  • Why does a choropleth map need a key?

    Because a shade or a pattern means nothing on its own.

    The key is what turns each shade into the range of values it stands for, so without one no region can be given a number.

  • A field at a concert is divided into 30 equal squares, and on the choropleth map the darkest squares all lie along one edge. What does that suggest?

    That the crowd was densest along that edge, since the darkest shading stands for the largest numbers of people.

    The stage was most likely on that side of the field, because that is where people gather.

  • Define cumulative frequency.

    Cumulative frequency is a running total of the frequencies, added up as you work down the table.

    Each row's figure is the number of data values up to and including the end of that row's class.

  • A cumulative frequency table gives 14 for x < 20 and 39 for x < 40 in a survey. What is the frequency of the class 20 \le x < 40 on its own?

    25, found by subtracting one cumulative total from the next, 39 - 14 = 25 values.

    A cumulative total includes everything below it, so the difference between consecutive rows recovers the frequency of a single class.

  • When is a cumulative frequency step polygon used instead of a cumulative frequency diagram?

    When the data is discrete, such as the number of eggs in a nest.

    A cumulative frequency diagram is used for grouped continuous data instead.

  • Why does a cumulative frequency step polygon rise in vertical jumps rather than along a smooth curve?

    Because the data is discrete, so nothing at all happens between one possible value and the next.

    The running total stays flat across the gap and then jumps at each value that actually occurs.

  • At which point of each class interval is the cumulative frequency plotted on a cumulative frequency diagram, and why?

    At the upper bound, the end of the class interval.

    The values in a class could be anywhere within it, so they cannot all be guaranteed to have been counted until you reach the top of that class.

  • A cumulative frequency diagram is drawn for times starting from the class 25 \le s < 30 upwards. Which extra point has to be added, and where?

    A point at the very start, placed at the lowest value in the table with a cumulative frequency of zero.

    Here that is the point \left(25 , 0\right) on the graph, because nobody had finished before 25 seconds.

  • What shape does a cumulative frequency diagram always have?

    A stretched S-shape that only ever rises.

    Because every figure on it is a running total, the curve can level off but can never come back down towards the horizontal axis.

  • Complete the quartile positions used to read a cumulative frequency diagram holding n values in total.

    \text{lower quartile at } \frac{n}{\_\_\_\_\_\_} \text{, upper quartile at } \frac{\_\_\_\_\_\_ n}{4}

    The completed positions are:

    \text{lower quartile at } \frac{n}{4} \text{, upper quartile at } \frac{3 n}{4}

  • True or False?

    On a cumulative frequency diagram of 60 values, the median is found halfway between the 30th and 31st values.

    False.

    From a cumulative frequency diagram you read straight across from the value \frac{n}{2} on the vertical axis, which is 30 here.

    Halfway between the 30th and 31st is the method for a list of raw values, not for a diagram.

  • How do you find the 10th percentile from a cumulative frequency diagram?

    Work out its position as \frac{n p}{100} where p is the percentile wanted, so with 60 values the 10th percentile sits at position 6.

    Read across from 6 on the cumulative frequency axis to the curve, then straight down to the horizontal axis.

  • 100 phone calls are shown on a cumulative frequency diagram, and reading up from 12 minutes gives 90. How many calls lasted longer than 12 minutes?

    About 10: the 90 counts the calls up to 12 minutes, so subtract it from the total, 100 - 90 = 10 calls.

    A cumulative frequency reading always gives the number below a value, so any 'more than' question needs subtracting from the total.

  • Why are the median and quartiles read from a cumulative frequency diagram only estimates?

    Because the diagram is built from grouped data, so the original values are unknown.

    Joining the points assumes the data is spread evenly across each class, which will not be exactly true.

  • Which five values do you need in order to draw a box plot?

    The five values are:

    • the lowest value that is not an outlier

    • the lower quartile

    • the median

    • the upper quartile

    • the highest value that is not an outlier

  • Complete the sentence about the parts of a box plot.

    The box covers the middle \_\_\_\_\_\_ of the data, and each whisker covers a further \_\_\_\_\_\_ of it.

    The completed sentence is:

    The box covers the middle 50% of the data, and each whisker covers a further 25% of it.

    The box runs from the lower quartile across to the upper quartile.

  • True or False?

    The box on a box plot can be drawn to any depth.

    True.

    The depth of the box carries no information at all and can be drawn to suit the page.

    Everything the box tells you is in where it sits along the scale and how wide it is, which is the interquartile range.

  • True or False?

    The median line always sits in the middle of the box on a box plot.

    False.

    The median line is drawn at the median value on the scale, which is usually not halfway between the two quartiles.

    Where it sits tells you whether the data bunches towards the lower quartile or the upper one.

  • How is an outlier shown on a box plot, and what happens to the whisker?

    An outlier is marked with a cross beyond the end of a whisker.

    The whisker then stops at the last value that is not an outlier, rather than running all the way out to the outlier itself.

  • What two features of two data sets do box plots let you compare?

    An average, using the medians, and a spread, using the interquartile ranges.

    A smaller interquartile range means the values are more tightly bunched, so that data set is the more consistent of the two.

  • Why are box plots useful for data containing a few extreme values, such as the prices of cars?

    Because they split the data at the quartiles, so the bulk of it still shows clearly even when a few values sit far out.

    One sports car among forty-nine family cars appears as an outlier rather than distorting the whole picture.

  • A box plot has a lower quartile of 2 and an upper quartile of 7.5. What is the interquartile range?

    The interquartile range is 7.5 - 2 = 5.5 on the scale, which is exactly the width of the box.

    It measures the spread of the middle half of the data.

  • A bar chart has gaps between its bars. Why does a histogram not?

    Because a histogram shows continuous data grouped into class intervals, and those intervals run straight into one another with no values missing in between.

    The bars therefore touch, with each bar spanning its own class interval along the horizontal axis.

  • True or False?

    The vertical axis of a histogram can simply be labelled 'frequency' when all the class widths are equal.

    True.

    With equal class widths the height of each bar is the frequency for that class, so nothing more complicated is needed.

    Frequency density is only required where the class widths are unequal.

  • Complete the sentence about reading a histogram with equal class widths.

    The class boundaries are read off the \_\_\_\_\_\_ axis, and the height of each bar on the \_\_\_\_\_\_ axis gives the frequency for that class.

    The completed sentence is:

    The class boundaries are read off the horizontal axis, and the height of each bar on the vertical axis gives the frequency for that class.

  • At which point of each class interval is a frequency polygon plotted, and why?

    At the midpoint of the class interval.

    The exact values in a class are unknown, so the midpoint is the best single value to stand for the whole class, and for a class of 120 \le t < 150 that point sits at 135.

  • When drawing a frequency polygon, which two lines must you not draw?

    Do not join the polygon down to the horizontal axis, unless one of the frequencies really is 0.

    And do not join the last point back to the first, so the shape is left open rather than closed, which is why it is not really a polygon at all.

  • What must be true before two histograms can be used to compare two data sets?

    They must use the same class intervals and the same frequency scale.

    Otherwise the bars are not measuring the same things, and comparing their heights tells you nothing reliable.

  • What three things should you consider when choosing how to represent a data set?

    The target audience, since a general audience needs a simpler visualisation than a group of experienced statisticians would.

    The nature of the data, since some diagrams suit discrete data and others suit continuous, grouped or bivariate data.

    The strengths and weaknesses of the representation itself, since each one makes some features clear and hides others.

  • Complete the two rules about matching a diagram to the data.

    Scatter diagrams are appropriate for \_\_\_\_\_\_ data, and histograms are appropriate for \_\_\_\_\_\_ data.

    The completed rules are:

    Scatter diagrams are appropriate for bivariate data, and histograms are appropriate for grouped data.

  • What does a table of data do well, and what does it do badly?

    A table holds the exact values, so nothing is rounded, estimated or lost.

    What it does badly is show patterns and trends, which is the job a chart or graph does much better.

  • Two data sets are given in different formats, one as a pie chart and one as a bar chart. What must you do before comparing them?

    Put both into the same form first, usually by working out a proportion or percentage from each.

    A bar chart's frequencies become percentages by dividing each one by the total, and those then compare directly with the shares shown on the pie chart.

  • Define misrepresentation in a statistical diagram.

    A misrepresentation is a diagram that gives a misleading or incorrect impression of the data.

    It may be the result of an honest mistake, or it may have been done deliberately in order to mislead.

  • True or False?

    A bar chart whose vertical scale does not start at zero is wrong.

    False.

    A scale does not have to start at zero, and there is often a sensible reason for it not to.

    But it can be misleading, because the differences between the bars then look far larger than the differences in the data really are.

  • Apart from a scale not starting at zero, name three features that can make a statistical diagram misleading.

    Scales that do not go up in equal steps, or that have parts missed out, which distorts the size and shape of everything plotted against them.

    Unlabelled axes or a missing key, which leave no way to tell what the diagram is showing.

    Bright colours that make some parts stand out more than others, and lines drawn so thick that the graph cannot be read precisely.

  • Why can a 3D pie chart be misleading?

    The tilt distorts the angles, so the sectors no longer look like the shares they actually represent.

    Sectors at the front appear larger than they are while those at the back appear smaller or are hidden, and a sector pulled out of the pie is hard to compare with the rest.

  • Why might someone deliberately choose a misleading way of presenting data?

    To make a change look larger or smaller than it really is, so that the diagram supports the case they are making.

    A company showing pay rises on a scale that does not start at zero makes the increases look far more generous than they were.

Sign up to unlock flashcards

or