Exam code: 9MA0
1/90Still learning
Know0
Define bivariate data.
Bivariate data is data collected on two variables, to look at how one affects the other.
Each value of one variable is paired with a value of the other, because both are recorded from the same item or person.
That pairing is what makes a scatter diagram possible, since each pair gives one point; the two variables are often related, but they do not have to be.

Join for free to unlock a full flashcard set, track what you know,
and turn revision into real progress.
Complete the names of the two variables on a scatter diagram:
The variable that can be controlled is the or independent variable, and is plotted on the
-axis.
The variable that is measured is the or dependent variable, and is plotted on the
-axis.
The completed names are:
The variable that can be controlled is the explanatory or independent variable, and is plotted on the -axis.
The variable that is measured is the response or dependent variable, and is plotted on the -axis.
So, for example, in a study of question packs completed against exam percentage, the number of packs is chosen by the student and goes on the -axis, while the percentage is the outcome and goes on the
-axis.
What two things must you say when describing the correlation on a scatter diagram?
Say the direction and the strength, and both are needed for full marks.
The direction is positive if both variables increase together, and negative if one increases as the other decreases.
The strength is strong if the points lie close to a straight line, and weak if they are scattered about it; if there is no pattern at all, say there is no correlation and do not describe a direction.
Was this flashcard helpful?
Define bivariate data.
Bivariate data is data collected on two variables, to look at how one affects the other.
Each value of one variable is paired with a value of the other, because both are recorded from the same item or person.
That pairing is what makes a scatter diagram possible, since each pair gives one point; the two variables are often related, but they do not have to be.
Complete the names of the two variables on a scatter diagram:
The variable that can be controlled is the or independent variable, and is plotted on the
-axis.
The variable that is measured is the or dependent variable, and is plotted on the
-axis.
The completed names are:
The variable that can be controlled is the explanatory or independent variable, and is plotted on the -axis.
The variable that is measured is the response or dependent variable, and is plotted on the -axis.
So, for example, in a study of question packs completed against exam percentage, the number of packs is chosen by the student and goes on the -axis, while the percentage is the outcome and goes on the
-axis.
What two things must you say when describing the correlation on a scatter diagram?
Say the direction and the strength, and both are needed for full marks.
The direction is positive if both variables increase together, and negative if one increases as the other decreases.
The strength is strong if the points lie close to a straight line, and weak if they are scattered about it; if there is no pattern at all, say there is no correlation and do not describe a direction.
Two variables show strong correlation. Why is that not enough to say that one causes the other?
Correlation says only that the two variables move together; it says nothing about why, and both could be driven by something else entirely, or the pairing could be a coincidence.
A causal relationship is one where a change in one variable genuinely produces the change in the other, and deciding whether there is one means looking at the context rather than at the data.
So, for example, temperature and ice cream sales in a park is plausibly causal, while global temperature and the number of monkeys kept as pets in the UK may well correlate, with plainly neither causing the other.
What does "least squares" mean in the least squares regression line?
It means the line is positioned to make the sum of the squares of the gaps between the line and the data points as small as possible.
The gaps are squared so that points above and below the line both count as errors rather than cancelling out, and so that a point a long way off counts for much more than several small misses.
For the regression line of on
, those gaps are measured vertically, which is why that line is used to predict
from
.
In the regression line , what does each of
and
tell you?
is the gradient, the change in
for each increase of one unit in
, and its sign also tells you the direction of the correlation without needing to see a diagram.
is the
-intercept: the value of
when
is zero.
So, for example, in for question packs completed against exam percentage, the score rises by 1.3 percentage points for each extra pack, starting from 18% with no packs completed.
True or False?
A larger value of in the regression line
means stronger correlation.
False.
is the gradient, so it tells you how fast
changes as
changes, not how closely the points sit to the line.
Data can climb very steeply while being scattered all over the place, giving a large and weak correlation, or climb gently with every point almost exactly on the line, giving a small
and very strong correlation.
Strength of correlation is measured separately, by the product moment correlation coefficient.
What is the difference between interpolation and extrapolation, and which of them should you avoid?
Interpolation is predicting from a value of the explanatory variable that lies inside the range of the data the line was calculated from, and it is reasonable.
Extrapolation is predicting from a value outside that range, and it should be avoided, because the line was fitted to a particular stretch of data and says nothing about what happens beyond it.
So, for example, a line fitted to between 10 and 60 question packs should not be used to predict the score for 80 packs; any prediction is more reliable when the original sample was larger.
Which point always lies on the regression line of on
?
The point , whose coordinates are the mean of the
values and the mean of the
values.
That is worth knowing for two reasons: it gives you a point to plot when drawing the line, and it lets you find a missing mean if you are given the equation of the line and the other mean.
By signing up you agree to our Terms and Privacy Policy