Further Correlation & Regression (OCR A Level Maths A: Statistics): Flashcards

Exam code: H240

1/17

0Still learning

Know0

  • Define product moment correlation coefficient.

Cards in this collection (17)

  • Define product moment correlation coefficient.

    The product moment correlation coefficient is a single number describing how close the points of a bivariate data set lie to a straight line, and it is written r.

    A positive value describes positive correlation, and a negative value describes negative correlation.

  • Complete the range of possible values of the product moment correlation coefficient:

    \_\_\_\_\_\_ \le r \le \_\_\_\_\_\_

    The completed range is:

    - 1 \le r \le 1

    Any value outside that range is impossible, which makes it a quick check on a coefficient read off a calculator or quoted in a question.

  • What do r = 1 and r = - 1 tell you about the points on a scatter diagram?

    Every point lies exactly on a straight line.

    For r = 1 that line rises as x increases, and for r = - 1 it falls; these are the only two values of r for which the fit is perfect.

  • True or False?

    A data set lying exactly on the line y = 2 x and one lying exactly on the line y = 20 x have the same value of r.

    True.

    Both have r = 1, because in each case every point lies exactly on a straight line that rises.

    The coefficient measures how close the points lie to a line, not how steep that line is, so changing the gradient does not change it at all.

  • Four bivariate data sets have product moment correlation coefficients of 0.1652, −0.7134, 0.8134 and −0.9993.

    Order them from the data set whose points lie closest to a straight line to the one whose points lie furthest from it.

    Order them by how far each value is from zero, ignoring the sign, because the sign gives the direction of the correlation rather than its strength.

    That gives −0.9993, then 0.8134, then −0.7134, then 0.1652.

    The last of those is close enough to zero that its scatter diagram would look almost patternless.

  • The points on a scatter diagram lie along a clear curve, and the product moment correlation coefficient is close to zero.

    What does that tell you?

    It tells you there is no linear relationship, which is not the same thing as no relationship at all.

    The coefficient only measures how close the points lie to a straight line, so data following a curve, such as exponential growth or decay, can give a value near zero while still showing a strong pattern.

    A straight-line regression model is therefore the wrong model for this data, even though the two variables are clearly related.

  • Define changing the variables.

    Changing the variables means replacing the original data with logarithms of it, so that a non-linear relationship becomes a linear one that a regression line can be fitted to.

    The coded variables are usually written X and Y in capitals to keep them apart from the original x and y values, and the process is also called coding the data.

  • A bivariate data set is thought to fit the model y = a x^{n} with a and n both constant.

    Which two logarithms are plotted against each other, and what do the gradient and intercept give you?

    Plot \log y on the vertical axis against \log x on the horizontal axis, because taking logarithms of both sides gives \log y = n \log x + \log a as the relationship.

    That is a straight line whose gradient is n and whose vertical intercept is \log a instead of a itself.

    A plot of \log y against \log x coming out straight is itself the evidence that a power model fits the data.

  • A bivariate data set is thought to fit the model y = k b^{x} with k and b both constant. Complete the equation that taking logarithms of both sides gives:

    \log y = \_\_\_\_\_\_ + x \_\_\_\_\_\_

    The completed equation is:

    \log y = \log k + x \log b

    Comparing that with the equation of a straight line shows that \log y is plotted against x itself rather than against \log x as before, giving a gradient of \log b and an intercept of \log k here.

    Which variables get logged is what separates the two models: a power model needs both of them logged, an exponential model only the y values.

  • For a data set coded with X = \log h and Y = \log t the regression line comes out as Y = - 3 . 5 X with no intercept.

    How do you recover the model relating the two original quantities?

    Substitute the codings back in to get \log t = - 3 . 5 \log h and then undo the logarithms.

    The power law turns - 3 . 5 \log h into \log \left(h^{- 3 . 5}\right) so that both sides are logarithms of a single quantity, and the model is therefore t = h^{- 3 . 5} with its constant equal to 1.

    Check which base was used before undoing anything, because a \log is reversed with a power of 10 and an \ln with a power of \text{e} instead.

  • Why is a hypothesis test needed to say anything about correlation in a population?

    Because finding the correlation coefficient of a whole population would mean collecting data on every individual in it, which a statistician rarely has the time or resources to do.

    Instead the coefficient is found from a sample, and the test asks whether a value that far from zero is good enough evidence that the whole population is correlated too.

  • Complete the two symbols used for the product moment correlation coefficient:

    The coefficient of a whole population is written \_\_\_\_\_\_ and the coefficient calculated from a sample is written \_\_\_\_\_\_ instead.

    The completed sentence is:

    The coefficient of a whole population is written \rho and the coefficient calculated from a sample is written r instead.

    \rho is the lower-case Greek letter rho, and Greek letters are used for population values throughout statistics while ordinary letters are used for sample values, which is the same convention as \mu and \bar{x} for a mean.

  • What are the null and alternative hypotheses for a hypothesis test on correlation?

    The null hypothesis is always \text{H}_{0} : \rho = 0, that there is no linear correlation in the population.

    The alternative hypothesis depends on the test: \text{H}_{1} : \rho > 0 or \text{H}_{1} : \rho < 0 for a one-tailed test, and \text{H}_{1} : \rho \neq 0 for a two-tailed test.

  • You are given the critical value for a hypothesis test on correlation. How do you decide whether to reject the null hypothesis?

    Compare the size of the sample coefficient with the size of the critical value, ignoring signs: the result is significant when \vert r \vert > \vert \text{critical value} \vert.

    A significant result means r lies in the critical region, and the null hypothesis is then rejected in favour of the alternative.

  • True or False?

    If the sample value of r does not lie in the critical region, the test has proved that there is no correlation in the population.

    False.

    A value outside the critical region means only that the sample gives insufficient evidence to reject the null hypothesis at that significance level.

    The test is run on one sample, and a larger sample or a different significance level could reach the opposite conclusion, so nothing is ever proved either way.

  • A table of critical values is given in a question on correlation. What must you check before reading a value from it?

    Find the row for your sample size n, and then the column for the significance level of your test.

    The table usually carries two rows of column headings, one for one-tailed tests and one for two-tailed tests, so check which kind of test you are running: the same column can be headed 5% for a one-tailed test and 10% for a two-tailed one.

  • In a test on correlation, what does the critical value tell you about a probability?

    It is the value of r for which, if the population coefficient really were zero, the probability of a sample giving a value at least that extreme is exactly the significance level of the test.

    So in a one-tailed test at the 1% level, a critical value of 0.7155 means a sample drawn from an uncorrelated population would give r of 0.7155 or more only 1% of the time.

    That is why a value beyond it counts as evidence against the null hypothesis: it would be a surprising thing to see by chance.

Sign up to unlock flashcards

or