Further Correlation & Regression (AQA A Level Maths: Statistics): Flashcards

Exam code: 7357

1/14

0Still learning

Know0

  • Define the product moment correlation coefficient.

Cards in this collection (14)

  • Define the product moment correlation coefficient.

    The product moment correlation coefficient, written r, is a number measuring how close the points of a set of bivariate data lie to a straight line.

    It measures linear correlation only, so data following a strong curved pattern can still give a value of r close to zero.

  • Complete the range of possible values of the product moment correlation coefficient:

    \_\_\_\_\_\_ \le r \le \_\_\_\_\_\_

    The completed range is:

    - 1 \le r \le 1

    A value of r = 1 means perfect positive correlation and r = - 1 means perfect negative correlation, with every point lying exactly on the line.

    Any value outside this range is impossible, so a coefficient given as 1 . 05 or - 1 . 2 is definitely wrong.

  • Two scatter diagrams each have every point lying exactly on a straight line of positive gradient, one steep and one shallow. What is r for each?

    Both have r = 1.

    The value of r measures only how close the points lie to a straight line, so once they lie exactly on one it has reached its largest value, however steep or shallow that line happens to be.

  • What does a value of r = 0 . 1652 tell you about the data?

    A value that close to zero means there is almost no linear correlation: the points are scattered with no straight-line pattern.

    How far r lies from zero is what matters, not simply whether it is positive, and 0 . 1652 is far nearer to 0 than to 1.

    By contrast r = 0 . 8134 would be fairly strong positive correlation, and r = - 0 . 9993 almost perfect negative correlation.

  • Why take logarithms of data that fits y = a x^{n} or y = k b^{x}?

    Because taking logarithms turns either of those relationships into a straight line, which is far easier to work with than a curve.

    Once the data is plotted as a straight line you can read off its gradient and its vertical intercept, and those two numbers give you the constants in the original model.

    Doing this is called changing the variables, or coding the data.

  • Complete the result of taking logarithms of both sides of y = a x^{n}:

    \log y = \_\_\_\_\_\_ + n \log \_\_\_\_\_\_

    The completed result is:

    \log y = \log a + n \log x

    So plotting \log y against \log x gives a straight line whose gradient is n and whose vertical intercept is \log a, not a itself.

    The addition law turns \log \left(a x^{n}\right) into \log a + \log x^{n}, and the power law then turns \log x^{n} into n \log x.

  • For y = k b^{x}, what do you plot against what, and what do the gradient and intercept give you?

    Plot \log y against x itself, which gives \log y = \log k + x \log b, a straight line.

    The gradient is \log b and the vertical intercept is \log k, so b and k are recovered by raising the base to each of those.

    Here x is not logged, because it sits in the exponent rather than in the base, and that is what separates this case from y = a x^{n}.

  • A coded line is Y = - 3 . 5 X, where X = \log h and Y = \log t. How do you recover the model connecting t and h?

    Substitute \log h for X and \log t for Y to get \log t = - 3 . 5 \log h, then use the power law to write the right-hand side as \log h^{- 3 . 5}.

    Both sides are now a single logarithm to the same base, so raising 10 to each side removes them and leaves t = h^{- 3 . 5}.

    Anything found from coded data comes out in terms of X and Y, so this reversing step is what turns it back into a statement about the original quantities.

  • True or False?

    Coding with natural logarithms instead of base 10 logarithms still produces a straight line.

    True.

    The laws of logarithms hold in any base, so \ln y = \ln a + n \ln x is just as straight a line as the base 10 version.

    What the base does change is how you reverse the coding: a value coded with \log is undone by raising 10 to it, and one coded with \ln by raising \text{e} to it.

  • Why is a hypothesis test needed to say anything about correlation in a whole population?

    Because finding the correlation coefficient of a whole population would mean collecting data on every individual in it, which is almost never possible.

    The coefficient is calculated from a sample instead, and the test asks whether a sample value that far from zero is good enough evidence that the population really is correlated.

  • In a hypothesis test for correlation, why are the hypotheses written in terms of \rho rather than r?

    Because \rho is the correlation coefficient of the whole population, and a hypothesis is always a claim about the population rather than about the sample in front of you.

    The sample coefficient r is the evidence used to test that claim, not the thing being claimed, so it belongs in the comparison rather than in the hypotheses.

    Writing \text{H}_{0} : r = 0 says the sample coefficient is zero, which is not what is being tested: you can already see r, so there is nothing left to hypothesise about it.

  • Complete the hypotheses for a test of whether a population is negatively correlated:

    \text{H}_{0} : \rho = \_\_\_\_\_\_ \text{ and } \text{H}_{1} : \rho \_\_\_\_\_\_ 0

    The completed hypotheses are:

    \text{H}_{0} : \rho = 0 \text{ and } \text{H}_{1} : \rho < 0

    The null hypothesis is always \rho = 0, whatever the test: it is the claim that there is no linear correlation in the population at all.

    Only the alternative changes, to \rho < 0 or \rho > 0 when a direction is being tested, and to \rho \neq 0 when any correlation at all would do.

  • True or False?

    A test for negative correlation gives r = - 0 . 45 against a critical value of 0 . 2787, so the result is not significant.

    False.

    Compare the two by size: 0 . 45 is greater than 0 . 2787, so r does lie in the critical region and the result is significant.

    Because the test is for negative correlation the critical value is really - 0 . 2787, and - 0 . 45 lies beyond it; deciding instead that - 0 . 45 is merely smaller than 0 . 2787 is the trap here.

  • A sample gives r = 0 . 3. Why is that not enough to say the population is correlated?

    Because a sample drawn from a population with no correlation at all will still give a value of r that is not exactly zero, simply through the randomness of which items happened to be picked.

    The question is therefore not whether r differs from zero, but whether it differs by more than chance would comfortably produce, and the critical value is where that line is drawn.

Sign up to unlock flashcards

or