Aller au contenu principal
Ecolint Campus des NationsMathématiques
Ecolint Campus des NationsMathematics
Year 10 · Bivariate Statistics

Solutions · Full Answer Key

Pack A answers · Pack B answers · Problem-solving worked solutions

Pack A — Answers

Bronze
1.Positive correlation
2.No correlation (or zero)
3.Strong positive correlation
4.(a) height (x-axis) (b) weight (y-axis)
5.Positive correlation: more sunshine, more sales
6.17
7.Gradient 4; y-intercept 7
8.No — both depend on a third variable (hot weather)
9.A student who scored higher than predicted for their revision hours
10.y=21y = 21
Silver
11.y=3x+3y = 3x + 3
12.124 cm
13.Each extra hour of revision adds 5 marks (on average)
14.Predicted mark with zero hours of revision = 40
15.Strong positive correlation
16.x=6x = 6
17.yˉ=19\bar{y} = 19
18.x=10x = 10: interpolation; x=30x = 30: extrapolation
19.Plot 1 (rr closer to 1)
20.Each hour of gaming predicts a 2-mark decrease; with zero gaming the predicted score is 70
Gold
21.y=2x+1y = 2x + 1
22.(a) 35 (b) 60 (extrapolation — unreliable)
23.Strong negative: more mileage → lower value
24.(a) Increase ∣r∣|r| — removes scatter (b) Outliers reduce correlation strength
25.y=2x+2y = 2x + 2
26.(a) Weak positive (b) Strong positive (c) Very strong negative (d) None
27.(4,11)(4, 11)
28.No — confounding variable (national wealth) drives both
29.Residual = −2-2 (observed minus predicted)
30.T=3°T = 3°C
Platinum
31.64% (r2=0.64r^2 = 0.64)
32.B is better fit (higher ∣r∣|r|)
33.70% of variation explained — moderate-strong predictive value
34.1.5
35.Strong positive rank agreement
36.There IS a relationship — but not linear
37.Extrapolation — and negative sales is nonsensical
38.There is a fairly strong negative relationship; the model explains about 49% of the variation
39.(a) 34 mpg (b) 10 mpg — likely extrapolation
40.No — strong tendency, not certainty

Pack B — Answers

Bronze
1.Negative correlation
2.No correlation
3.Strong negative correlation
4.(a) hours studied (b) test score
5.Positive correlation: more revision, higher score
6.26
7.Gradient −2-2; y-intercept 15
8.No — both depend on size of fire
9.A student who scored lower than predicted
10.y=21y = 21
Silver
11.y=3x+4y = 3x + 4
12.120 cm
13.Each extra hour of revision adds 8 marks (on average)
14.Predicted mark with zero revision = 50
15.Moderate negative correlation
16.x=9x = 9
17.yˉ=22\bar{y} = 22
18.x=4x = 4: interpolation; x=20x = 20: extrapolation
19.Plot 1 (∣r∣=0.85|r| = 0.85 vs 0.1)
20.Each hour of gaming predicts a 1.5-mark decrease; intercept 80
Gold
21.y=3x+4y = 3x + 4
22.(a) 40 (b) 70 (extrapolation)
23.Strong positive: older athletes have slower sprints
24.(a) Possibly decrease ∣r∣|r| (b) An aligned outlier may anchor a strong correlation
25.y=3x−3y = 3x - 3
26.(a) Weak negative (b) Strong positive (c) None (d) Moderate positive
27.(5,14)(5, 14)
28.No — population size drives both
29.Residual = 0
30.T=−6°T = -6°C
Platinum
31.36%
32.B is better fit
33.40% of variation explained — weak predictive value
34.1.5
35.Moderate negative rank agreement
36.Strong non-linear relationship hidden by rr
37.Same — extrapolation gives nonsense
38.There is a moderate positive relationship; the model explains about 25% of the variation
39.(a) 45 mpg (b) 5 mpg — extrapolation
40.No — moderate tendency only

Problem-solving — Worked Solutions

1Problem 1
Answer
(a) Strong positive (b) xˉ=7.7\bar{x} = 7.7, yˉ=61.5\bar{y} = 61.5 (c) Slope ≈ 4.3; y≈4.3x+28y \approx 4.3x + 28 (d) ≈ 45.4 (interpolation, reasonable)
Full working
(a) Strong positive correlation: marks increase steadily with hours.

(b) xˉ=(2+3+5+6+7+8+9+10+12+15)/10=77/10=7.7\bar{x} = (2+3+5+6+7+8+9+10+12+15)/10 = 77/10 = 7.7. yˉ=(35+40+50+55+60+65+70+70+80+90)/10=615/10=61.5\bar{y} = (35+40+50+55+60+65+70+70+80+90)/10 = 615/10 = 61.5.

(c) Roughly, from (2, 35) to (15, 90): gradient ≈(90−35)/(15−2)=55/13≈4.23\approx (90 - 35)/(15 - 2) = 55/13 \approx 4.23. Using the mean point: y−61.5=4.23(x−7.7)⇒y≈4.23x+28.93y - 61.5 = 4.23(x - 7.7) \Rightarrow y \approx 4.23x + 28.93. Round to y≈4.3x+29y \approx 4.3x + 29.

(d) At x=4x = 4: y≈4.3(4)+29=46.2y \approx 4.3(4) + 29 = 46.2. Since 4 is within the data range (2–15), this is **interpolation** and reasonably reliable.
2Problem 2
Answer
No — coincidence / lurking variables. The lesson: correlation does not imply causation.
Full working
(a) **No.** This is a famous example of a **spurious correlation** — two unrelated time series can correlate by chance over a small range of years.

(b) Two explanations:
- **Coincidence.** With many possible variables, some pairs will appear correlated purely by chance over short time windows.
- **Confounding by time.** Both quantities may simply increase or decrease over time (population growth, more films and pools available). The correlation reflects time, not a causal link.

(c) Correlation does **not** imply causation. Strong rr alone is not evidence of a causal mechanism; you need:
- A plausible mechanism.
- Replication across data.
- Controlling for confounding variables.

This is why scientists do **controlled experiments**, not just correlations.
3Problem 3
Answer
(a) Each 1 m up loses 0.0065°C; sea level baseline 15°C (b) 5.25°C (c) ≈ −43.5°C — probably extrapolation
Full working
(a) Gradient −0.0065-0.0065: temperature drops by 0.0065 °C per metre of altitude (i.e., about 6.5 °C per km). Intercept 15 °C: predicted temperature at sea level (a=0a = 0).

(b) T=−0.0065(1500)+15=−9.75+15=5.25T = -0.0065(1500) + 15 = -9.75 + 15 = 5.25°C.

(c) T=−0.0065(9000)+15=−58.5+15=−43.5T = -0.0065(9000) + 15 = -58.5 + 15 = -43.5°C. This is **extrapolation** way beyond the typical data range (most measurements would be from aa ≈ 0 to 4000 m). The atmosphere at 9000 m has different physics (lapse rate varies with altitude); the linear model is probably less reliable. The actual temperature near the summit of Everest is around −30°-30° to −40°-40°C — the prediction is in the right region.
4Problem 4
Answer
(a) (7,25)(7, 25) is the outlier (b) With: ≈ 2.5; Without: ≈ 2.0 (c) Outlier inflates gradient and reduces rr
Full working
(a) Plot the points: (1,3), (2,5), (3,7) ... all roughly follow y=2x+1y = 2x + 1. Point (7,25)(7, 25) is way above this trend (expected y≈15y \approx 15). **(7, 25) is the outlier.**

(b) With outlier (8 points): gradient via approximation through start and end: (16−3)/(8−1)=13/7≈1.86(16 - 3)/(8 - 1) = 13/7 ≈ 1.86. With outlier averaging, the gradient is pulled up.

Without (7, 25): remaining 7 points fit y≈2x+1y \approx 2x + 1 closely → gradient ≈ 2.

(c) The outlier:
- **Distorts** the LOBF (raises gradient slightly).
- **Reduces** rr — adds scatter.
- Could be a data-entry error or a legitimate anomaly. Always check the source before deleting.
5Problem 5
Answer
(a) Plausibly causal (sleep affects cognition) (b) Confounding (hot weather) (c) Coincidence (d) Confounding (size of fire)
Full working
(a) **Plausibly causal**: well-rested students perform better on cognitive tasks. Evidence from controlled studies supports a direct effect, though revision habits matter too.

(b) **Confounding by weather**: hot/sunny weeks mean more sunscreen *and* more swimming (hence drownings). Sunscreen doesn't cause drownings.

(c) **Coincidence**: no plausible mechanism linking London weather and a football team's results. Likely spurious.

(d) **Confounding by fire size**: bigger fires call for more trucks AND cause more damage. The trucks don't cause damage; fire size does.
6Problem 6
Answer
(a) 90 kg (b) W=0.5H+5W = 0.5 H + 5 (c) ✓ — same 90 kg
Full working
(a) W=50(1.7)+5=85+5=90W = 50(1.7) + 5 = 85 + 5 = 90 kg.

(b) Convert: Hm=Hcm/100H_{\text{m}} = H_{\text{cm}}/100. So W=50(Hcm/100)+5=0.5Hcm+5W = 50(H_{\text{cm}}/100) + 5 = 0.5 H_{\text{cm}} + 5.

(c) At Hcm=170H_{\text{cm}} = 170: W=0.5(170)+5=90W = 0.5(170) + 5 = 90 kg ✓.

**Lesson:** the gradient depends on units, but the intercept is unchanged (when the conversion is purely multiplicative on xx).
7Problem 7
Answer
(a) No, alone (b) Physical mechanism, experiments, models (c) Unlikely given mechanism evidence
Full working
(a) No — correlation alone never proves causation, no matter how strong.

(b) Strengthening evidence:
- **Physical mechanism**: CO2CO_2 absorbs infrared radiation, a known physical effect (greenhouse effect).
- **Predictive models**: climate models incorporating the greenhouse effect reproduce observed temperature changes.
- **Experimental confirmation**: lab experiments demonstrate the greenhouse effect of CO2CO_2.
- **Multiple datasets**: temperature increase verified across many independent sources.

(c) Given the strong physical mechanism *and* the consistency of the data across many measurements, coincidence is implausible. This is a case where the correlation **plus** a well-understood mechanism gives high confidence in causation.
8Problem 8
Answer
(a) Strong negative (b) y=−2x+40y = -2x + 40 (c) y=28y = 28
Full working
(a) Strong negative correlation: as xx increases, yy decreases, with relatively little scatter.

(b) Through (10,20)(10, 20) with gradient −2-2: y−20=−2(x−10)⇒y=−2x+40y - 20 = -2(x - 10) \Rightarrow y = -2x + 40.

(c) y(6)=−2(6)+40=28y(6) = -2(6) + 40 = 28.
9Problem 9
Answer
(a) r=1r = 1 (b) r≈0r ≈ 0 (c) rr around −0.7-0.7 (d) r≈0r ≈ 0
Full working
(a) r=+1r = +1 (perfect positive linear).

(b) r≈0r \approx 0 (no linear trend).

(c) Negative correlation, moderately strong — rr around −0.7-0.7 to −0.8-0.8.

(d) **r≈0r \approx 0** because rr measures *linear* association. A symmetric parabola has r=0r = 0 even though there's a clear (non-linear) relationship. **This is a key warning: a low rr doesn't rule out a relationship — only a linear one.**
10Problem 10
Answer
(a) £8000 and £3000 (b) £-7000 — nonsensical and extrapolation (c) 16 years — partly realistic (very old cars near 0 value)
Full working
(a) Age 0: P=8000P = 8000 (new car). Age 10: P=−5000+8000=3000P = -5000 + 8000 = 3000.

(b) Age 30: P=−15000+8000=−7000P = -15000 + 8000 = -7000 — a **negative** price, which is nonsensical. The model breaks down because it's a linear extrapolation outside the fitting range (0–10).

(c) Set P=0P = 0: 0=−500a+8000⇒a=160 = -500a + 8000 \Rightarrow a = 16 years. At 16 years old the model predicts zero value. Some cars do retain very low value at this age, so the prediction is **plausible** for typical cars at the boundary, but the model can't extend further: a 30-year-old classic car may even *appreciate*, not depreciate further.
11Problem 11
Answer
See working — Bea's method is more accurate and repeatable.
Full working
**Anya (by eye).**
- ✗ **Less accurate**: humans pick lines based on visual judgement; bias can creep in.
- ✗ **Not repeatable**: two students may draw slightly different lines.
- ✓ **Quick** for rough estimates.
- ✓ **Robust to outliers** if drawn carefully.

**Bea (calculator regression).**
- ✓ **More accurate**: minimises sum of squared residuals, mathematically optimal for linear fit.
- ✓ **Repeatable**: same data → same line every time.
- ✓ **Can be extended** to compute rr, r2r^2, residuals, etc.
- ✗ **Sensitive to outliers**: a single bad point can pull the line strongly.
- ✗ Requires understanding when linear regression is appropriate.

**For a research report**, use the least-squares regression. For quick sketches, by-eye is fine.
12Problem 12
Answer
(a) See sketch (b) No — growth slows over time (c) Initially rapid, then slows (sub-linear)
Full working
(a) Sketch the points. Initially they rise steeply, but the rate of increase slows down (the curve flattens).

(b) **Not entirely appropriate.** The data show a clear non-linear pattern: the gradient decreases as days increase. A linear fit would over-predict heights at early days and at late days (it would be straightest in the middle).

Computing increments per day-range:
- Days 5–20 (15 days): height +18, so +1.2 cm/day.
- Days 20–50 (30 days): height +30, so +1 cm/day.
- Days 50–80 (30 days): height +15, so +0.5 cm/day.

The rate is **clearly decreasing**.

(c) The plant grows rapidly when young (about 1.2 cm/day) and slows as it ages (down to 0.5 cm/day by days 50–80). This is sub-linear (logarithmic-style) growth — characteristic of biological growth approaching a maximum height.

**Better model**: a logistic or square-root function. A linear fit would still capture the general trend but with poor predictive accuracy near the extremes.