Predicting at x* and the limits of the line
◈ 5 cardsAt x* the interval for the MEAN response uses s√(1/n + (x* − x̄)²/SS_xx); the prediction interval for ONE new unit adds 1 inside the root and is always wider; both widen away from x̄, and neither rescues a line pushed outside its data.
Two questions at the same x*
Module 4 predicted for a client earning $75,000 and left it as a point. Two different questions hide behind that number:
- What is the average credit score of all clients earning $75,000? — the mean response $\mu_{y \mid 75}$, a parameter on the population line.
- What score will this one new client have? — a single future observation, the line’s value plus that client’s own .
Both intervals are centred on and both use ; they differ only in the standard error.
The extra 1 is the whole difference. The first root measures how uncertain the line is at (the intercept and slope were estimated); the second adds the scatter of an individual around the line — one full — because a single client does not sit on the line even when the line is known exactly.
Worked example — Northfield at $75,000
, , , , , .
The shared piece: .
Mean response. .
With 95 % confidence, the average score of Northfield clients earning $75,000 is between 690.6 and 707.7.
Prediction. .
A new client earning $75,000 will have a score between 674 and 724, with 95 % confidence. Three times wider — the individual’s own scatter dominates the line’s uncertainty here, as it does whenever $n$ is small and $x^*$ is near $\bar{x}$.
Both widen away from x̄
The term is zero at and grows quadratically away from it: the line is pinned best at the centre of the data and swings more at the ends, so both intervals are narrowest at and flare outward. At , the edge of the data, the mean-response half-width is already instead of 8.6.
The limits of the line
The intervals are honest only inside the data. Three reasons a lender should not set credit limits from :
- . Every interval above is wide because the line rests on eight points; the slope itself is only known to ±0.35.
- Extrapolation. Incomes ran from $35,000 to $110,000. Outside that range the intervals are formulas, not evidence — nothing in the data says the relationship stays linear.
- One predictor, observational data. Income is not the only driver of a credit score — age, existing debt, payment history — and the line cannot separate income’s association from theirs. A significant slope (L13.2) is association, not a mechanism to lend against.