CSED105 · LECTURE 4

Learning a line

Fit a line by hand. Then let gradient descent update one parameter. Keep the data fixed and change only the learning rate.

ŷ = 0.0000 × x
Parameter a0.0000
Training MSE92.2500
Updates0

Data and predictions

Observed dataPredictionPrediction error

Loss during this trial

The curve begins at update 0. Moving a or changing the learning rate starts a new trial.

Try to reduce MSE. Manual control covers −2 ≤ a ≤ 8; training may move outside this range.

Changing α resets a to 0 for a fair comparison. Each update uses all four training points.

25-minute experiment

  1. Fit by hand · 5 min. Move a. Find a line with low MSE. What changes on the plot?
  2. One update · 5 min. Reset a = 0. With α = 0.05, predict the direction of the next change, then make one update.
  3. Compare · 10 min. For each α = 0.01, 0.05, 0.20, start at a = 0, run exactly 20 updates, and record the result below.
  4. Explain · 5 min. Which setting reduced the error quickly? Did taking larger steps always help?

The same data as Lecture 3

x1234
y461114

Illustrative data. Our model is ŷ = ax, with no intercept. There is one parameter to learn: a. All four points are used for training.

What does the computer calculate?

L(a) = [(a − 4)² + (2a − 6)² + (3a − 11)² + (4a − 14)²] / 4

L′(a) = 15a − 52.5
anew = a − α(15a − 52.5)

For this dataset and this mean-squared loss, the gradient tells us which way the loss increases locally. Subtracting a small multiple of it can reduce the loss.

What this experiment does not test

A lower training error means a better fit to these four points. It does not establish performance on new data. We will use separate evaluation data in a later lesson.

Your comparison

Record each learning rate after 20 updates from a = 0. Compare rows with the same starting a and update count.

αStarting aUpdatesFinal aTraining MSEAction