Fit a line by hand. Then let gradient descent update one parameter. Keep the data fixed and change only the learning rate.
Data and predictions
Loss during this trial
The curve begins at update 0. Moving a or changing the learning rate starts a new trial.
Try to reduce MSE. Manual control covers −2 ≤ a ≤ 8; training may move outside this range.
Changing α resets a to 0 for a fair comparison. Each update uses all four training points.
25-minute experiment
- Fit by hand · 5 min. Move a. Find a line with low MSE. What changes on the plot?
- One update · 5 min. Reset a = 0. With α = 0.05, predict the direction of the next change, then make one update.
- Compare · 10 min. For each α = 0.01, 0.05, 0.20, start at a = 0, run exactly 20 updates, and record the result below.
- Explain · 5 min. Which setting reduced the error quickly? Did taking larger steps always help?
The same data as Lecture 3
| x | 1 | 2 | 3 | 4 |
|---|---|---|---|---|
| y | 4 | 6 | 11 | 14 |
Illustrative data. Our model is ŷ = ax, with no intercept. There is one parameter to learn: a. All four points are used for training.
What does the computer calculate?
L(a) = [(a − 4)² + (2a − 6)² + (3a − 11)² + (4a − 14)²] / 4
L′(a) = 15a − 52.5
anew = a − α(15a − 52.5)
For this dataset and this mean-squared loss, the gradient tells us which way the loss increases locally. Subtracting a small multiple of it can reduce the loss.
What this experiment does not test
A lower training error means a better fit to these four points. It does not establish performance on new data. We will use separate evaluation data in a later lesson.
Your comparison
Record each learning rate after 20 updates from a = 0. Compare rows with the same starting a and update count.
| α | Starting a | Updates | Final a | Training MSE | Action |
|---|