Stochastic Gradient Descent with Momentum
Problem statement
Simulate stochastic gradient descent with momentum for one parameter vector. Start with zero velocity. For each gradient row in order, update every coordinate using:
velocity = momentum * velocity + gradientweight = weight - learningRate * velocity
Return a two-row matrix: the final weights followed by the final velocity. Do not modify the input arrays.
Function
sgdWithMomentum(initialWeights: double[], gradients: double[][], learningRate: double, momentum: double) → double[][]Examples
Example 1
initialWeights = [1,2]gradients = [[0.5,-1],[1,1]]learningRate = 0.1momentum = 0.9return = [[0.805,2.09],[1.45,0.1]]After the first step the velocity is [0.5,-1] and weights are [0.95,2.1]. The second velocity is [1.45,0.1].
Example 2
initialWeights = [0]gradients = [[2],[2]]learningRate = 0.5momentum = 0return = [[-2],[2]]With zero momentum, each step uses only the current gradient.
Example 3
initialWeights = [3,-4]gradients = [[0,0]]learningRate = 0.25momentum = 0.5return = [[3,-4],[0,0]]A zero gradient and zero initial velocity leave the weights unchanged.
Constraints
1 <= initialWeights.length <= 100.1 <= gradients.length <= 1000, and every gradient row has the parameter dimension.- Every initial weight and gradient value is finite and has absolute value at most
10^6. 0 < learningRate <= 1and0 <= momentum < 1.