FastPrepStochastic Gradient Descent with Momentum

Stochastic Gradient Descent with Momentum

Meta logoMeta● MediumFULLTIMEONSITE INTERVIEW
Learn

Problem statement

Simulate stochastic gradient descent with momentum for one parameter vector. Start with zero velocity. For each gradient row in order, update every coordinate using:

  1. velocity = momentum * velocity + gradient
  2. weight = weight - learningRate * velocity

Return a two-row matrix: the final weights followed by the final velocity. Do not modify the input arrays.

Function

sgdWithMomentum(initialWeights: double[], gradients: double[][], learningRate: double, momentum: double) → double[][]

Examples

Example 1

initialWeights = [1,2]gradients = [[0.5,-1],[1,1]]learningRate = 0.1momentum = 0.9return = [[0.805,2.09],[1.45,0.1]]

After the first step the velocity is [0.5,-1] and weights are [0.95,2.1]. The second velocity is [1.45,0.1].

Example 2

initialWeights = [0]gradients = [[2],[2]]learningRate = 0.5momentum = 0return = [[-2],[2]]

With zero momentum, each step uses only the current gradient.

Example 3

initialWeights = [3,-4]gradients = [[0,0]]learningRate = 0.25momentum = 0.5return = [[3,-4],[0,0]]

A zero gradient and zero initial velocity leave the weights unchanged.

Constraints

  • 1 <= initialWeights.length <= 100.
  • 1 <= gradients.length <= 1000, and every gradient row has the parameter dimension.
  • Every initial weight and gradient value is finite and has absolute value at most 10^6.
  • 0 < learningRate <= 1 and 0 <= momentum < 1.

More Meta problems

See Meta hiring insights
public double[][] sgdWithMomentum(double[] initialWeights, double[][] gradients, double learningRate, double momentum) {
    // Write your code here.
}
initialWeights[1,2]
gradients[[0.5,-1],[1,1]]
learningRate0.1
momentum0.9
expected[[0.805,2.09],[1.45,0.1]]
Checking account…