Linear Layer Forward and Backward Passes
Problem statement
Implement the forward and backward passes of a single-output linear layer under half mean-squared error.
For each sample i, compute prediction[i] = bias + sum_j(features[i][j] * weights[j]). The scalar loss is sum_i((prediction[i] - targets[i])^2) / (2 * B), where B is the sample count.
Return a matrix whose rows may have different lengths:
- Row
0: all predictions. - Row
1: the single loss value. - Row
2: the gradient with respect to the weights. - Row
3: the single gradient with respect to the bias. - Row
4 + i: the gradient with respect to feature rowi.
Return analytic, unrounded values and do not update the parameters.
Function
linearLayerForwardBackward(features: double[][], weights: double[], bias: double, targets: double[]) → double[][]Examples
Example 1
features = [[1,2],[3,4]]weights = [2,-1]bias = 0.5targets = [0,1]return = [[0.5,2.5],[0.625],[2.5,3.5],[1.0],[0.5,-0.25],[1.5,-0.75]]The prediction errors are 0.5 and 1.5. Dividing by the batch size gives output gradients 0.25 and 0.75, from which every returned gradient follows.
Example 2
features = [[1,-1]]weights = [0,0]bias = 0targets = [0]return = [[0],[0],[0,0],[0],[0,0]]A zero prediction equal to the target produces zero loss and zero gradients.
Example 3
features = [[2]]weights = [3]bias = -1targets = [1]return = [[5],[8],[8],[4],[12]]The error is 4, so the half-squared loss is 8, the weight gradient is 4 * 2 = 8, and the feature gradient is 4 * 3 = 12.
Constraints
1 <= features.length <= 64.1 <= features[i].length <= 64, and the matrix is rectangular.weights.length = features[i].lengthandtargets.length = features.length.- All input values are finite and lie in
[-10, 10].