FastPrepPadded Vector Batch Pipeline

Padded Vector Batch Pipeline

Siemens logoSiemens● MediumFULLTIMEPHONE SCREEN
Learn

Problem statement

Process vectors in consecutive batches of at most batchSize. Within each retained batch, pad every vector on the right with padValue to the longest vector length in that batch.

If dropLast is true, discard the final batch when it contains fewer than batchSize vectors. Return all retained padded vectors in their original order. Padding is local to each batch, so rows from different batches may have different lengths.

Function

runDataPipeline(vectors: int[][], batchSize: int, padValue: int, dropLast: boolean) → int[][]

Examples

Example 1

vectors = [[1,2],[3],[4,5,6]]batchSize = 2padValue = 0dropLast = falsereturn = [[1,2],[3,0],[4,5,6]]

The first two vectors share width two; the one-vector final batch keeps width three.

Example 2

vectors = [[1],[2,3],[4]]batchSize = 2padValue = -1dropLast = truereturn = [[1,-1],[2,3]]

The incomplete final batch is dropped.

Example 3

vectors = [[7],[8,9]]batchSize = 5padValue = 4dropLast = falsereturn = [[7,4],[8,9]]

An incomplete batch is retained when dropping is disabled.

Constraints

  • 1 <= vectors.length <= 10^5 and every vector is nonempty.
  • 1 <= batchSize <= 10^5.
  • The total number of input elements is at most 2 * 10^5.
See Siemens hiring insights
public int[][] runDataPipeline(int[][] vectors, int batchSize, int padValue, boolean dropLast) {
    // Write your solution here.
}
vectors[[1,2],[3],[4,5,6]]
batchSize2
padValue0
dropLastfalse
expected[[1,2],[3,0],[4,5,6]]
Checking account…