Padded Vector Batch Pipeline
Problem statement
Process vectors in consecutive batches of at most batchSize. Within each retained batch, pad every vector on the right with padValue to the longest vector length in that batch.
If dropLast is true, discard the final batch when it contains fewer than batchSize vectors. Return all retained padded vectors in their original order. Padding is local to each batch, so rows from different batches may have different lengths.
Function
runDataPipeline(vectors: int[][], batchSize: int, padValue: int, dropLast: boolean) → int[][]Examples
Example 1
vectors = [[1,2],[3],[4,5,6]]batchSize = 2padValue = 0dropLast = falsereturn = [[1,2],[3,0],[4,5,6]]The first two vectors share width two; the one-vector final batch keeps width three.
Example 2
vectors = [[1],[2,3],[4]]batchSize = 2padValue = -1dropLast = truereturn = [[1,-1],[2,3]]The incomplete final batch is dropped.
Example 3
vectors = [[7],[8,9]]batchSize = 5padValue = 4dropLast = falsereturn = [[7,4],[8,9]]An incomplete batch is retained when dropping is disabled.
Constraints
1 <= vectors.length <= 10^5and every vector is nonempty.1 <= batchSize <= 10^5.- The total number of input elements is at most
2 * 10^5.