Group-Relative Policy Advantages
Problem statement
Compute group-relative advantages for sampled policy responses. Response i belongs to groupIds[i] and receives the integer reward correctnessRewards[i] + formatRewards[i].
Within each group, compute the population mean and population standard deviation of the combined rewards. The advantage of a response is (reward - mean) / standardDeviation. When a group's standard deviation is zero, every response in that group has advantage zero.
Examples
Example 1
groupIds = [0,0,1,1,1]correctnessRewards = [1,0,1,1,0]formatRewards = [1,0,0,1,0]return = ["1.000000","-1.000000","0.000000","1.224745","-1.224745"]Each prompt group is normalized independently after correctness and format rewards are combined.
Unlock this recently reported problem
FastPrep Pro gives you full access to interview problems reported within the last week.
- Full problem statement and constraints
- 1 more worked example, explained
- Guided hints and editorial
- Run your code on real test cases
$99 billed yearly — or $19 month-to-month. Cancel anytime.