FastPrepGroup-Relative Policy Advantages

Group-Relative Policy Advantages

Amazon logoAmazon● MediumFULLTIMEONSITE INTERVIEW

Problem statement

Compute group-relative advantages for sampled policy responses. Response i belongs to groupIds[i] and receives the integer reward correctnessRewards[i] + formatRewards[i].

Within each group, compute the population mean and population standard deviation of the combined rewards. The advantage of a response is (reward - mean) / standardDeviation. When a group's standard deviation is zero, every response in that group has advantage zero.

The problem statement continues
Pro

Examples

Example 1

groupIds = [0,0,1,1,1]correctnessRewards = [1,0,1,1,0]formatRewards = [1,0,0,1,0]return = ["1.000000","-1.000000","0.000000","1.224745","-1.224745"]

Each prompt group is normalized independently after correctness and format rewards are combined.

FastPrep Pro
Reported in 1 Amazon interview this week

Unlock this recently reported problem

FastPrep Pro gives you full access to interview problems reported within the last week.

  • Full problem statement and constraints
  • 1 more worked example, explained
  • Guided hints and editorial
  • Run your code on real test cases
$8.25/month

$99 billed yearly — or $19 month-to-month. Cancel anytime.

Free plan — 2 of 2 free unlocks used this week
See Amazon hiring insights
CodePython 3
Run and Submit unlock with Pro
FastPrep Pro
Reported in 1 Amazon interview this week

Unlock this recently reported problem

FastPrep Pro gives you full access to interview problems reported within the last week.

  • Full problem statement and constraints
  • 1 more worked example, explained
  • Guided hints and editorial
  • Run your code on real test cases
$8.25/month

$99 billed yearly — or $19 month-to-month. Cancel anytime.

Free plan — 2 of 2 free unlocks used this week