Incremental Hard Attention with a KV Cache
Problem statement
Implement a deterministic incremental attention adapter. The input contains n token embeddings and n position embeddings, each of dimension d, plus square key and value weight matrices.
At step t:
The problem statement continues
ProExamples
Example 1
tokenEmbeddings = [[1,0],[1,1]]positionEmbeddings = [[0,0],[0,0]]keyWeights = [[1,0],[0,1]]valueWeights = [[0,1],[1,0]]return = [[0,1],[1,1]]At each step, the newest key has the unique highest score. Values come from the separate value projection, not from the attention output.
FastPrep Pro
Reported in 1 OpenAI interview this weekUnlock this recently reported problem
FastPrep Pro gives you full access to interview problems reported within the last week.
- Full problem statement and constraints
- 1 more worked example, explained
- Guided hints and editorial
- Run your code on real test cases
$8.25/month
$99 billed yearly — or $19 month-to-month. Cancel anytime.
Free plan — 2 of 2 free unlocks used this week