Attention & TransformersMedium
Softmax Attention Weights
~12 mincode completion
Implement softmax_attention(scores) that converts a 1D scores vector to attention weights.
Examples
Equal scores: uniform weights
- Input
- softmax_attention([0, 0, 0, 0])
- Output
- [0.25, 0.25, 0.25, 0.25]
Dominant score takes most weight
- Input
- softmax_attention([0, 10, 0])
- Output
- [0.0000454, 0.9999092, 0.0000454]
Hints
Hint 1
Subtract the row max before exponentiating to keep the result stable.
Hint 2
Do not forget to subtract the max. That step is easy to skip.
Requirements
scores: 1D array of raw scoresReturn 1D array of attention weights summing to 1.
Constraints
Allowed library: NumPy only
Time limit: 200 ms, Memory: 64 MB
Where this shows up
~12 min
••••••••••••••••
8 employers weight this skill
4 frontier labs, 3 AI product companies, 1 enterprise vendor. Top match scores 92.
Python
import numpy as np
def softmax_attention(scores: np.ndarray) -> np.ndarray:
"""
Convert raw attention scores to a probability distribution.
Args:
scores: 1D array of raw scores
Returns:
1D array of attention weights summing to 1.
"""
# YOUR CODE HERE
pass