A Shared rotation matrix learned for a single concept which maps hidden states into a feature space where the causal variable is linearly separable.
Position specific Gates for each token position .
A Mask that identifies which dimensions encode the causal feature. At each position the intervention mask is:
Loss for Gates:
Should make each dimension of the learned concept subspace be “claimed” by only one token position at any given time so the model cannot active all positions at once.