residual stream is dim order ~100 or 10K in big model different layers can store information in different subspaces attention heads are independent and additive