residual stream is dim order ~100 or 10K in big model

  • different layers can store information in different subspaces

attention heads are independent and additive