Fact — formula — Knowledge Tree

When the number of in-context examples D increases, the prediction loss for both single-head and multi-head attention in transformers is in O(1/D), but the prediction loss for multi-head attention has a smaller multiplicative constant.

Authors

Person: Samuel Tesfazgi, Leonhard Sprandl, Sandra Hirche Organization: AISTATS
Track: Poster Session 3 - aistats 2026

Sources

Track: Poster Session 3 - aistats 2026 virtual.aistats.org Samuel Tesfazgi, Leonhard Sprandl, Sandra Hirche · AISTATS via serper

Referenced by nodes (2)

Transformers concept
In-Context Learning concept