Talks

invited talks and presentations in reversed chronological order.

Modern Position Encoding in Transformers: RoPE/Yarn and PaTH

Remote

Introduction of RoPE/Yarn and PaTH for modern position encoding in transformers.

Toward More Expressive yet Scalable RNNs: DeltaNet and Its Variants

Remote

Introduction of DeltaNet and its variants for more expressive yet scalable RNNs.

Linear attention and beyond or What's Next for Mamba? Towards More Expressive Recurrent Update Rules

Remote

Discussion on developments in linear attention and its variants.

Linear Transformers for Efficient Sequence Modeling

HazyResearch @ Standford

Introduction of linear transformers for efficient sequence modeling.

Gated linear Recurrence for Efficient Sequence Modeling

Cornell Tech

Introduction of gated linear recurrence for efficient sequence modeling.