update publications (#1308)

2024-01-17 11:06:46 -08:00 · 2024-01-17 11:06:46 -08:00 · 139b93db61
commit 139b93db61
parent ca37d632c9
1 changed files with 2 additions and 0 deletions
--- a/PUBLICATIONS.md
+++ b/PUBLICATIONS.md
@ -2,6 +2,8 @@

 ## 2023

+- ["A Case Study in CUDA Kernel Fusion: Implementing FlashAttention-2 on NVIDIA Hopper Architecture using the CUTLASS Library"](https://arxiv.org/abs/2312.11918). Ganesh Bikshandi and Jay Shah. _arXiv_, December 2023.
+
 - ["FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning"](https://arxiv.org/abs/2307.08691). Tri Dao. _Technical Report_, July 2023.

 - ["MegaBlocks: Efficient Sparse Training with Mixture-of-Experts"](https://arxiv.org/abs/2211.15841). Trevor Gale, Deepak Narayanan, Cliff Young, Matei Zaharia. _Proceedings of the Sixth Machine Learning and Systems_, May 2023.