推文
@uqbariss · 2026-10-11 00:15
What if the biggest AI bottleneck isn't computation, but moving data? FlashAttention accelerated exact attention by reducing memory transfers between GPU HBM and on-chip SRAM. No bigger model. No magical new GPU. Just a smarter way to move bytes. The next AI performance breakthrough might be hiding in memory, not mathematics. Paper: https://arxiv.org/abs/2205.14135
曝光 40 · 评论 0 · 点赞 0 · 书签 0 · 曝光/时 12.277839050560926