[attention] Investigate overlapping matmul and softmax #91

antiagainst · 2024-08-05T03:37:29Z

In Flash Attention 3 we see a technique to overlap matmul and softmax from different waves to maximize mfma utilization. We should consider how to use it for current attention. Need to understand hardware scheduler and see how to work with/around it, like using s_setprio instructions.

The text was updated successfully, but these errors were encountered:

antiagainst added this to Turbine: SDXL on CDNA Aug 5, 2024

antiagainst converted this from a draft issue Aug 5, 2024

antiagainst added this to the SDXL-MLPerf milestone Aug 5, 2024

antiagainst moved this to Todo in Turbine: SDXL on CDNA Aug 5, 2024

antiagainst assigned Groverkss and raikonenfnu Aug 5, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[attention] Investigate overlapping matmul and softmax #91

[attention] Investigate overlapping matmul and softmax #91

antiagainst commented Aug 5, 2024 •

edited

Loading

[attention] Investigate overlapping matmul and softmax #91

[attention] Investigate overlapping matmul and softmax #91

Comments

antiagainst commented Aug 5, 2024 • edited Loading

antiagainst commented Aug 5, 2024 •

edited

Loading