← Back to Bounties
G
Groq
groq.com ↗Optimize CUDA Kernel for Transformer Attention
AdvancedML Systems⏱ 12–24 hours● OPEN0 applicants
The Challenge
Implement a fused attention kernel in CUDA that outperforms naive attention by at least 2x on an A100 GPU. Bonus for flash-attention style tiling. Benchmarks required.
Required Skills
cudapythonc++pytorch
What to Submit
GITHUB REPO
Reward
$1,000
USD
Details
DeadlineFeb 28, 2025
Spots0 / 25 filled
Time est.12–24 hours
Sign in to Apply →
Free account required. GitHub login takes 10 seconds.