Skip to Main Content

PyTorch Helion Hackathon

Mar 14, 2026 · San Francisco, CA

We implemented 4 GPU kernels using Helion DSL targeting NVIDIA B200 (Blackwell) GPUs: 1. causal_conv1d: Depthwise causal 1D convolution (core Mamba/Mamba-2…

Axiom

Here's how each kernel works: 1. FP8 Quantization (fp8_quant) Memory-bound, fully parallel. Takes float32 activations, groups them (64-128…

WarpDrive

using helion with NVIDIA ComputeIQ and TileIR.

Shiprail

fully autokernel generated

cob-farmers

causal_conv1d: Causal depthwise 1D convolution with static per-shape configs and a stable fused convolution path; this is the highest-confidence officially…

Team SEWR

GPU modders

GPU modders

causal_conv1d - brief summary reading, it's functionality not known and performance not known fp8_quant_py - brief summary reading, it's functionality not…

vectorblock.io

--- Causal Depthwise Conv1D (Mamba/SSM) Core operator in Mamba state-space models. Each channel is convolved independently with a causal filter…

Voyager

I added two PRs for improvements: https://github.com/pytorch/helion/pull/1705 https://github.com/pytorch/helion/pull/1704

Hell

gated_deltanet_recompute_w_u/submission/submission.py in the repo computes the Gated DeltaNet WY-transform forward outputs (w, u) from raw inputs (k, v, beta…

Clio AI

Efficient kernels

Campanile Compilers

We removed all the extra compute in all the kernels. We adjusted tile size and autotuned the kernels and used hl.dot instead of separate outer products for…

Team Manifold - includes NVIDIA employee

Leaderboard name: bloomberg9383 All 5 kernels implemented in Helion DSL for Gated DeltaNet on B200: 1. fp8_quant: FP8 group quantization — computes…

KernIt_Raj

1. Causal Conv1d — Depthwise Causal Convolution Implements the core Mamba/Mamba-2 causal convolution where each channel is convolved independently with past…

ZzCompute

KernalForge is our Helion hackathon project for optimizing production-style GPU kernels on NVIDIA B200 hardware. We built and benchmarked all five required…

KernelForge

implemented all 4 scored Helion kernels for Gated DeltaNet and causal conv1d, optimized for B200 GPUs with per-shape autotuned configs. causal_conv1d…

usagi

Problem 3 (gated_deltanet_chunk_fwd_h): Implemented inter-chunk state recurrence kernel for Gated DeltaNet using Helion DSL with hardcoded configs optimized…

RackSavant AI
Team Beaker (CodingMaster)

I autotuned the kernels and tried to do some DMA-Compute overlapping, though that wasn't really effective.

Noob
RAP

Optimized code for each kernel is listed under respective folder

Buba Shrimp

https://github.com/narain1/helion-hack/blob/main/submission_fp8_quant.py

heisenberg

Not sure where the f8_quant submission is supposed to go, so here it is: https://github.com/dpiresearch/Helion_20260314/blob/main/helion/fp8_quant_py/submissio…

DPI Research

Autotuned kernels

LilyErnest

4 kernels submitted with optimizations

Pradeep

We implemented four kernels for the Helion challenge on NVIDIA B200 GPUs, prioritizing FP32 accumulation for precision and 3D tiling for better hardware…

Battlestars

Created a LLM-based-search for autotuner. It takes in context of the kernel in scoring and performs a beam-search with the candidate kernel…

Cappucino