PyTorch Helion Hackathon
Mar 14, 2026 · San Francisco, CA
We implemented 4 GPU kernels using Helion DSL targeting NVIDIA B200 (Blackwell) GPUs: 1. causal_conv1d: Depthwise causal 1D convolution (core Mamba/Mamba-2…
Here's how each kernel works: 1. FP8 Quantization (fp8_quant) Memory-bound, fully parallel. Takes float32 activations, groups them (64-128…
using helion with NVIDIA ComputeIQ and TileIR.
fully autokernel generated
causal_conv1d: Causal depthwise 1D convolution with static per-shape configs and a stable fused convolution path; this is the highest-confidence officially…
GPU modders
causal_conv1d - brief summary reading, it's functionality not known and performance not known fp8_quant_py - brief summary reading, it's functionality not…
--- Causal Depthwise Conv1D (Mamba/SSM) Core operator in Mamba state-space models. Each channel is convolved independently with a causal filter…
I added two PRs for improvements: https://github.com/pytorch/helion/pull/1705 https://github.com/pytorch/helion/pull/1704
gated_deltanet_recompute_w_u/submission/submission.py in the repo computes the Gated DeltaNet WY-transform forward outputs (w, u) from raw inputs (k, v, beta…
Efficient kernels
We removed all the extra compute in all the kernels. We adjusted tile size and autotuned the kernels and used hl.dot instead of separate outer products for…
Leaderboard name: bloomberg9383 All 5 kernels implemented in Helion DSL for Gated DeltaNet on B200: 1. fp8_quant: FP8 group quantization — computes…
1. Causal Conv1d — Depthwise Causal Convolution Implements the core Mamba/Mamba-2 causal convolution where each channel is convolved independently with past…
KernalForge is our Helion hackathon project for optimizing production-style GPU kernels on NVIDIA B200 hardware. We built and benchmarked all five required…
implemented all 4 scored Helion kernels for Gated DeltaNet and causal conv1d, optimized for B200 GPUs with per-shape autotuned configs. causal_conv1d…
Problem 3 (gated_deltanet_chunk_fwd_h): Implemented inter-chunk state recurrence kernel for Gated DeltaNet using Helion DSL with hardcoded configs optimized…
I autotuned the kernels and tried to do some DMA-Compute overlapping, though that wasn't really effective.
f
Optimized code for each kernel is listed under respective folder
https://github.com/narain1/helion-hack/blob/main/submission_fp8_quant.py
Not sure where the f8_quant submission is supposed to go, so here it is: https://github.com/dpiresearch/Helion_20260314/blob/main/helion/fp8_quant_py/submissio…
Autotuned kernels
4 kernels submitted with optimizations
We implemented four kernels for the Helion challenge on NVIDIA B200 GPUs, prioritizing FP32 accumulation for precision and 3D tiling for better hardware…
Created a LLM-based-search for autotuner. It takes in context of the kernel in scoring and performs a beam-search with the candidate kernel…