feedd.›AI
AIMarkTechPost · 79d ago

A Coding Guide to NVIDIA’s Tile-Based GPU Programming: From cuTile and Triton Kernels to Flash Attention

NVIDIA's tile-based GPU programming uses TileGym to operate on whole data tiles rather than single threads, with implementations falling back from cuTile to Triton kernels depending on available hardware. The tutorial covers vector addition, fused GELU, row-wise softmax, tiled matrix multiplication, and flash attention operations validated against PyTorch.

Read full story →
More from AI

Hyundai Motor Group activated its Data Flywheel system and demonstrated Level 2++ autonomous driving technology at an event in South Korea on September 13, 2026. The Group presented a dual-track autonomous driving roadmap alongside development strategies and implementation plans at 42dot headquarters in Gyeonggi Province.

01

A fruit fly connectome with 166,700 neurons and 25.6 million connections was integrated into a 1.2 billion parameter frozen language model, training only 278,528 additional parameters. The connectome integration achieved 0.0222 nat per token improvement, but control models without the biological wiring performed slightly better across all random seeds.

02

OpenAI committed to embedding independent evaluators with employee-level access to its systems, matching Anthropic's earlier pledge on the same day. Both companies endorsed slowing frontier AI development pace, with OpenAI promising to share additional details soon.

03

Get feedd. daily

Top stories in your inbox every morning. Pick what you want.

No spam. Unsubscribe anytime.