CORTEXA
← Browse

Gaetan Hains

1 paper indexed

arxivcs.LGcs.AI2026-07-21

MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

Lenore Mulin, Gaetan Hains

We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step. The artifacts are: (1)~a s…

View free PDFSource page