[{"data":1,"prerenderedAt":47},["ShallowReactive",2],{"chapter:kernels\u002Ffast\u002Flayout.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":40,"prev":41,"next":44},"kernels","\u002Fkernels\u002Ffast\u002Flayout","Memory Coalescing and Tensor Layout","Making It Fast","fast\u002Flayout.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Ffast\u002Flayout.md","\u003Cp>Most kernels are memory-bound. For a memory-bound kernel, the single largest\nfactor is not how much you load — it is the \u003Cem>order\u003C\u002Fem> you load it in.\u003C\u002Fp>\n\u003Cp>This chapter is about that, and it is the one most likely to make a real\ndifference to a kernel you have already written.\u003C\u002Fp>\n\u003Ch2 id=\"coalescing\">Coalescing\u003C\u002Fh2>\n\u003Cp>When the lanes of a program read memory, the hardware tries to service them with\nas few transactions as possible. Memory comes in fixed-size chunks; if the\naddresses your lanes want all fall in one chunk, that is one transaction. If\nthey are scattered, it is one transaction per chunk touched.\u003C\u002Fp>\n\u003Cp>Contiguous access:\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"text\">\u003Ccode>lanes:      0    1    2    3    4    5    6    7\naddresses:  0    1    2    3    4    5    6    7     → 1 transaction\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Strided access, stride 8:\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"text\">\u003Ccode>lanes:      0    1    2    3    4    5    6    7\naddresses:  0    8   16   24   32   40   48   56     → 8 transactions\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Same number of values. Eight times the memory traffic, because each transaction\nfetches a whole chunk and you use one value from it.\u003C\u002Fp>\n\u003Cp>This is why \u003Ccode>T::arange(0, BLOCK_SIZE) + block_start\u003C\u002Fcode> appears in every simple\nkernel. Consecutive lanes get consecutive addresses, which is the pattern the\nhardware is built for.\u003C\u002Fp>\n\u003Ch2 id=\"what-it-actually-costs\">What it actually costs\u003C\u002Fh2>\n\u003Cp>\u003Ccode>examples\u002Fcoalescing.rs\u003C\u002Fcode> measures it with two kernels that read \u003Cstrong>exactly the\nsame 64 MB\u003C\u002Fstrong> and do exactly the same arithmetic. One walks along rows, the other\ndown columns of the same row-major matrix. The only line that differs:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">arange\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> +\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pid \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">   \u002F\u002F rows:    stride 1\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> rows \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> COLS\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> +\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> col\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">                   \u002F\u002F columns: stride COLS\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> run\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --release\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-triton\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --features\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> cuda\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --example\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> coalescing\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>On an \u003Cstrong>RTX 5070 (sm_120), CUDA 13.3, driver 610.43.02\u003C\u002Fstrong>, over a 4096×4096\nmatrix, 65536 programs of 256 elements each, identical for both:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>access\u003C\u002Fth>\n\u003Cth style=\"text-align:right\">time\u003C\u002Fth>\n\u003Cth style=\"text-align:right\">bandwidth\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>row-major (stride 1)\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">111.6 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">601.1 GB\u002Fs\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>column (stride 4096)\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">267.1 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">251.3 GB\u002Fs\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>\u003Cstrong>2.4× slower for the same bytes and the same arithmetic.\u003C\u002Fstrong> No extra work, no\ndifferent algorithm — the order alone.\u003C\u002Fp>\n\u003Cp>Two things worth drawing out.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>601 GB\u002Fs is the ceiling.\u003C\u002Fstrong> Chapter 16’s block-size sweep plateaued at\n596 GB\u002Fs on the same card with a completely different kernel. Two unrelated\nmemory-bound kernels landing within 1% of each other is what a hardware limit\nlooks like — and it is the number to compare any new kernel against.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2.4× is smaller than the naive prediction, and that is instructive.\u003C\u002Fstrong> With\n128-byte lines and 4-byte elements, 32 elements share a line. A strided read\nwhere every lane lands in a different line should fetch 32× the data, so you\nmight expect something near a 32× slowdown. You get 2.4×.\u003C\u002Fp>\n\u003Cp>The reason is reuse. Neighbouring programs read neighbouring columns, so a line\nfetched for column \u003Ccode>c\u003C\u002Fcode> still holds columns \u003Ccode>c+1 … c+31\u003C\u002Fcode>, and by the time those\nprograms run it is often still in L2. The cache recovers most of what the access\npattern threw away.\u003C\u002Fp>\n\u003Cp>That is worth internalising in both directions: a strided pattern is genuinely\nexpensive, and a back-of-the-envelope traffic calculation will usually\n\u003Cem>overstate\u003C\u002Fem> it, because it ignores the cache. Which is the argument for\nmeasuring rather than predicting.\u003C\u002Fp>\n\u003Ch2 id=\"seeing-it-in-a-real-kernel\">Seeing it in a real kernel\u003C\u002Fh2>\n\u003Cp>The naive matmul from Chapter 11 has both patterns, side by side:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\">\u002F\u002F Row m of A — contiguous.\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> a_offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> k_offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">+\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> m \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> K\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\">\u002F\u002F Column n of B — stride N.\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> b_col_offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> k_offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> N\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> +\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> n\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cem>From \u003Ccode>kernels\u002Fteeny-kernels\u002Fsrc\u002Fmath\u002Fgemm.rs\u003C\u002Fcode>.\u003C\u002Fem>\u003C\u002Fp>\n\u003Cp>\u003Ccode>A\u003C\u002Fcode> is row-major, so a row is contiguous and reads well. A \u003Cem>column\u003C\u002Fem> of a\nrow-major matrix has its elements \u003Ccode>N\u003C\u002Fcode> apart, so the same loop reads it badly.\nThe multiply by \u003Ccode>N\u003C\u002Fcode> in the offset expression is the tell — any offset expression\nwhose lane-varying term is multiplied by something is strided.\u003C\u002Fp>\n\u003Cp>Two ways out, and tiling is the reason the tiled version of this kernel exists:\nload a tile once into fast memory, then read it in whatever order you like.\u003C\u002Fp>\n\u003Ch2 id=\"strides-and-reading-an-index-expression\">Strides, and reading an index expression\u003C\u002Fh2>\n\u003Cp>For a row-major array, the stride of a dimension is the product of all the\ndimensions after it. For \u003Ccode>[B, C, H, W]\u003C\u002Fcode>:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Dimension\u003C\u002Fth>\n\u003Cth>Stride\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>W\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>1\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>H\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>W\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>C\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>H * W\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>B\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>C * H * W\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>So the flat index is \u003Ccode>b*(C*H*W) + c*(H*W) + h*W + w\u003C\u002Fcode>, which is exactly what the\nconv kernels compute:\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"text\">\u003Ccode>x[b, c_in, oh, ow] = x_flat[b*(C_IN*M) + c_in*M + oh*OW + ow]\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The trick for reading these quickly: \u003Cstrong>find which term varies with the lane\nindex\u003C\u002Fstrong>. If it is the one with stride 1, the access is contiguous. Anything else\nis strided, and the multiplier tells you how badly.\u003C\u002Fp>\n\u003Ch2 id=\"nchw-versus-nhwc\">NCHW versus NHWC\u003C\u002Fh2>\n\u003Cp>This is where layout choice becomes a design decision rather than an\nobservation.\u003C\u002Fp>\n\u003Cp>A batch of images has four dimensions: batch \u003Ccode>N\u003C\u002Fcode>, channels \u003Ccode>C\u003C\u002Fcode>, height \u003Ccode>H\u003C\u002Fcode>,\nwidth \u003Ccode>W\u003C\u002Fcode>. Two orderings are in common use.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>NCHW\u003C\u002Fstrong> stores all of one channel’s pixels together. \u003Ccode>[N][C][H][W]\u003C\u002Fcode>, with \u003Ccode>W\u003C\u002Fcode>\ncontiguous. It is what PyTorch uses by default, and what this tree’s conv\nkernels assume.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>NHWC\u003C\u002Fstrong> stores all of one pixel’s channels together. \u003Ccode>[N][H][W][C]\u003C\u002Fcode>, with \u003Ccode>C\u003C\u002Fcode>\ncontiguous.\u003C\u002Fp>\n\u003Cp>Neither is better in general. Which one wins depends on what varies across your\nlanes.\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Operation\u003C\u002Fth>\n\u003Cth>Wants\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>Convolution over spatial extent\u003C\u002Ftd>\n\u003Ctd>NCHW — neighbouring pixels of a channel are adjacent\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Per-channel scale and bias\u003C\u002Ftd>\n\u003Ctd>NHWC — a pixel’s channels are adjacent\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Matmul-style contraction over channels\u003C\u002Ftd>\n\u003Ctd>NHWC — the reduction dimension is contiguous\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Tensor Core paths\u003C\u002Ftd>\n\u003Ctd>Usually NHWC\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>This is why a fused conv+batchnorm is more interesting than it looks. The\nconvolution wants NCHW; the batch norm, which applies one scale per channel,\nwants NHWC. Fusing them means one of the two runs against the grain — but that\nis still far better than writing the intermediate to memory in one layout and\nreading it back in the other.\u003C\u002Fp>\n\u003Cp>The tree makes this explicit in \u003Ccode>channel_bias_add\u003C\u002Fcode>, whose doc comment says it\ntreats a \u003Ccode>(B, C, H, W)\u003C\u002Fcode> tensor as \u003Ccode>NC\u003C\u002Fcode> with \u003Ccode>N = B*H*W\u003C\u002Fcode> — a reinterpretation\nthat makes the channel dimension the fast one for that operation, without moving\nany data.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Converting between layouts costs a full pass over memory.\u003C\u002Fstrong> It is worth it only\nif the destination layout saves more than one pass. Usually the answer is to\npick one layout for the whole model and live with it.\u003C\u002Fp>\n\u003Ch2 id=\"block-pointers-and-descriptors\">Block pointers and descriptors\u003C\u002Fh2>\n\u003Cp>For tiled access there are two addressing modes that describe the tile rather\nthan computing every address.\u003C\u002Fp>\n\u003Cp>\u003Ccode>make_block_ptr\u003C\u002Fcode> takes shape, strides, offsets, a block shape, and an order:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> ptr \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">make_block_ptr\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">base\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">strides\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">offsets\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">block_shape\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">order\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> tile \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">load\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">ptr\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;[\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> Some\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">PaddingOption\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Zero\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> false\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The \u003Ccode>boundary_check\u003C\u002Fcode> argument — the one that is \u003Ccode>&amp;[]\u003C\u002Fcode> in every kernel in Part 2\n— becomes useful here: it names the dimensions to check, and out-of-range lanes\nget \u003Ccode>padding_option\u003C\u002Fcode> instead of a mask you built by hand.\u003C\u002Fp>\n\u003Cp>\u003Ccode>make_tensor_descriptor\u003C\u002Fcode> is the TMA form from Chapter 11, and on hardware that\nhas it the copy is done by dedicated hardware rather than by the arithmetic\nunits.\u003C\u002Fp>\n\u003Cp>Both let the compiler see the whole access pattern at once, which is more than\nit can infer from arbitrary offset arithmetic.\u003C\u002Fp>\n\u003Ch2 id=\"the-alignment-constraint\">The alignment constraint\u003C\u002Fh2>\n\u003Cp>TMA requires rows aligned to 16 bytes — four elements for \u003Ccode>f32\u003C\u002Fcode>. A tensor whose\nlast dimension is not a multiple of four therefore cannot be addressed directly.\u003C\u002Fp>\n\u003Cp>The SDK handles this by padding: \u003Ccode>RuntimeOp::forward_output_row_stride\u003C\u002Fcode> returns\nthe stride the buffer actually has, rounded up, and the executor allocates\naccordingly and passes the real stride to \u003Ccode>pack_args\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>Which means: \u003Cstrong>use the \u003Ccode>output_row_stride\u003C\u002Fcode> argument, not\n\u003Ccode>output_shape.last()\u003C\u002Fcode>\u003C\u002Fstrong>. They are usually equal and occasionally not, and when\nthey differ, computing your own gives a kernel that reads the wrong addresses.\nThe same applies to \u003Ccode>backward_grad_output_row_stride\u003C\u002Fcode> on the backward path.\u003C\u002Fp>\n\u003Ch2 id=\"compiler-hints\">Compiler hints\u003C\u002Fh2>\n\u003Cp>Three methods promise the compiler something it cannot prove:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Method\u003C\u002Fth>\n\u003Cth>Promise\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>T::multiple_of(x, values)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>These values are multiples of these constants\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>T::max_contiguous(x, values)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>This many elements are contiguous along each dimension\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>T::max_constancy(x, values)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>This many elements are constant along each dimension\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>They generate no code. They let the compiler emit wider loads it would otherwise\nhave to guard.\u003C\u002Fp>\n\u003Cp>They are also unchecked promises. If you tell it offsets are multiples of 16 and\nthey are not, you get wrong results with no diagnostic. Use them when you know\nsomething structural — a padded stride, an aligned base — and not otherwise.\u003C\u002Fp>\n\u003Ch2 id=\"a-checklist\">A checklist\u003C\u002Fh2>\n\u003Col>\n\u003Cli>\u003Cstrong>Find the lane-varying term\u003C\u002Fstrong> in each offset expression.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Check its multiplier.\u003C\u002Fstrong> Stride 1 is contiguous; anything else is not.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>If it is strided, can you tile?\u003C\u002Fstrong> Load once, reuse from fast memory.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>If not, can you change the layout?\u003C\u002Fstrong> Sometimes the fix is upstream.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Use the passed row stride\u003C\u002Fstrong>, never the shape’s last dimension.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Measure.\u003C\u002Fstrong> Chapter 18. A traffic calculation tells you the worst case; the\ncache decides how much of it you actually pay. On this card the gap between\nthe two was 32× predicted against 2.4× measured.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>Next: how to time it.\u003C\u002Fp>\n",[12,16,19,22,25,28,31,34,37],{"id":13,"text":14,"level":15},"coalescing","Coalescing",2,{"id":17,"text":18,"level":15},"what-it-actually-costs","What it actually costs",{"id":20,"text":21,"level":15},"seeing-it-in-a-real-kernel","Seeing it in a real kernel",{"id":23,"text":24,"level":15},"strides-and-reading-an-index-expression","Strides, and reading an index expression",{"id":26,"text":27,"level":15},"nchw-versus-nhwc","NCHW versus NHWC",{"id":29,"text":30,"level":15},"block-pointers-and-descriptors","Block pointers and descriptors",{"id":32,"text":33,"level":15},"the-alignment-constraint","The alignment constraint",{"id":35,"text":36,"level":15},"compiler-hints","Compiler hints",{"id":38,"text":39,"level":15},"a-checklist","A checklist",false,{"title":42,"titleHtml":42,"route":43},"Choosing a Block Size","\u002Fkernels\u002Ffast\u002Fblock-size",{"title":45,"titleHtml":45,"route":46},"Measuring","\u002Fkernels\u002Ffast\u002Fmeasuring",1786271830052]