[{"data":1,"prerenderedAt":45},["ShallowReactive",2],{"chapter:kernels\u002Ffast\u002Fblock-size.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":38,"prev":39,"next":42},"kernels","\u002Fkernels\u002Ffast\u002Fblock-size","Choosing a Block Size","Making It Fast","fast\u002Fblock-size.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Ffast\u002Fblock-size.md","\u003Cp>Chapter 6 gave you three rules of thumb and told you to measure. This chapter is\nwhat you are actually choosing between, and what the SDK does and does not let\nyou control.\u003C\u002Fp>\n\u003Ch2 id=\"two-different-numbers\">Two different numbers\u003C\u002Fh2>\n\u003Cp>The word “block” does two jobs, and conflating them causes real confusion.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The const generic\u003C\u002Fstrong> — \u003Ccode>BLOCK_SIZE\u003C\u002Fcode>, \u003Ccode>BLOCK_OW\u003C\u002Fcode>, \u003Ccode>BLOCK_M\u003C\u002Fcode> — is how much\n\u003Cem>data\u003C\u002Fem> one program covers. It is baked into the kernel at compile time, and it\ndetermines the shape of every tensor inside the body.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The launch block\u003C\u002Fstrong> — \u003Ccode>CudaLaunchConfig::block\u003C\u002Fcode> — is how many \u003Cem>threads\u003C\u002Fem> the\nhardware gives that program.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>They are never the same thing, and you only choose one of them.\u003C\u002Fstrong> \u003Ccode>teenyc\u003C\u002Fcode>\ndecides the thread count and records it in the PTX as a \u003Ccode>.reqntid\u003C\u002Fcode> directive.\nThe driver enforces it: launch with any other block dimension and you get\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"text\">\u003Ccode>Error: CUDA error: 1 (invalid argument)\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>That is not a warning or a slowdown. It is a hard failure, and it is what you\nhit the first time you assume \u003Ccode>BLOCK_SIZE\u003C\u002Fcode> is a thread count.\u003C\u002Fp>\n\u003Cp>The sweep below makes it plain — six different \u003Ccode>BLOCK_SIZE\u003C\u002Fcode> values, and the\ncompiler picks 128 threads for every one of them.\u003C\u002Fp>\n\u003Cp>For a tiled kernel the two are visibly different. \u003Ccode>conv2d_bn_silu\u003C\u002Fcode> has\n\u003Ccode>BLOCK_OW = 16\u003C\u002Fcode> — sixteen output columns per tile — and its bench launches 128\nthreads:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> cfg \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> CudaLaunchConfig\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> {\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">    grid\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">grid\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">    block\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">128\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">    cluster\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">};\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Those 128 threads cooperate on a 16-wide tile. The compiler decides how. Nothing\nrequires the two numbers to match, and nothing checks that your choice is\nsensible.\u003C\u002Fp>\n\u003Ch2 id=\"what-you-control\">What you control\u003C\u002Fh2>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">pub\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> struct\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> CudaLaunchConfig\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> {\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    pub\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> grid\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">u32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 3\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">     \u002F\u002F how many programs\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    pub\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> block\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">u32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 3\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">    \u002F\u002F threads per program\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    pub\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> cluster\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">u32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 3\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003Cspan style=\"color:#7F877D;font-style:italic\">  \u002F\u002F programs per cluster (Hopper and later)\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">}\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Plus the const generics, chosen when you build the kernel. That is the whole\nsurface.\u003C\u002Fp>\n\u003Cp>\u003Ccode>RuntimeOp\u003C\u002Fcode> exposes the same two through \u003Ccode>block()\u003C\u002Fcode> and \u003Ccode>grid()\u003C\u002Fcode>, and the helpers\nin \u003Ccode>teeny_cuda::testing\u003C\u002Fcode> build a config for you:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Helper\u003C\u002Fth>\n\u003Cth>Use\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>launch_config_with_grid(grid_x, &amp;program)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Cstrong>The safe one.\u003C\u002Fstrong> Grid you computed, threads from the PTX\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>launch_config_from_program(n, &amp;program)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Threads from the PTX, grid from the element count — correct only when one program handles exactly one thread’s worth of data\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>launch_config(n_elements, block_size)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Both from you. Fails unless \u003Ccode>block_size\u003C\u002Fcode> happens to equal what the compiler chose\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>Prefer the first. You know how much data one program covers — that is your\n\u003Ccode>BLOCK_SIZE\u003C\u002Fcode> — so compute the grid from it and let the metadata supply the\nthreads:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> grid \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> N\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">div_ceil\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> as\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> usize\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> cfg \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> teeny_cuda\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">testing\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">launch_config_with_grid\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">grid\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">program\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Ccode>launch_config\u003C\u002Fcode> is a trap in waiting. It works whenever your \u003Ccode>BLOCK_SIZE\u003C\u002Fcode>\ncoincides with the compiler’s thread count, which for a simple elementwise\nkernel at 128 it usually does — and then silently stops working when you change\nthe constant.\u003C\u002Fp>\n\u003Cp>The loader parses the thread count out of \u003Ccode>.reqntid\u003C\u002Fcode> like this:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> threads \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> program\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">metadata\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">threads_per_block\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">().\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">max\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#6FBF98\">CudaLaunchConfig\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> {\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">    grid\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">n_elements \u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\">as\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> u32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">).\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">div_ceil\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">threads\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">    block\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">threads\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">    cluster\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">program\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">metadata\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">num_ctas\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">max\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#8A9088\">}\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>So the compiler’s own choice is available, and using it is usually right.\u003C\u002Fp>\n\u003Ch2 id=\"what-you-do-not-control\">What you do not control\u003C\u002Fh2>\n\u003Cp>This is the part that differs sharply from Python Triton.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>\u003Ccode>num_warps\u003C\u002Fcode> is not settable.\u003C\u002Fstrong> In Python it is a launch argument: how many\nwarps cooperate on one program. Here it exists only as a value parsed \u003Cem>out\u003C\u002Fem> of\ncompiled PTX — derived from \u003Ccode>.reqntid\u003C\u002Fcode>, rounded up to whole warps — and the\nstruct holding it is \u003Ccode>pub(crate)\u003C\u002Fcode>. You cannot read it directly, let alone set\nit; the only access is through \u003Ccode>launch_config_from_program\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>\u003Ccode>num_stages\u003C\u002Fcode> does not exist at all.\u003C\u002Fstrong> In Python it controls the depth of the\ncompiler’s software pipeline — how many loop iterations’ loads are in flight at\nonce. There is no equivalent anywhere in this SDK.\u003C\u002Fp>\n\u003Cp>Both are recorded as item 3 in\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fblob\u002Fmain\u002Fbooks\u002Fkernels\u002FKNOWN-GAPS.md\" target=\"_blank\" rel=\"noopener noreferrer\">\u003Ccode>KNOWN-GAPS.md\u003C\u002Fcode>\u003C\u002Fa>.\nIt is why this chapter is narrower than a Triton performance guide would be:\nmuch of what such a guide tells you to tune is not exposed.\u003C\u002Fp>\n\u003Cp>What you \u003Cem>can\u003C\u002Fem> tune is the thread count and the tile shape, which between them\nstill cover most of the available performance.\u003C\u002Fp>\n\u003Ch2 id=\"occupancy\">Occupancy\u003C\u002Fh2>\n\u003Cp>Occupancy is how many programs a card can keep resident at once. It matters\nbecause it is how memory latency gets hidden: when one program stalls waiting\nfor a load, another runs.\u003C\u002Fp>\n\u003Cp>Each program consumes:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Registers\u003C\u002Fstrong>, for the tensors in its body. Bigger blocks and more live\ntensors mean more registers.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Shared memory\u003C\u002Fstrong>, for reductions and tiles.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The card has a fixed budget of each per multiprocessor. Divide, and you get how\nmany programs fit.\u003C\u002Fp>\n\u003Cp>That is the tension. A larger tile does more arithmetic per byte loaded — good\n— but uses more registers, so fewer programs are resident, so there is less\nother work to hide latency with — bad. The optimum is somewhere in the middle\nand it moves with the kernel and the card.\u003C\u002Fp>\n\u003Cp>Some of the inputs are visible in the PTX metadata the loader parses: the shared\nmemory a kernel wants, its global scratch requirements, the thread count. They\nare not exposed as a public API, so reading them means either the PTX itself or\n\u003Ccode>nvdisasm\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Maximum occupancy is not the goal.\u003C\u002Fstrong> A kernel at 50% occupancy with good\narithmetic intensity routinely beats one at 100% that is memory-starved.\nOccupancy is a diagnostic, not a target.\u003C\u002Fp>\n\u003Ch3 id=\"the-other-kind-of-occupancy-problem\">The other kind of occupancy problem\u003C\u002Fh3>\n\u003Cp>Everything above is about \u003Cem>resources\u003C\u002Fem> — registers and shared memory limiting how\nmany programs fit. There is a second, simpler failure that looks identical from\na benchmark and is completely different underneath: \u003Cstrong>not launching enough\nprograms to fill the machine.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>A card has some number of streaming multiprocessors. If your grid is 40 programs\nand the card has 48 SMs, most of the machine is idle regardless of how efficient\nyour kernel is. No amount of register tuning helps.\u003C\u002Fp>\n\u003Cp>This is a real case in this tree. Profiling YOLO26n found conv layers running at\n8% achieved occupancy with 100% \u003Cem>theoretical\u003C\u002Fem> occupancy — the giveaway that\nresources were not the limit. Deep layers have small spatial extent and many\nchannels, and a fixed tile size produced as few as 40 blocks.\u003C\u002Fp>\n\u003Cp>The fix was to pick the tile size from the shape:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> lowering \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> TritonLowering\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">default\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">().\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">with_sm_count\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Some\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">48\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">));\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>With \u003Ccode>sm_count\u003C\u002Fcode> set, the lowering chooses the largest candidate tile whose\nresulting grid still clears a small multiple of the SM count — smaller tiles\nwhere that means more blocks, larger where the shape already provides enough.\nLeft at \u003Ccode>None\u003C\u002Fcode>, the default, tile sizes stay fixed exactly as before.\u003C\u002Fp>\n\u003Cp>Two things to take from this.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Check your grid size before tuning anything else.\u003C\u002Fstrong> It is one division, and it\nrules out the most embarrassing cause.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The SM count is a parameter, not a query.\u003C\u002Fstrong> It sits alongside \u003Ccode>ptx_version\u003C\u002Fcode> in\n\u003Ccode>Options\u003C\u002Fcode> and is deliberately not read from the local device, because\nahead-of-time compilation routinely targets a card that is not the one doing the\ncompiling — Chapter 23.\u003C\u002Fp>\n\u003Ch2 id=\"a-measured-sweep\">A measured sweep\u003C\u002Fh2>\n\u003Cp>\u003Ccode>examples\u002Fblock_size.rs\u003C\u002Fcode> compiles the vector-add kernel once per block size and\ntimes each on 32M elements — three 128 MB buffers, far past any cache, so this\nis memory and nothing else.\u003C\u002Fp>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> run\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --release\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-triton\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --features\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> cuda\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --example\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> block_size\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>On an \u003Cstrong>RTX 5070 (sm_120), CUDA 13.3, driver 610.43.02\u003C\u002Fstrong>:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth style=\"text-align:right\">\u003Ccode>BLOCK_SIZE\u003C\u002Fcode>\u003C\u002Fth>\n\u003Cth style=\"text-align:right\">threads\u003C\u002Fth>\n\u003Cth style=\"text-align:right\">time\u003C\u002Fth>\n\u003Cth style=\"text-align:right\">bandwidth\u003C\u002Fth>\n\u003Cth style=\"text-align:right\">grid\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd style=\"text-align:right\">32\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">128\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">1089.0 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">369.8 GB\u002Fs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">1048576\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd style=\"text-align:right\">64\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">128\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">707.9 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">568.8 GB\u002Fs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">524288\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd style=\"text-align:right\">128\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">128\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">676.6 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">595.1 GB\u002Fs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">262144\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd style=\"text-align:right\">256\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">128\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">675.1 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">\u003Cstrong>596.5 GB\u002Fs\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">131072\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd style=\"text-align:right\">512\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">128\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">677.0 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">594.8 GB\u002Fs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">65536\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd style=\"text-align:right\">1024\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">128\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">681.7 µs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">590.6 GB\u002Fs\u003C\u002Ftd>\n\u003Ctd style=\"text-align:right\">32768\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>Three things to read off it.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The threads column never moves.\u003C\u002Fstrong> Six compilations, 128 threads every time.\nThis is the distinction at the top of the chapter, measured.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>32 is genuinely bad, and for a specific reason.\u003C\u002Fstrong> Each program has 128 threads\nbut only 32 elements to work on, so three quarters of every program’s threads\nhave nothing to do. The bandwidth loss — 370 against 596 GB\u002Fs, a 38% drop — is\nalmost exactly that idle fraction.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Everything from 128 up is the same.\u003C\u002Fstrong> 595, 596, 595, 591 GB\u002Fs: a 1% spread,\nwhich is noise. Once each program has enough work to occupy its threads, this\nkernel is limited by memory and nothing you do to the block size will change\nthat.\u003C\u002Fp>\n\u003Cp>That flat plateau is the useful result. It says \u003Cem>stop tuning\u003C\u002Fem> — the kernel is at\nthe card’s practical limit for this access pattern, and effort is better spent\nsomewhere else. Compare the plateau against your card’s datasheet bandwidth; if\nyou are near it, you are done.\u003C\u002Fp>\n\u003Cp>A sweep that looks like this is the common case for a simple memory-bound\nkernel. Sweeps that do \u003Cem>not\u003C\u002Fem> plateau are the interesting ones.\u003C\u002Fp>\n\u003Ch2 id=\"a-procedure\">A procedure\u003C\u002Fh2>\n\u003Cp>Given no autotuner, this is the honest loop:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Start at 128 or 256 threads\u003C\u002Fstrong>, a power of two, a multiple of 32.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Get it correct.\u003C\u002Fstrong> Do not tune a kernel that is wrong.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Sweep.\u003C\u002Fstrong> 64, 128, 256, 512. It is one constant and a rebuild each — the\n\u003Ccode>id\u003C\u002Fcode> from Chapter 8 keeps the compiled variants apart automatically.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Measure properly.\u003C\u002Fstrong> Chapter 18. An unwarmed first run measures the\ncompiler.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Write the winner down, with the card it won on.\u003C\u002Fstrong> Six months later nobody\nremembers whether 256 was measured or guessed.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Re-measure on a new target.\u003C\u002Fstrong> Chapter 24: correctness transfers between\ncards, performance does not.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>For tiled kernels the sweep is two- or three-dimensional and gets expensive\nquickly. Fix the tile shape from the problem — the fused conv kernels use 32 for\nall three of \u003Ccode>BLOCK_M\u003C\u002Fcode>, \u003Ccode>BLOCK_N\u003C\u002Fcode>, \u003Ccode>BLOCK_K\u003C\u002Fcode> — then sweep the thread count\nalone.\u003C\u002Fp>\n\u003Ch2 id=\"where-the-numbers-in-this-tree-came-from\">Where the numbers in this tree came from\u003C\u002Fh2>\n\u003Cp>The shape-based dispatch thresholds in \u003Ccode>kernels\u002Fteeny-kernels\u002Fsrc\u002Fgraph\u002Fmod.rs\u003C\u002Fcode>\n— GEMM kernel for 1×1 convolutions with at least 32 output channels, tiled above\n16 — are hand-picked constants from exactly this process. \u003Cem>Which\u003C\u002Fem> kernel runs is\nstill a hand-picked threshold; only the tile size inside it is now derived, and\nonly when \u003Ccode>sm_count\u003C\u002Fcode> is set.\u003C\u002Fp>\n\u003Cp>The bench beside them exists to check they still hold. Its doc comment says so:\nthe shapes it benchmarks were chosen to straddle those thresholds, so the\nmeasurements say whether the dispatch still picks the right kernel.\u003C\u002Fp>\n\u003Cp>That is the pattern to copy. A tuned constant with no benchmark defending it\nbecomes folklore within a release or two.\u003C\u002Fp>\n\u003Cp>Next: the choice that usually matters more than the block size.\u003C\u002Fp>\n",[12,16,19,22,25,29,32,35],{"id":13,"text":14,"level":15},"two-different-numbers","Two different numbers",2,{"id":17,"text":18,"level":15},"what-you-control","What you control",{"id":20,"text":21,"level":15},"what-you-do-not-control","What you do not control",{"id":23,"text":24,"level":15},"occupancy","Occupancy",{"id":26,"text":27,"level":28},"the-other-kind-of-occupancy-problem","The other kind of occupancy problem",3,{"id":30,"text":31,"level":15},"a-measured-sweep","A measured sweep",{"id":33,"text":34,"level":15},"a-procedure","A procedure",{"id":36,"text":37,"level":15},"where-the-numbers-in-this-tree-came-from","Where the numbers in this tree came from",false,{"title":40,"titleHtml":40,"route":41},"Compile-Time Parameters and Dtype Dispatch","\u002Fkernels\u002Fpatterns\u002Fspecialisation",{"title":43,"titleHtml":43,"route":44},"Memory Coalescing and Tensor Layout","\u002Fkernels\u002Ffast\u002Flayout",1786271830016]