[{"data":1,"prerenderedAt":36},["ShallowReactive",2],{"chapter:kernels\u002Ffirst-kernel\u002Fkernel-body.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":29,"prev":30,"next":33},"kernels","\u002Fkernels\u002Ffirst-kernel\u002Fkernel-body","The Kernel Body","Your First Kernel","first-kernel\u002Fkernel-body.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Ffirst-kernel\u002Fkernel-body.md","\u003Cp>Chapter 5 showed a whole working kernel. This chapter takes the first three\nlines apart: how a program finds its slice of the data, and how you choose how\nbig that slice should be.\u003C\u002Fp>\n\u003Ch2 id=\"finding-your-slice\">Finding your slice\u003C\u002Fh2>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\">    \u002F\u002F Which slice of the vector is this program responsible for?\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pid \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">program_id\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Axis\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">X\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> block_start \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pid \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\">    \u002F\u002F The indices it will touch: block_start, block_start + 1, and so on.\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">arange\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> +\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> block_start\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Ccode>T::program_id(Axis::X)\u003C\u002Fcode> is the only thing that differs between the copies of\nyour kernel that are running. Everything else — the code, the constants, the\npointers — is identical across all of them. That one integer is the whole\nmechanism by which parallel work gets divided up.\u003C\u002Fp>\n\u003Cp>The rest is arithmetic. If each program handles \u003Ccode>BLOCK_SIZE\u003C\u002Fcode> elements, then\nprogram \u003Ccode>pid\u003C\u002Fcode> starts at \u003Ccode>pid * BLOCK_SIZE\u003C\u002Fcode>, and the indices it touches are that\nstart plus \u003Ccode>0, 1, 2, …\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>\u003Ccode>T::arange(0, BLOCK_SIZE)\u003C\u002Fcode> produces those offsets. Note the bounds: it is a\nhalf-open range, like Rust’s \u003Ccode>0..n\u003C\u002Fcode>, so \u003Ccode>arange(0, 128)\u003C\u002Fcode> gives 0 through 127.\u003C\u002Fp>\n\u003Ch3 id=\"axes\">Axes\u003C\u002Fh3>\n\u003Cp>The grid can have up to three dimensions, and \u003Ccode>Axis\u003C\u002Fcode> selects which one you are\nasking about:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> row \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">program_id\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Axis\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">X\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> col \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">program_id\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Axis\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Y\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Use \u003Ccode>Axis::X\u003C\u002Fcode> alone until you have a reason not to. Two- and three-dimensional\ngrids are convenient for tiled kernels, but a flat grid with the index arithmetic\ndone by hand is equally fast and much easier to reason about. Several kernels in\nthis tree do exactly that — \u003Ccode>detect_decode\u003C\u002Fcode> in vision-rs launches a flat grid and\nrecovers two coordinates with a divide and a remainder:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> a_tiles \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">cdiv\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">A\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_A\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pid_b   \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">program_id\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Axis\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">X\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> \u002F\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> a_tiles\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> a_tile  \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">program_id\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Axis\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">X\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> %\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> a_tiles\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Ccode>T::num_programs(axis)\u003C\u002Fcode> tells a program how many others there are, which you\nneed when a program must loop over more data than one block holds.\u003C\u002Fp>\n\u003Ch2 id=\"choosing-a-block-size\">Choosing a block size\u003C\u002Fh2>\n\u003Cp>\u003Ccode>BLOCK_SIZE\u003C\u002Fcode> is the number of elements one program handles. It is fixed when the\nkernel is built:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> kernel \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> VectorAdd\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">f32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">new\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">128\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Three rules, in order of importance.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Make it a multiple of 32.\u003C\u002Fstrong> The hardware executes threads in groups of 32\ncalled warps, always, and a partial warp still occupies a full one. A block size\nof 100 does the work of 128 and wastes a quarter of it.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Make it a power of two.\u003C\u002Fstrong> Reductions — sums, maxima — are implemented as trees\nthat halve the working set at each step. A power of two divides evenly all the\nway down; anything else needs padding.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Start at 128 or 256.\u003C\u002Fstrong> Small blocks mean more programs, each doing less work,\nand the fixed cost of starting one starts to dominate. Large blocks mean each\nprogram needs more registers, and past a point the card can keep fewer of them\nin flight at once, so there is less work available to hide memory latency.\nBetween 128 and 512 is the usual sweet spot for simple kernels.\u003C\u002Fp>\n\u003Cp>Beyond that, measure. Chapter 16 covers what the number actually controls, and\nChapter 18 shows how to time the alternatives.\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>There is no autotuner. Python Triton has \u003Ccode>@triton.autotune\u003C\u002Fcode>, which sweeps a\nlist of configurations at run time and caches the winner per input shape.\nteenygrad has no equivalent, so the choice is yours and it is static. See\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fblob\u002Fmain\u002Fbooks\u002Fkernels\u002FKNOWN-GAPS.md\" target=\"_blank\" rel=\"noopener noreferrer\">\u003Ccode>KNOWN-GAPS.md\u003C\u002Fcode>\u003C\u002Fa>.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Ch2 id=\"why-it-is-a-const-generic\">Why it is a const generic\u003C\u002Fh2>\n\u003Cp>\u003Ccode>BLOCK_SIZE\u003C\u002Fcode> is a const generic parameter, not a function argument:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">pub\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> fn\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\"> vector_add\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> Triton\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> D\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> Num\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\"> const\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> i32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>(\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>That is not a stylistic choice. Chapter 3 explained that the kernel body is\ncaptured as text and compiled separately — and that compiler needs the block\nsize to be a literal in the text it receives. It is, quite literally: the\ngenerated entry point instantiates the kernel with the number baked in, and\nChapter 8 shows the result.\u003C\u002Fp>\n\u003Cp>This buys real things. The compiler can unroll loops whose trip count it knows,\nsize registers exactly, and fold the bounds arithmetic. It also means a kernel\nwith two block sizes is two compiled kernels — which is fine, and is how\nspecialisation works throughout this SDK. Chapter 15 covers it properly.\u003C\u002Fp>\n\u003Cp>The trade is that you cannot decide the block size from data at run time. You\ncan pick between pre-built kernels, but each one is built for its own constant.\u003C\u002Fp>\n\u003Ch2 id=\"choosing-the-grid\">Choosing the grid\u003C\u002Fh2>\n\u003Cp>The block size is baked into the kernel. The grid — how many programs to start —\nis computed on the CPU at launch, and it is the one number that depends on your\nactual data:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> cfg \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> teeny_cuda\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">testing\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">launch_config\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">N\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>which is a division, rounded up:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#E6E8E3\">grid \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> [(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">n_elements \u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\">as\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> u32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">).\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">div_ceil\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">block_size \u003C\u002Fspan>\u003Cspan style=\"color:#FF5F9E\">as\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> u32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> 1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">]\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Rounding up is what makes the last program partial. 1000 elements in blocks of\n128 gives 8 programs, and the eighth has only 104 real elements to work on. The\nother 24 lanes must not touch memory.\u003C\u002Fp>\n\u003Cp>Which is the next chapter.\u003C\u002Fp>\n",[12,16,20,23,26],{"id":13,"text":14,"level":15},"finding-your-slice","Finding your slice",2,{"id":17,"text":18,"level":19},"axes","Axes",3,{"id":21,"text":22,"level":15},"choosing-a-block-size","Choosing a block size",{"id":24,"text":25,"level":15},"why-it-is-a-const-generic","Why it is a const generic",{"id":27,"text":28,"level":15},"choosing-the-grid","Choosing the grid",false,{"title":31,"titleHtml":31,"route":32},"Vector Add, End to End","\u002Fkernels\u002Ffirst-kernel\u002Fvector-add",{"title":34,"titleHtml":34,"route":35},"Loads, Stores, and Masks","\u002Fkernels\u002Ffirst-kernel\u002Floads-stores-masks",1786271829671]