[{"data":1,"prerenderedAt":35},["ShallowReactive",2],{"chapter:kernels\u002Forientation\u002Fblock-not-thread.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":28,"prev":29,"next":32},"kernels","\u002Fkernels\u002Forientation\u002Fblock-not-thread","You Program a Block, Not a Thread","Orientation","orientation\u002Fblock-not-thread.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Forientation\u002Fblock-not-thread.md","\u003Cp>There are two ways to write the “one worker’s job” from the last chapter. This\nchapter is about which one teenygrad uses, and why the difference matters more\nthan it first appears.\u003C\u002Fp>\n\u003Ch2 id=\"the-two-models\">The two models\u003C\u002Fh2>\n\u003Cp>In CUDA, the unit you write for is \u003Cstrong>one thread handling one element\u003C\u002Fstrong>. Your\ncode says “I am thread number 4,721, so I will add element 4,721”. The hardware\nstarts millions of these. If you want a thread to cooperate with its neighbours\n— to sum a row, say — you arrange that yourself, with shared memory and explicit\nbarriers.\u003C\u002Fp>\n\u003Cp>This is called SIMT: Single Instruction, Multiple Threads.\u003C\u002Fp>\n\u003Cp>In Triton, and so in teenygrad, the unit you write for is \u003Cstrong>one program handling\na block of elements\u003C\u002Fstrong>. Your code says “I am program number 36, so I will handle\nelements 4,608 through 4,735”. You do not write anything per-element. You write\noperations on the whole block at once, and the compiler works out how to spread\nthat across the actual hardware threads.\u003C\u002Fp>\n\u003Cp>If you have used NumPy, you already have the right instinct. \u003Ccode>a + b\u003C\u002Fcode> on two\narrays does not make you write a loop. Neither does a Triton kernel.\u003C\u002Fp>\n\u003Ch2 id=\"seeing-it\">Seeing it\u003C\u002Fh2>\n\u003Cp>Here is the beginning of the vector-add kernel from Chapter 5:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\">    \u002F\u002F Which slice of the vector is this program responsible for?\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pid \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">program_id\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Axis\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">X\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> block_start \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> pid \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">*\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\">    \u002F\u002F The indices it will touch: block_start, block_start + 1, and so on.\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">    let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> offsets \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">arange\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> +\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> block_start\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Two of those three values are plain integers. The third is not.\u003C\u002Fp>\n\u003Cp>\u003Ccode>pid\u003C\u002Fcode> is a single integer — which program am I. \u003Ccode>block_start\u003C\u002Fcode> is a single\ninteger — where my slice begins. But \u003Ccode>offsets\u003C\u002Fcode> is a \u003Cstrong>tensor\u003C\u002Fstrong>: \u003Ccode>arange(0, 128)\u003C\u002Fcode>\nproduces all 128 indices at once, and adding \u003Ccode>block_start\u003C\u002Fcode> shifts every one of\nthem. From there on, every operation in the kernel works on 128 values in\nparallel without a loop in sight.\u003C\u002Fp>\n\u003Cp>The picture, for a vector of 1000 elements in blocks of 128:\u003C\u002Fp>\n\u003Cpre class=\"mermaid\" data-mermaid>flowchart TD\n    G[&quot;grid: 8 programs&quot;]\n    G --&gt; P0[&quot;program 0&lt;br\u002F&gt;offsets 0..127&quot;]\n    G --&gt; P1[&quot;program 1&lt;br\u002F&gt;offsets 128..255&quot;]\n    G --&gt; Pd[&quot;…&quot;]\n    G --&gt; P7[&quot;program 7&lt;br\u002F&gt;offsets 896..1023&lt;br\u002F&gt;24 lanes masked off&quot;]\n\u003C\u002Fpre>\n\u003Cp>Each box runs the same code. Each computes a different \u003Ccode>pid\u003C\u002Fcode>, and therefore a\ndifferent \u003Ccode>offsets\u003C\u002Fcode>. Nothing coordinates them, and they may run in any order or\nall at once.\u003C\u002Fp>\n\u003Ch2 id=\"why-this-is-the-better-default\">Why this is the better default\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>Cooperation comes free.\u003C\u002Fstrong> Summing a row in CUDA means a shared-memory\nreduction: allocate scratch, have each thread write its partial, synchronise,\nhave half the threads combine pairs, synchronise again, repeat. In Triton it is\n\u003Ccode>T::sum(x, None, false)\u003C\u002Fcode>. The compiler emits the same shuffles and barriers; you\ndo not write them, and you cannot get them subtly wrong.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The compiler can see what you meant.\u003C\u002Fstrong> Because you said “load these 128\ncontiguous addresses” rather than “load this one address” a hundred and\ntwenty-eight times, the compiler knows the access pattern and can combine it\ninto the smallest number of memory transactions. Recovering that from\nper-thread code is much harder.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>It is far less code.\u003C\u002Fstrong> Most of the ceremony in a CUDA kernel is bookkeeping —\nindex arithmetic, bounds checks, barriers — that the block model removes.\u003C\u002Fp>\n\u003Cp>The cost is control. There are hand-tuned CUDA kernels that beat what this model\nwill produce, and if you need one of those, you need CUDA. For nearly everything\nelse the trade is worth it.\u003C\u002Fp>\n\u003Ch2 id=\"the-words-and-where-they-leak\">The words, and where they leak\u003C\u002Fh2>\n\u003Cp>Four terms, used consistently from here on:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Term\u003C\u002Fth>\n\u003Cth>What it is\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Cstrong>program\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>One instance of your kernel. What your code describes.\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Cstrong>block\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>The slice of data one program handles. \u003Ccode>BLOCK_SIZE\u003C\u002Fcode> elements.\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Cstrong>grid\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>How many programs to launch. Computed on the CPU, before launch.\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Cstrong>lane\u003C\u002Fstrong>\u003C\u002Ftd>\n\u003Ctd>One element’s position within a block.\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>You will also meet CUDA’s vocabulary, because teenygrad sits on top of CUDA and\ndoes not hide it completely. Three terms in particular:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>A \u003Cstrong>thread\u003C\u002Fstrong> is the hardware’s unit. A program is implemented as a group of\nthreads. When you set a block size of 128, you are also saying “128 threads”.\u003C\u002Fli>\n\u003Cli>A \u003Cstrong>warp\u003C\u002Fstrong> is 32 threads that execute in lockstep, always. This is why block\nsizes are multiples of 32 — a block of 100 wastes most of a fourth warp.\u003C\u002Fli>\n\u003Cli>A \u003Cstrong>CTA\u003C\u002Fstrong> (cooperative thread array) is CUDA’s name for what Triton calls a\nprogram. It shows up in this SDK’s API: \u003Ccode>RuntimeOp::grid\u003C\u002Fcode> is documented as the\n“number of CTAs to launch”.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>You do not need to think in threads to write kernels. You do need to recognise\nthe words when the API or an error message uses them.\u003C\u002Fp>\n\u003Ch2 id=\"one-thing-to-carry-forward\">One thing to carry forward\u003C\u002Fh2>\n\u003Cp>When you read a kernel in this book, read it as \u003Cstrong>one\u003C\u002Fstrong> program. Ask “what is\nthis one doing, and how does it know which part is its?” — never “how do all of\nthem work together?”, because they do not. They are independent by construction.\u003C\u002Fp>\n\u003Cp>Next: what actually happens to your Rust when you write it.\u003C\u002Fp>\n",[12,16,19,22,25],{"id":13,"text":14,"level":15},"the-two-models","The two models",2,{"id":17,"text":18,"level":15},"seeing-it","Seeing it",{"id":20,"text":21,"level":15},"why-this-is-the-better-default","Why this is the better default",{"id":23,"text":24,"level":15},"the-words-and-where-they-leak","The words, and where they leak",{"id":26,"text":27,"level":15},"one-thing-to-carry-forward","One thing to carry forward",true,{"title":30,"titleHtml":30,"route":31},"What a Kernel Is","\u002Fkernels\u002Forientation\u002Fwhat-a-kernel-is",{"title":33,"titleHtml":33,"route":34},"From Rust to PTX","\u002Fkernels\u002Forientation\u002Frust-to-ptx",1786271829627]