[{"data":1,"prerenderedAt":39},["ShallowReactive",2],{"chapter:kernels\u002Ffirst-kernel\u002Fcompiling.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":31,"prev":32,"next":36},"kernels","\u002Fkernels\u002Ffirst-kernel\u002Fcompiling","Compiling and Reading the Output","Your First Kernel","first-kernel\u002Fcompiling.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Ffirst-kernel\u002Fcompiling.md","\u003Cp>The pipeline from Chapter 3 has been a diagram until now. This chapter opens it\nup, because a kernel you can read the output of is a kernel you can debug.\u003C\u002Fp>\n\u003Ch2 id=\"compiling\">Compiling\u003C\u002Fh2>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> ptx_path \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\"> compile_kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Target\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">new\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">env\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">capability\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> false\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)?;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Three arguments:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>the kernel\u003C\u002Fstrong>, which supplies the source text and its id,\u003C\u002Fli>\n\u003Cli>\u003Cstrong>the target\u003C\u002Fstrong>, a compute capability such as \u003Ccode>sm_89\u003C\u002Fcode>,\u003C\u002Fli>\n\u003Cli>\u003Cstrong>\u003Ccode>force\u003C\u002Fcode>\u003C\u002Fstrong>, which recompiles even when a cached result exists.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>It returns a path, to a file with a \u003Ccode>.o\u003C\u002Fcode> extension that contains PTX text —\n\u003Ccode>\u002F\u002F Generated by LLVM NVPTX Back-End\u003C\u002Fcode> is its first line. Beside it, same name,\nsit two more: \u003Ccode>.mlir\u003C\u002Fcode> and \u003Ccode>.rs\u003C\u002Fcode>. The MLIR is the most useful of the three, and\nthe \u003Ccode>.rs\u003C\u002Fcode> is the captured source from Chapter 8.\u003C\u002Fp>\n\u003Cp>Compilation is cached under \u003Ccode>TEENYC_CACHE_DIR\u003C\u002Fcode>, keyed by the kernel’s \u003Ccode>id\u003C\u002Fcode>. The\nfirst call shells out to \u003Ccode>teenyc\u003C\u002Fcode>; later calls with the same id return\nimmediately. This matters when you benchmark: an uncached first run measures the\ncompiler, not the kernel. Chapter 18 comes back to it.\u003C\u002Fp>\n\u003Ch2 id=\"reading-the-mlir\">Reading the MLIR\u003C\u002Fh2>\n\u003Cp>This is the compiler’s own record of what it understood your kernel to mean. It\nis close enough to the source to check line by line, and specific enough to show\nyou what actually happens.\u003C\u002Fp>\n\u003Cp>Here is the real MLIR for \u003Ccode>elemwise_add_forward\u003C\u002Fcode> — the library kernel that\n\u003Ccode>vector_add\u003C\u002Fcode> is a copy of — with \u003Ccode>f32\u003C\u002Fcode> and a block size of 128:\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"mlir\">\u003Ccode>%c128_i32 = arith.constant 128 : i32\n%0 = tt.get_program_id x : i32\n%1 = arith.muli %0, %c128_i32 : i32\n%2 = tt.make_range {end = 128 : i32, start = 0 : i32} : tensor&lt;128xi32&gt;\n%3 = tt.splat %1 : i32 -&gt; tensor&lt;128xi32&gt;\n%4 = arith.addi %2, %3 : tensor&lt;128xi32&gt;\n%5 = tt.splat %arg3 : i32 -&gt; tensor&lt;128xi32&gt;\n%6 = arith.cmpi slt, %4, %5 : tensor&lt;128xi32&gt;\n%7 = tt.splat %arg0 : !tt.ptr&lt;f32&gt; -&gt; tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;\n%8 = tt.addptr %7, %4 : tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;, tensor&lt;128xi32&gt;\n%9 = tt.load %8, %6 : tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;\n%10 = tt.splat %arg1 : !tt.ptr&lt;f32&gt; -&gt; tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;\n%11 = tt.addptr %10, %4 : tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;, tensor&lt;128xi32&gt;\n%12 = tt.load %11, %6 : tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;\n%13 = tt.splat %arg2 : !tt.ptr&lt;f32&gt; -&gt; tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;\n%14 = tt.addptr %13, %4 : tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;, tensor&lt;128xi32&gt;\n%15 = arith.addf %9, %12 : tensor&lt;128xf32&gt;\ntt.store %14, %15, %6 {operandSegmentSizes = array&lt;i32: 1, 1, 1&gt;} : tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;\ntt.return\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cem>From \u003Ccode>kernels\u002Fteeny-kernels\u002Ftests\u002Fsnapshots\u002Ftest_elemwise_add__elemwise_add_forward_mlir.snap\u003C\u002Fcode>.\u003C\u002Fem>\u003C\u002Fp>\n\u003Cp>Now line it up against what you wrote:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Your Rust\u003C\u002Fth>\n\u003Cth>The MLIR\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>T::program_id(Axis::X)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>tt.get_program_id x\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>pid * BLOCK_SIZE\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>arith.muli %0, %c128_i32\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>T::arange(0, BLOCK_SIZE)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>tt.make_range {start = 0, end = 128}\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>+ block_start\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>tt.splat\u003C\u002Fcode> then \u003Ccode>arith.addi\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>offsets.lt(n_elements)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>tt.splat\u003C\u002Fcode> then \u003Ccode>arith.cmpi slt\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>a_ptr.add_offsets(offsets)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>tt.splat\u003C\u002Fcode> then \u003Ccode>tt.addptr\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>T::load(..., Some(in_bounds), ...)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>tt.load %8, %6\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>a + b\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>arith.addf\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>T::store(...)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>tt.store %14, %15, %6\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>Almost one to one. Four things are worth noticing.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>\u003Ccode>BLOCK_SIZE\u003C\u002Fcode> is gone.\u003C\u002Fstrong> It is \u003Ccode>128\u003C\u002Fcode>, a constant, everywhere it appears — in\nthe multiply, in the range, in every tensor type. This is what “baked in at\ncompile time” means concretely.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Types carry the shape.\u003C\u002Fstrong> \u003Ccode>tensor&lt;128xi32&gt;\u003C\u002Fcode> is a block of 128 integers;\n\u003Ccode>tensor&lt;128x!tt.ptr&lt;f32&gt;&gt;\u003C\u002Fcode> is a block of 128 pointers to \u003Ccode>f32\u003C\u002Fcode>. The type system\nis tracking your blocks all the way down, and a mismatch here is a mismatch you\nwould have got as a Rust type error first.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>\u003Ccode>tt.splat\u003C\u002Fcode> is broadcasting.\u003C\u002Fstrong> Whenever a scalar meets a block, it is copied\nacross every lane. \u003Ccode>block_start\u003C\u002Fcode> is one integer in your Rust and a\n\u003Ccode>tensor&lt;128xi32&gt;\u003C\u002Fcode> here. Three of these appear because three scalars —\n\u003Ccode>block_start\u003C\u002Fcode>, \u003Ccode>n_elements\u003C\u002Fcode>, and each base pointer — get broadcast.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The mask is an argument to the memory operations.\u003C\u002Fstrong> \u003Ccode>tt.load %8, %6\u003C\u002Fcode> — address\ntensor, then mask. Nothing branches. That is the whole implementation of the\nbounds check from Chapter 7: not a jump, just an operand.\u003C\u002Fp>\n\u003Ch2 id=\"the-two-functions\">The two functions\u003C\u002Fh2>\n\u003Cp>The full file has two functions, not one. Your kernel appears under a long\nmangled name, and beside it sits:\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"mlir\">\u003Ccode>tt.func public @elemwise_add_forward_entry_point(...) {\n  tt.call @_RINvCslSnLtkXmXla_85elemwise_add_forward_7ef7...(%arg0, %arg1, %arg2, %arg3)\n  tt.return\n}\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>That is the wrapper from Chapter 8, doing its one job: giving the loader a\npredictable symbol to look up. \u003Ccode>CudaProgram::try_from_ptx\u003C\u002Fcode> resolves\n\u003Ccode>{name}_entry_point\u003C\u002Fcode> and gets a function pointer to the thing that calls your\nkernel.\u003C\u002Fp>\n\u003Ch2 id=\"snapshot-tests\">Snapshot tests\u003C\u002Fh2>\n\u003Cp>The output above is not a screenshot. It is a committed test fixture, and the\npattern behind it is worth stealing:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> kernel \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> ElemwiseAddForward\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">f32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">new\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">BLOCK_SIZE\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> target \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> Target\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">new\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Capability\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Sm89\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> ptx_path \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> PathBuf\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">from\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">compile_kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">kernel\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> &#x26;\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">target\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> true\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)?);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> mlir \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> std\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">fs\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">read_to_string\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">ptx_path\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">with_extension\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">\"\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\">mlir\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">\"\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">))?;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">assert_debug_snapshot!\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">\"\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\">elemwise_add_forward_mlir\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">\"\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> mlir\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">trim\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">());\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>This needs \u003Ccode>teenyc\u003C\u002Fcode> but \u003Cstrong>not a GPU\u003C\u002Fstrong> — it compiles, it does not run. So it is\nthe strongest check available on a machine without a card, and it makes any\nchange to the generated code visible in a diff. If you change a kernel and the\nsnapshot moves in a way you did not expect, something happened that you did not\nintend.\u003C\u002Fp>\n\u003Cp>Note \u003Ccode>Target::new(Capability::Sm89)\u003C\u002Fcode> — a fixed capability rather than the local\ndevice’s, so the target is the same everywhere.\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>\u003Cstrong>These snapshots are not portable between machines.\u003C\u002Fstrong> The MLIR embeds the\nmangled Rust symbol of your kernel, and that name contains rustc’s crate\ndisambiguator — a hash of the crate’s build environment, not of its source.\nCheck out this tree on a second machine and every MLIR snapshot fails, with a\ndiff whose only difference is \u003Ccode>Cs&lt;something&gt;\u003C\u002Fcode> against \u003Ccode>Cs&lt;something else&gt;\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>The kernel body in the diff will be byte-identical. If that is all you see,\nnothing is wrong with your kernel. Normalising the disambiguator before\ncomparing would fix it; today nothing does.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Cp>So treat a snapshot diff as a question, not a verdict: look at what actually\nchanged inside it before believing it.\u003C\u002Fp>\n\u003Ch2 id=\"when-something-is-wrong\">When something is wrong\u003C\u002Fh2>\n\u003Cp>A rough order to work through:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Does it compile?\u003C\u002Fstrong> A \u003Ccode>teenyc\u003C\u002Fcode> failure is about the captured text. Check\nthat you have not used anything outside the DSL — Chapter 3’s second\nconsequence.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Does the MLIR match your intent?\u003C\u002Fstrong> Count the loads and stores. Check the\nconstants. A missing mask operand on a \u003Ccode>tt.load\u003C\u002Fcode> is visible immediately.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Are the numbers wrong at the edges?\u003C\u002Fstrong> Suspect the mask, or the \u003Ccode>other\u003C\u002Fcode>\nfill value for a reduction.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Are the numbers wrong everywhere?\u003C\u002Fstrong> Suspect the index arithmetic, or the\nargument order at the launch site — nothing checks that the tuple you pass\nmatches the parameters your kernel declares.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch2 id=\"the-end-of-part-2\">The end of Part 2\u003C\u002Fh2>\n\u003Cp>You can write a kernel, run it, and read what the compiler made of it.\u003C\u002Fp>\n\u003Cp>That is the whole mechanism. Everything after this is patterns built on it: how\nto reduce across a row, how to tile a matrix multiply, how to fuse work into a\nkernel that has already paid for its loads.\u003C\u002Fp>\n\u003Cp>Part 3 starts with the first kernel where programs have to do more than mind\ntheir own slice.\u003C\u002Fp>\n",[12,16,19,22,25,28],{"id":13,"text":14,"level":15},"compiling","Compiling",2,{"id":17,"text":18,"level":15},"reading-the-mlir","Reading the MLIR",{"id":20,"text":21,"level":15},"the-two-functions","The two functions",{"id":23,"text":24,"level":15},"snapshot-tests","Snapshot tests",{"id":26,"text":27,"level":15},"when-something-is-wrong","When something is wrong",{"id":29,"text":30,"level":15},"the-end-of-part-2","The end of Part 2",false,{"title":33,"titleHtml":34,"route":35},"What #[kernel] Generates","What \u003Ccode>#[kernel]\u003C\u002Fcode> Generates","\u002Fkernels\u002Ffirst-kernel\u002Fkernel-macro",{"title":37,"titleHtml":37,"route":38},"Softmax: Your First Reduction","\u002Fkernels\u002Fpatterns\u002Fsoftmax",1786271829691]