[{"data":1,"prerenderedAt":35},["ShallowReactive",2],{"chapter:kernels\u002Forientation\u002Frust-to-ptx.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":28,"prev":29,"next":32},"kernels","\u002Fkernels\u002Forientation\u002Frust-to-ptx","From Rust to PTX","Orientation","orientation\u002Frust-to-ptx.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Forientation\u002Frust-to-ptx.md","\u003Cp>Your kernel is a Rust function. Your GPU does not run Rust. This chapter is\nabout what happens in between.\u003C\u002Fp>\n\u003Cp>It is worth reading before you write a kernel rather than after, because the\nanswer is genuinely unusual, and almost every rule in Parts 2 and 3 is a\nconsequence of it.\u003C\u002Fp>\n\u003Ch2 id=\"the-surprise\">The surprise\u003C\u002Fh2>\n\u003Cp>Here is the thing to know:\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>Your kernel function is never called by your program. It is captured as\n\u003Cstrong>text\u003C\u002Fstrong>, and compiled by a different compiler.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Cp>When you write \u003Ccode>#[kernel]\u003C\u002Fcode> on a function, the macro keeps the function — it\ncompiles normally, as ordinary Rust, and the Rust compiler type-checks it. But\nthe macro also converts the function’s source code back into a string and stores\nthat string in a generated struct.\u003C\u002Fp>\n\u003Cp>That string is the real artefact. It gets handed to \u003Ccode>teenyc\u003C\u002Fcode>, a separate\ncompiler binary, which compiles it for the GPU.\u003C\u002Fp>\n\u003Cp>So the function is type-checked twice, by two different compilers, and executed\nby neither of them in the way you would expect.\u003C\u002Fp>\n\u003Ch2 id=\"why-do-it-that-way\">Why do it that way\u003C\u002Fh2>\n\u003Cp>Because the two compilers need different things from the same code.\u003C\u002Fp>\n\u003Cp>The Rust compiler you already have is very good at checking that your kernel\nmakes sense: that you did not pass a float where an index belongs, that the\ntensor ranks line up, that the dtype you loaded is the dtype you stored. Doing\nthat check is exactly what \u003Ccode>T: Triton\u003C\u002Fcode> and those \u003Ccode>where\u003C\u002Fcode> clauses are for. None\nof it needs a GPU.\u003C\u002Fp>\n\u003Cp>But that compiler cannot emit GPU code. Producing something a graphics card will\nrun means going through MLIR and the Triton compiler passes, and that is what\n\u003Ccode>teenyc\u003C\u002Fcode> — a modified rustc — exists to do.\u003C\u002Fp>\n\u003Cp>Capturing the source is the seam between the two. You get Rust’s type checking\non kernels, from an ordinary \u003Ccode>cargo check\u003C\u002Fcode>, without the GPU toolchain being\ninvolved at all.\u003C\u002Fp>\n\u003Ch2 id=\"the-pipeline\">The pipeline\u003C\u002Fh2>\n\u003Cpre class=\"mermaid\" data-mermaid>flowchart TD\n    A[&quot;your #[kernel] fn&lt;br\u002F&gt;&lt;i&gt;Rust&lt;\u002Fi&gt;&quot;] --&gt;|&quot;macro captures source text&quot;| B[&quot;kernel source + generated entry point&lt;br\u002F&gt;&lt;i&gt;a String&lt;\u002Fi&gt;&quot;]\n    B --&gt;|teenyc| C[&quot;MLIR, Triton dialect&lt;br\u002F&gt;&lt;i&gt;tt.load, tt.store, tensors&lt;\u002Fi&gt;&quot;]\n    C --&gt;|&quot;Triton passes&quot;| D[&quot;MLIR, GPU dialects&lt;br\u002F&gt;&lt;i&gt;layouts, coalescing, pipelining&lt;\u002Fi&gt;&quot;]\n    D --&gt; E[&quot;LLVM IR&quot;]\n    E --&gt;|&quot;NVPTX backend&quot;| F[&quot;PTX&lt;br\u002F&gt;&lt;i&gt;NVIDIA assembly&lt;\u002Fi&gt;&quot;]\n    F --&gt;|&quot;driver JIT, at load time&quot;| G[&quot;SASS&lt;br\u002F&gt;&lt;i&gt;machine code for your card&lt;\u002Fi&gt;&quot;]\n\u003C\u002Fpre>\n\u003Cp>The stages that matter to you:\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The captured string.\u003C\u002Fstrong> Source text, plus a small wrapper the macro generates.\nChapter 8 shows it in full.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>MLIR in the Triton dialect.\u003C\u002Fstrong> Your kernel, still recognisable, expressed as\noperations on tensors. This is the most useful thing to look at when a kernel\nmisbehaves, and Chapter 9 reads one.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The Triton passes.\u003C\u002Fstrong> Where the compiler decides how your block maps onto real\nthreads, how loads are combined, and where data lives. This is the part you do\nnot control directly, and mostly should not want to.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>PTX.\u003C\u002Fstrong> NVIDIA’s portable assembly. This is what \u003Ccode>compile_kernel\u003C\u002Fcode> gives you back\n— a \u003Ccode>.ptx\u003C\u002Fcode> file on disk.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>SASS.\u003C\u002Fstrong> The actual machine code, produced by the driver when the PTX is\nloaded, targeted at the exact chip present. You never see this, and it is why a\nsingle PTX file works across several card generations.\u003C\u002Fp>\n\u003Ch2 id=\"what-this-costs-you\">What this costs you\u003C\u002Fh2>\n\u003Cp>Four consequences, all of which you will meet.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The kernel body is compiled in a different world.\u003C\u002Fstrong> \u003Ccode>teenyc\u003C\u002Fcode> compiles your\ncaptured text against a small generated environment — not against your crate,\nand not against the real standard library. So a \u003Ccode>println!\u003C\u002Fcode> in a kernel body will\ntype-check happily in your editor and then fail in an unfamiliar compiler. You\ncannot call your own helper functions from a kernel body either, unless they\nare part of that environment.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Everything must be knowable from the text.\u003C\u002Fstrong> This is why \u003Ccode>BLOCK_SIZE\u003C\u002Fcode> is a\nconst generic rather than an argument, and why \u003Ccode>T::reduce\u003C\u002Fcode> takes a plain \u003Ccode>fn\u003C\u002Fcode>\npointer instead of a closure — a closure that captured a variable could not be\nwritten out as source. Chapter 13 comes back to this.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Compilation happens at run time, not build time.\u003C\u002Fstrong> \u003Ccode>cargo build\u003C\u002Fcode> does not\nproduce PTX. The first time your program calls \u003Ccode>compile_kernel\u003C\u002Fcode>, it shells out\nto \u003Ccode>teenyc\u003C\u002Fcode>. Results are cached on disk, keyed by the kernel’s identity, so it\nhappens once rather than every launch.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>You need \u003Ccode>teenyc\u003C\u002Fcode> installed to run anything, but not to build anything.\u003C\u002Fstrong> This\nis genuinely useful: \u003Ccode>cargo check\u003C\u002Fcode> on a crate full of kernels works on a laptop\nwith no GPU and no toolchain. Only running needs the rest.\u003C\u002Fp>\n\u003Ch2 id=\"where-the-pieces-live\">Where the pieces live\u003C\u002Fh2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Piece\u003C\u002Fth>\n\u003Cth>Crate\u003C\u002Fth>\n\u003Cth>What it does\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>#[kernel]\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>teeny-macros\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Captures the source, generates the struct and entry point\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Triton\u003C\u002Fcode> trait\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>teeny-triton\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>The operations a kernel body may use\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>compile_kernel\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>teeny-compiler\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Finds \u003Ccode>teenyc\u003C\u002Fcode>, runs it, caches the PTX\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>teenyc\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>separate toolchain\u003C\u002Ftd>\n\u003Ctd>The modified rustc that emits GPU code\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Device\u003C\u002Fcode>, \u003Ccode>launch\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>teeny-cuda\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Loads the PTX and runs it\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>Two environment variables are worth knowing now, because they are how you fix\nthings when the toolchain is not where it should be:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Ccode>TEENYC_PATH\u003C\u002Fcode> — the \u003Ccode>teenyc\u003C\u002Fcode> binary to use. Without it, the compiler looks for\na single rustup toolchain whose name contains \u003Ccode>teenyc\u003C\u002Fcode>, and fails clearly if\nthere are none or several.\u003C\u002Fli>\n\u003Cli>\u003Ccode>TEENYC_CACHE_DIR\u003C\u002Fcode> — where compiled PTX is cached. Defaults to\n\u003Ccode>\u002Ftmp\u002Fteenyc_cache\u003C\u002Fcode>.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Next: getting all of that installed.\u003C\u002Fp>\n",[12,16,19,22,25],{"id":13,"text":14,"level":15},"the-surprise","The surprise",2,{"id":17,"text":18,"level":15},"why-do-it-that-way","Why do it that way",{"id":20,"text":21,"level":15},"the-pipeline","The pipeline",{"id":23,"text":24,"level":15},"what-this-costs-you","What this costs you",{"id":26,"text":27,"level":15},"where-the-pieces-live","Where the pieces live",true,{"title":30,"titleHtml":30,"route":31},"You Program a Block, Not a Thread","\u002Fkernels\u002Forientation\u002Fblock-not-thread",{"title":33,"titleHtml":33,"route":34},"Setting Up","\u002Fkernels\u002Forientation\u002Fsetting-up",1786271829640]