[{"data":1,"prerenderedAt":41},["ShallowReactive",2],{"chapter:kernels\u002Fin-a-model\u002Fportability.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":34,"prev":35,"next":38},"kernels","\u002Fkernels\u002Fin-a-model\u002Fportability","What Is Portable","Kernels in a Real Model","in-a-model\u002Fportability.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Fin-a-model\u002Fportability.md","\u003Cp>A book about a portable-looking abstraction owes you a straight answer about how\nportable it actually is. This chapter is that answer, as of the code this book\nwas written against.\u003C\u002Fp>\n\u003Ch2 id=\"the-short-version\">The short version\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>Your kernel bodies are portable. Everything around them is CUDA.\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp>The \u003Ccode>Triton\u003C\u002Fcode> trait is a genuine abstraction — a kernel written against it names\nno vendor and no device. But there is exactly one driver crate, \u003Ccode>teeny-cuda\u003C\u002Fcode>,\nand the launch path, the buffers, the compilation target and the capability enum\nare all NVIDIA.\u003C\u002Fp>\n\u003Cp>So the portability is real but latent: the kernels are ready for a second\nbackend that does not exist yet.\u003C\u002Fp>\n\u003Ch2 id=\"line-by-line\">Line by line\u003C\u002Fh2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Layer\u003C\u002Fth>\n\u003Cth>Portable?\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>Kernel body (\u003Ccode>Triton\u003C\u002Fcode> trait)\u003C\u002Ftd>\n\u003Ctd>Yes — no vendor in the API\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>#[kernel]\u003C\u002Fcode> and the generated struct\u003C\u002Ftd>\n\u003Ctd>Yes — device-independent\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>CustomOp\u003C\u002Fcode>, the graph, \u003Ccode>SymTensor\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Yes — no device concepts\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>RuntimeOp\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Mostly — \u003Ccode>block\u003C\u002Fcode>\u002F\u003Ccode>grid\u003C\u002Fcode> are a CUDA-shaped model\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>compile_kernel\u003C\u002Fcode>, \u003Ccode>Target\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>No — \u003Ccode>driver::cuda\u003C\u002Fcode>, \u003Ccode>target::cuda\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Capability\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>No — \u003Ccode>sm_75\u003C\u002Fcode>…\u003Ccode>sm_120\u003C\u002Fcode>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Device\u003C\u002Fcode>, \u003Ccode>Buffer\u003C\u002Fcode>, \u003Ccode>launch\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>No — \u003Ccode>teeny-cuda\u003C\u002Fcode> is the only implementation\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>PTX\u003C\u002Ftd>\n\u003Ctd>No — NVIDIA’s assembly\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>The seam is clean and in a sensible place. A second driver would need a new\n\u003Ccode>Device\u003C\u002Fcode>\u002F\u003Ccode>Buffer\u003C\u002Fcode>\u002F\u003Ccode>Program\u003C\u002Fcode> implementation, a target description, and a\ncompilation path. It would not need your kernels rewritten.\u003C\u002Fp>\n\u003Ch2 id=\"what-exists-today\">What exists today\u003C\u002Fh2>\n\u003Cp>Compilation backends, in \u003Ccode>teeny-compiler\u003C\u002Fcode>:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>\u003Ccode>llvm\u003C\u002Fcode>\u003C\u002Fstrong> — the real one. Source → MLIR → Triton passes → LLVM → PTX.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>\u003Ccode>ndarray\u003C\u002Fcode>\u003C\u002Fstrong> — a CPU path, on by default, for running graphs without a GPU.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Device drivers, in \u003Ccode>drivers\u002F\u003C\u002Fcode>:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>\u003Ccode>teeny-cuda\u003C\u002Fcode>\u003C\u002Fstrong> — the only one.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>There is no Vulkan backend, no ROCm backend, no dedicated CPU driver. The\nexisting teenygrad book is straight about this: \u003Ccode>teeny-cpu\u003C\u002Fcode> and \u003Ccode>teeny-vulkan\u003C\u002Fcode>\nare roadmap items, and the \u003Ccode>ndarray\u003C\u002Fcode> path is the current CPU story — which is a\ndifferent thing from a driver crate.\u003C\u002Fp>\n\u003Ch2 id=\"portable-within-cuda\">Portable within CUDA\u003C\u002Fh2>\n\u003Cp>Between NVIDIA generations, most things do carry:\u003C\u002Fp>\n\u003Cp>\u003Cstrong>PTX is forward-compatible.\u003C\u002Fstrong> Built for \u003Ccode>sm_75\u003C\u002Fcode>, it runs on anything newer,\nbecause the driver compiles it for the actual chip at load time.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Instructions are not backward-compatible.\u003C\u002Fstrong> A kernel that uses tensor\ndescriptors — the TMA path from Chapter 11 — needs hardware that has TMA. Build\nit for an older capability and you get either different code or a failure.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Block sizes do not transfer.\u003C\u002Fstrong> A block size tuned on an A100 is not the right\none for an Orin: different register files, different memory bandwidth, different\ncore counts. Nothing warns you; the kernel just runs slower than it should.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>PTX versions can be rejected.\u003C\u002Fstrong> Chapter 23’s \u003Ccode>sm_120a\u003C\u002Fcode> case: a driver refusing\na PTX version newer than it knows. Forward compatibility has limits at both\nends.\u003C\u002Fp>\n\u003Cp>So “portable across NVIDIA” means “will run”, not “will run well”. Retune per\ntarget, or accept that you have optimised for one card.\u003C\u002Fp>\n\u003Ch2 id=\"what-would-not-survive-a-second-backend\">What would not survive a second backend\u003C\u002Fh2>\n\u003Cp>If a Vulkan or ROCm driver arrived tomorrow, these would need attention:\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Anything assuming warps of 32.\u003C\u002Fstrong> AMD’s wavefronts are 64. Chapter 6’s\n“multiple of 32” rule is NVIDIA’s number, and a block size chosen around it is\nNVIDIA-shaped.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Tensor Core specifics.\u003C\u002Fstrong> \u003Ccode>InputPrecision::TF32\u003C\u002Fcode> names an NVIDIA feature. Other\nvendors have matrix units with different precision modes.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>PTX-level anything.\u003C\u002Fstrong> \u003Ccode>T::inline_asm_elementwise\u003C\u002Fcode> takes an assembly string.\nUsing it ends portability at that line, deliberately and obviously.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>The capability enum.\u003C\u002Fstrong> \u003Ccode>Capability\u003C\u002Fcode> is a list of SM versions. A second vendor\nneeds a different type, or that one generalised.\u003C\u002Fp>\n\u003Cp>None of this is unusual — it is the normal cost of a portable layer over\nhardware that is not actually alike. It is worth knowing which of your choices\nare the portable kind.\u003C\u002Fp>\n\u003Ch2 id=\"practical-advice\">Practical advice\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>Write kernels against \u003Ccode>Triton\u003C\u002Fcode> only.\u003C\u002Fstrong> Avoid inline assembly unless you have\nmeasured that you need it, and mark it loudly where you do.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Keep the tuning constants together.\u003C\u002Fstrong> Block sizes and tile shapes are the\nper-device numbers. If they are named constants in one place, retuning for a new\ncard is an afternoon. If they are scattered through kernel bodies, it is a week.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Do not build portability you cannot test.\u003C\u002Fstrong> With one driver, an abstraction\n“for the second backend” is untested by construction, and untested abstractions\nare usually wrong. The \u003Ccode>Triton\u003C\u002Fcode> trait is enough.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Assume you will retune.\u003C\u002Fstrong> Correctness transfers between NVIDIA cards.\nPerformance does not.\u003C\u002Fp>\n\u003Ch2 id=\"end-of-part-5\">End of Part 5\u003C\u002Fh2>\n\u003Cp>Your kernel can now be a node in a model, produce gradients, and be built for a\nboard you have never touched.\u003C\u002Fp>\n\u003Cp>Part 6 is reference material: the Python Triton translation table, the compile\nerrors you will actually hit, a glossary, and a worked port of a Python kernel.\u003C\u002Fp>\n",[12,16,19,22,25,28,31],{"id":13,"text":14,"level":15},"the-short-version","The short version",2,{"id":17,"text":18,"level":15},"line-by-line","Line by line",{"id":20,"text":21,"level":15},"what-exists-today","What exists today",{"id":23,"text":24,"level":15},"portable-within-cuda","Portable within CUDA",{"id":26,"text":27,"level":15},"what-would-not-survive-a-second-backend","What would not survive a second backend",{"id":29,"text":30,"level":15},"practical-advice","Practical advice",{"id":32,"text":33,"level":15},"end-of-part-5","End of Part 5",false,{"title":36,"titleHtml":36,"route":37},"Building for Another Target","\u002Fkernels\u002Fin-a-model\u002Fcross-building",{"title":39,"titleHtml":39,"route":40},"Python Triton to Rust","\u002Fkernels\u002Freference\u002Ftranslation-table",1786271830119]