[{"data":1,"prerenderedAt":30},["ShallowReactive",2],{"chapter:kernels\u002Fintroduction.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":7,"part":8,"sourcePath":9,"editUrl":10,"html":11,"toc":12,"hasMermaid":26,"prev":8,"next":27},"kernels","\u002Fkernels","Writing GPU Kernels in Rust","Introduction",null,"introduction.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Fintroduction.md","\u003Cp>This book teaches you to write GPU kernels in Rust using teenygrad.\u003C\u002Fp>\n\u003Cp>A kernel is a small program that runs on a graphics card. You write one when the\noperation you need is not already fast — because it does not exist, or because\nit exists as three separate operations that each read and write memory when one\ncould have done the job in a single pass.\u003C\u002Fp>\n\u003Cp>Most GPU kernel work today happens in Python, through Triton. teenygrad gives\nyou the same programming model in Rust: you write a plain Rust function, and the\nteenygrad toolchain turns it into machine code your GPU runs.\u003C\u002Fp>\n\u003Ch2 id=\"who-this-is-for\">Who this is for\u003C\u002Fh2>\n\u003Cp>You are comfortable in Rust. You have never written a GPU kernel, and you may\nnever have used CUDA either.\u003C\u002Fp>\n\u003Cp>That is the whole prerequisite. Every GPU term is explained the first time it\nappears. You should not have to open the Python Triton documentation to follow\nany chapter here — if you do, that is a bug in this book, and there is an “Edit\nthis page” link at the bottom of every page.\u003C\u002Fp>\n\u003Ch2 id=\"what-you-will-be-able-to-do\">What you will be able to do\u003C\u002Fh2>\n\u003Cp>By the end of Part 2 you will have compiled and run your own kernel, and seen\nthe numbers it produced.\u003C\u002Fp>\n\u003Cp>By the end of Part 3 you will have written a softmax, a matrix multiply, and a\nkernel that fuses several operations into one pass over memory.\u003C\u002Fp>\n\u003Cp>By the end of the book you will have measured a kernel against alternatives,\nattached one to a model’s computation graph, given it a backward pass so it can\nbe trained through, and built it for a different GPU.\u003C\u002Fp>\n\u003Ch2 id=\"how-the-code-in-this-book-works\">How the code in this book works\u003C\u002Fh2>\n\u003Cp>Every code sample in this book is real code from the teenygrad repository.\nNothing is retyped into the prose — the samples are pulled straight out of the\nfiles, so a chapter cannot drift from code that builds.\u003C\u002Fp>\n\u003Cp>They come from two places. The teaching examples are runnable programs under\n\u003Ccode>kernels\u002Fteeny-triton\u002Fexamples\u002F\u003C\u002Fcode>. Later chapters teach from the library’s own\nkernels in \u003Ccode>kernels\u002Fteeny-kernels\u002Fsrc\u002F\u003C\u002Fcode> instead, because a kernel that ships is\na better thing to learn from than a copy of one.\u003C\u002Fp>\n\u003Cp>The examples you can run yourself:\u003C\u002Fp>\n\u003Cpre data-lang=\"bash\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7FB6D9\">cargo\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> run\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> -p\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> teeny-triton\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --features\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> cuda\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\"> --example\u003C\u002Fspan>\u003Cspan style=\"color:#D8A76B\"> vector_add\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>The \u003Ccode>cuda\u003C\u002Fcode> feature is what says “I have a GPU and the CUDA toolkit”. Without it\nthe examples are not built at all, so the rest of the workspace still compiles\non a laptop.\u003C\u002Fp>\n\u003Ch2 id=\"a-note-on-where-this-book-is-going\">A note on where this book is going\u003C\u002Fh2>\n\u003Cp>Parts 1 and 2 are the ones that matter most. If you cannot get from a clean\nmachine to a working kernel using only those chapters, the rest of the book has\nnot earned your time. They are written to be read in order, once.\u003C\u002Fp>\n\u003Cp>Parts 3 to 6 are closer to reference material. Read the chapter you need.\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>This book is being written. Chapters greyed out in the sidebar are drafted in\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fblob\u002Fmain\u002Fbooks\u002Fkernels\u002FOUTLINE.md\" target=\"_blank\" rel=\"noopener noreferrer\">\u003Ccode>OUTLINE.md\u003C\u002Fcode>\u003C\u002Fa>,\nalongside the exact API each one will use. Gaps between what the book wants to\nteach and what the SDK can currently do are recorded in\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fblob\u002Fmain\u002Fbooks\u002Fkernels\u002FKNOWN-GAPS.md\" target=\"_blank\" rel=\"noopener noreferrer\">\u003Ccode>KNOWN-GAPS.md\u003C\u002Fcode>\u003C\u002Fa>.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n",[13,17,20,23],{"id":14,"text":15,"level":16},"who-this-is-for","Who this is for",2,{"id":18,"text":19,"level":16},"what-you-will-be-able-to-do","What you will be able to do",{"id":21,"text":22,"level":16},"how-the-code-in-this-book-works","How the code in this book works",{"id":24,"text":25,"level":16},"a-note-on-where-this-book-is-going","A note on where this book is going",false,{"title":28,"titleHtml":28,"route":29},"What a Kernel Is","\u002Fkernels\u002Forientation\u002Fwhat-a-kernel-is",1786271829610]