[{"data":1,"prerenderedAt":20},["ShallowReactive",2],{"chapter:teenygrad\u002Fkernels-and-backends\u002Fbackends.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":12,"prev":13,"next":17},"teenygrad","\u002Fteenygrad\u002Fkernels-and-backends\u002Fbackends","CPU, CUDA, and Vulkan Backends","Kernels & Backends","kernels-and-backends\u002Fbackends.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fteenygrad\u002Fsrc\u002Fkernels-and-backends\u002Fbackends.md","\u003Cp>Device backends live under \u003Ccode>drivers\u002F\u003C\u002Fcode> in the workspace. Today, only\n\u003Ca href=\"\u002Fapi\u002Fteenygrad\u002Fteenygrad\u002Fteeny_cuda\u002F\">\u003Ccode>teeny-cuda\u003C\u002Fcode>\u003C\u002Fa> (NVIDIA, via \u003Ccode>bindgen\u003C\u002Fcode>-generated\ndriver bindings) is implemented. \u003Ccode>teeny-cpu\u003C\u002Fcode> and \u003Ccode>teeny-vulkan\u003C\u002Fcode> are on the\n\u003Ca href=\"\u002Fteenygrad\u002Fappendix\u002Ffaq-and-roadmap\">roadmap\u003C\u002Fa> but don’t exist yet — the \u003Ccode>ndarray\u003C\u002Fcode>-backed CPU path in\n\u003Ccode>teeny-compiler\u003C\u002Fcode> (the \u003Ccode>ndarray\u003C\u002Fcode> feature, on by default) is the current CPU story, distinct from a\ndedicated \u003Ccode>teeny-cpu\u003C\u002Fcode> driver crate.\u003C\u002Fp>\n\u003Cp>\u003Ccode>teeny-kernels\u003C\u002Fcode>’ \u003Ccode>cuda\u003C\u002Fcode> feature (on by default) enables the \u003Ccode>teeny-cuda\u003C\u002Fcode> backend for its kernel\nimplementations. Building it requires the CUDA toolkit — see\n\u003Ca href=\"\u002Fteenygrad\u002Fgetting-started\u002Finstallation\">Installation &amp; Toolchain\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>Kernel \u003Cem>definitions\u003C\u002Fem> (the actual per-op logic — attention, matmul, elementwise ops, etc.) are\nbackend-agnostic, written once against the \u003Ccode>teeny-triton\u003C\u002Fcode> DSL and compiled per-backend — see\n\u003Ca href=\"\u002Fteenygrad\u002Fkernels-and-backends\u002Fwriting-a-kernel\">Writing a Triton Kernel\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>\u003Cstrong>TODO\u003C\u002Fstrong>: document the \u003Ccode>teeny-cuda\u003C\u002Fcode> device\u002Fruntime API directly (device selection, memory\nmanagement, profiling via \u003Ccode>cuda_profiler_start\u003C\u002Fcode>\u002F\u003Ccode>cuda_profiler_stop\u003C\u002Fcode>) once a driver-level\nchapter is warranted.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n",[],false,{"title":14,"titleHtml":15,"route":16},"Building Models with nn","Building Models with \u003Ccode>nn\u003C\u002Fcode>","\u002Fteenygrad\u002Fnn-layers\u002Fbuilding-models",{"title":18,"titleHtml":18,"route":19},"Writing a Triton Kernel","\u002Fteenygrad\u002Fkernels-and-backends\u002Fwriting-a-kernel",1786271829003]