[{"data":1,"prerenderedAt":19},["ShallowReactive",2],{"chapter:kernels\u002Freference\u002Fglossary.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":12,"prev":13,"next":16},"kernels","\u002Fkernels\u002Freference\u002Fglossary","Glossary","Reference","reference\u002Fglossary.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Freference\u002Fglossary.md","\u003Cp>Every GPU term this book uses, in one sentence each. Chapter references point at\nwhere the term is introduced properly.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Accumulator\u003C\u002Fstrong> — A tensor held in registers across a loop, summing partial\nresults, written to memory once at the end. Chapter 11.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Anchor\u003C\u002Fstrong> — Not a GPU term: an mdbook comment marking a region of a source file\nso a chapter can include exactly those lines.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Arithmetic intensity\u003C\u002Fstrong> — Arithmetic performed per byte loaded. Raising it is\nhow a memory-bound kernel becomes compute-bound. Chapter 11.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Atomic\u003C\u002Fstrong> — A read-modify-write that cannot be interleaved with another\nprogram’s. Chapter 14.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Backward kernel\u003C\u002Fstrong> — The kernel computing an operation’s gradient, given the\ngradient of its output. Chapter 22.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Block\u003C\u002Fstrong> — The slice of data one program handles, \u003Ccode>BLOCK_SIZE\u003C\u002Fcode> elements wide.\nChapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Block pointer\u003C\u002Fstrong> — An addressing mode carrying shape, strides and a tile shape,\nbuilt with \u003Ccode>make_block_ptr\u003C\u002Fcode>, as an alternative to explicit offsets. Chapter 17.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Broadcast\u003C\u002Fstrong> — Copying a scalar across every lane of a block so it can combine\nwith a tensor. \u003Ccode>tt.splat\u003C\u002Fcode> in the MLIR. Chapter 9.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Capability\u003C\u002Fstrong> — An NVIDIA GPU generation, written \u003Ccode>sm_75\u003C\u002Fcode> through \u003Ccode>sm_120\u003C\u002Fcode>.\nKernels are compiled for one. Chapter 4.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Coalescing\u003C\u002Fstrong> — Combining the memory accesses of many lanes into as few\ntransactions as possible. Contiguous access coalesces; strided access does not.\nChapter 17.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Compute-bound\u003C\u002Fstrong> — Limited by arithmetic rather than by memory. Chapter 1.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Const generic\u003C\u002Fstrong> — A compile-time constant parameter. How block sizes reach a\nkernel, because the value must be a literal in the captured source. Chapter 6.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Contention\u003C\u002Fstrong> — Many programs hitting the same address at once, forcing the\nhardware to serialise them. Chapter 14.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>CTA\u003C\u002Fstrong> — Cooperative Thread Array. CUDA’s name for what Triton calls a program;\nappears in \u003Ccode>RuntimeOp\u003C\u002Fcode>’s documentation. Chapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Dtype\u003C\u002Fstrong> — An element type: \u003Ccode>f32\u003C\u002Fcode>, \u003Ccode>i32\u003C\u002Fcode>, \u003Ccode>bool\u003C\u002Fcode>. \u003Ccode>DtypeRepr\u003C\u002Fcode> is its runtime,\ntype-erased form. Chapter 15.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Entry point\u003C\u002Fstrong> — The \u003Ccode>extern &quot;C&quot;\u003C\u002Fcode> wrapper the macro generates, giving the\nloader a predictable symbol, \u003Ccode>{name}_entry_point\u003C\u002Fcode>. Chapter 8.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Epilogue\u003C\u002Fstrong> — Work done to a result while it is still in registers, before\nstoring. Chapter 12.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Fusion\u003C\u002Fstrong> — Combining several operations into one kernel so intermediate\nresults never reach memory. Chapters 1 and 12.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Grid\u003C\u002Fstrong> — How many programs to launch. Computed on the CPU, from the data size.\nChapter 6.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Identity\u003C\u002Fstrong> — The value that leaves a reduction unchanged: 0 for a sum, 1 for a\nproduct, −∞ for a maximum. What masked lanes must be filled with. Chapter 10.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Kernel\u003C\u002Fstrong> — A function that runs on the GPU, executed by many programs at once.\nChapter 1.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Lane\u003C\u002Fstrong> — One element’s position within a block. Chapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Launch\u003C\u002Fstrong> — Starting a grid of programs running a compiled kernel. Chapter 5.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Lowering\u003C\u002Fstrong> — Turning a graph of operations into a DAG of compilable kernels.\nChapter 20.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Mask\u003C\u002Fstrong> — A boolean tensor saying which lanes are real. The bounds check.\nChapter 7.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Memory-bound\u003C\u002Fstrong> — Limited by moving data rather than by arithmetic. Most\nkernels. Chapter 1.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>MLIR\u003C\u002Fstrong> — The intermediate representation \u003Ccode>teenyc\u003C\u002Fcode> produces, and the most\nuseful view into what your kernel compiled to. Chapter 9.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Monomorphization\u003C\u002Fstrong> — Generating a separate compiled copy per concrete type or\nconstant. Chapter 15.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Occupancy\u003C\u002Fstrong> — How many programs a card can keep in flight at once, limited by\nregisters and shared memory per program. Chapter 16.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Program\u003C\u002Fstrong> — One instance of your kernel. What your code describes. Chapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Program ID\u003C\u002Fstrong> — The index identifying which program you are, and hence which\nslice is yours. Chapter 6.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>PTX\u003C\u002Fstrong> — NVIDIA’s portable assembly, and what \u003Ccode>compile_kernel\u003C\u002Fcode> produces.\nCompiled to machine code by the driver at load time. Chapter 3.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Race\u003C\u002Fstrong> — Two programs reading and writing the same address with no ordering,\nso one update is lost. Chapter 14.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Reduction\u003C\u002Fstrong> — Combining many values into one: sum, maximum, count. Chapter 10.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Register\u003C\u002Fstrong> — The fastest storage, private to a lane. Where tensors live inside\na kernel. Chapter 11.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>SASS\u003C\u002Fstrong> — The real machine code for a specific chip, produced from PTX by the\ndriver. Never seen directly. Chapter 3.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Scan\u003C\u002Fstrong> — A prefix operation: each output holds the reduction of everything up\nto it. Chapter 13.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Scatter\u003C\u002Fstrong> — Writing to indices computed from data rather than from the program\nid. The usual reason to need atomics. Chapter 14.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Shared memory\u003C\u002Fstrong> — Memory shared between the lanes of one program, faster than\nglobal and slower than registers. Used by reductions; not directly exposed.\nChapter 10.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>SIMT\u003C\u002Fstrong> — Single Instruction, Multiple Threads. CUDA’s model, where you write\nfor one thread. Contrast with Triton’s block model. Chapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Specialisation\u003C\u002Fstrong> — Compiling a separate kernel per constant or dtype, so the\ncompiler can use the known values. Chapter 15.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Stride\u003C\u002Fstrong> — The distance in elements between consecutive entries along a\ndimension. 1 along a row-major row; the row length along a column. Chapter 17.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Symbolic shape\u003C\u002Fstrong> — A shape with unknown dimensions, written \u003Ccode>None\u003C\u002Fcode>, resolved\nwhen real data arrives. Chapter 20.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Tensor\u003C\u002Fstrong> — In a kernel body, a block of values operated on as a unit. Not the\nframework’s \u003Ccode>SymTensor\u003C\u002Fcode>, which is a graph node handle. Chapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Tensor Core\u003C\u002Fstrong> — Hardware doing a small matrix multiply as one instruction.\nReached through \u003Ccode>T::dot\u003C\u002Fcode>. Chapter 11.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Tensor descriptor\u003C\u002Fstrong> — A TMA addressing mode built from shape, strides and a\ntile shape, loading tiles without explicit offsets. Chapter 11.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Tile\u003C\u002Fstrong> — A rectangular piece of a larger array that one program works on.\nChapter 11.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>TMA\u003C\u002Fstrong> — Tensor Memory Accelerator. Hardware moving tiles between global and\nshared memory without occupying the arithmetic units. Imposes 16-byte alignment.\nChapters 11 and 21.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Thread\u003C\u002Fstrong> — The hardware’s unit of execution. A program is implemented as a\ngroup of them. Chapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Warp\u003C\u002Fstrong> — 32 threads executing in lockstep. Why block sizes are multiples of\n32. Chapter 2.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>\u003Ccode>teenyc\u003C\u002Fcode>\u003C\u002Fstrong> — The modified rustc that compiles captured kernel source into GPU\ncode. Chapter 3.\u003C\u002Fp>\n",[],false,{"title":14,"titleHtml":14,"route":15},"Common Compile Errors","\u002Fkernels\u002Freference\u002Fcompile-errors",{"title":17,"titleHtml":17,"route":18},"Appendix: Porting a Python Triton Kernel","\u002Fkernels\u002Freference\u002Fporting",1786271830195]