[{"data":1,"prerenderedAt":38},["ShallowReactive",2],{"chapter:kernels\u002Fpatterns\u002Fatomics.json":3},{"project":4,"route":5,"title":6,"titleHtml":6,"navTitle":6,"part":7,"sourcePath":8,"editUrl":9,"html":10,"toc":11,"hasMermaid":31,"prev":32,"next":35},"kernels","\u002Fkernels\u002Fpatterns\u002Fatomics","Atomics","Real Patterns","patterns\u002Fatomics.md","https:\u002F\u002Fgithub.com\u002Fteenygrad\u002Fteenygrad\u002Fedit\u002Fmain\u002Fbooks\u002Fkernels\u002Fsrc\u002Fpatterns\u002Fatomics.md","\u003Cp>Every kernel so far has had a simple guarantee: each program writes to memory\nthat no other program writes to. Programs never disagree, because they never\ntouch the same address.\u003C\u002Fp>\n\u003Cp>Some problems cannot be written that way. This chapter is about the tool for\nthose, and about why it is the last tool to reach for.\u003C\u002Fp>\n\u003Ch2 id=\"the-problem\">The problem\u003C\u002Fh2>\n\u003Cp>Suppose several programs each need to add something to the same counter.\u003C\u002Fp>\n\u003Cpre class=\"code-panel\" data-lang=\"text\">\u003Ccode>program A: read count (5) ... add 1 ... write 6\nprogram B: read count (5) ... add 1 ... write 6\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Both read 5. Both write 6. One increment vanished. This is a \u003Cstrong>race\u003C\u002Fstrong>, and on a\nGPU with thousands of programs in flight it is not a rare edge case — it is what\nhappens.\u003C\u002Fp>\n\u003Cp>The read, the modify and the write have to be one indivisible step. That is what\nan atomic is.\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#6FBF98\">T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">atomic_add\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">ptr\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">add_offsets\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">offsets\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> values\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2 id=\"where-they-are-genuinely-needed\">Where they are genuinely needed\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>Scatter.\u003C\u002Fstrong> When the \u003Cem>output\u003C\u002Fem> index is computed from data rather than from\n\u003Ccode>program_id\u003C\u002Fcode>, two programs can land on the same place and you cannot prove\notherwise. Negative log-likelihood loss is exactly this: the gradient goes to\nthe target class, and the target comes from a label tensor.\u003C\u002Fp>\n\u003Cp>The library’s kernel does it in two steps — a normal masked store for the bulk\nof the row, then an atomic for the one data-dependent position:\u003C\u002Fp>\n\u003Cpre data-lang=\"rust\" class=\"shiki teeny-datasheet\" style=\"background-color:#16181a;color:#e6e8e3\" tabindex=\"0\">\u003Ccode>\u003Cspan class=\"line\">\u003Cspan style=\"color:#7F877D;font-style:italic\">\u002F\u002F Step 2: subtract dy at target position via atomic_add(-dy)\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> base\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Tensor\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">i32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> =\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">full\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">i32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>(&#x26;[\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> row_base\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> flat_off\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">:\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">Tensor\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">&#x3C;\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">i32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">>\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> =\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> base \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">+\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> tgt\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#FF5F9E\">let\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> neg_dy \u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">=\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">full\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(&#x26;[\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">],\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> -\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">1\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#B79AD4\">0_\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\">f32\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">)\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\"> *\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> dy\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">;\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003Cspan style=\"color:#6FBF98\">T\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">::\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">atomic_add\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">dx_ptr\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">.\u003C\u002Fspan>\u003Cspan style=\"color:#7FB6D9\">add_offsets\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">(\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\">flat_off\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">),\u003C\u002Fspan>\u003Cspan style=\"color:#E6E8E3\"> neg_dy\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">,\u003C\u002Fspan>\u003Cspan style=\"color:#6FBF98\"> None\u003C\u002Fspan>\u003Cspan style=\"color:#8A9088\">);\u003C\u002Fspan>\u003C\u002Fspan>\n\u003Cspan class=\"line\">\u003C\u002Fspan>\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cstrong>Gradient accumulation in a backward pass.\u003C\u002Fstrong> If a value was read by many\noutputs in the forward pass, its gradient is the sum of many contributions.\nConvolution backward is the standard case, and several of this tree’s \u003Ccode>conv\u003C\u002Fcode>\nand \u003Ccode>pad\u003C\u002Fcode> backward kernels use atomics for exactly this.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Histograms and counters.\u003C\u002Fstrong> Bin from data, increment. The definition of the\nproblem is a race.\u003C\u002Fp>\n\u003Ch2 id=\"the-full-set\">The full set\u003C\u002Fh2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Method\u003C\u002Fth>\n\u003Cth>Operation\u003C\u002Fth>\n\u003Cth>Dtypes\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>atomic_add\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>*p += v\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>numeric\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>atomic_max\u003C\u002Fcode>, \u003Ccode>atomic_min\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>*p = max\u002Fmin(*p, v)\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>numeric\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>atomic_and\u003C\u002Fcode>, \u003Ccode>atomic_or\u003C\u002Fcode>, \u003Ccode>atomic_xor\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>bitwise\u003C\u002Ftd>\n\u003Ctd>integer\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>atomic_xchg\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>*p = v\u003C\u002Fcode>, returns old\u003C\u002Ftd>\n\u003Ctd>any\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>atomic_cas\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>compare and swap\u003C\u002Ftd>\n\u003Ctd>any\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>All return the \u003Cstrong>previous\u003C\u002Fstrong> value, which is what makes \u003Ccode>atomic_add\u003C\u002Fcode> usable as a\n“claim me a slot” primitive: the value you get back is your index.\u003C\u002Fp>\n\u003Cp>All except \u003Ccode>atomic_cas\u003C\u002Fcode> take a mask, so only the lanes you want participate.\u003C\u002Fp>\n\u003Ch2 id=\"ordering-and-scope\">Ordering and scope\u003C\u002Fh2>\n\u003Cp>The last two arguments control how strongly the operation is ordered against\neverything else.\u003C\u002Fp>\n\u003Cp>\u003Ccode>MemSem\u003C\u002Fcode> — memory semantics:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Value\u003C\u002Fth>\n\u003Cth>Meaning\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Relaxed\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Atomic, but no ordering guarantee about anything else\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Acquire\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Later operations cannot move before this one\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Release\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Earlier operations cannot move after this one\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>AcqRel\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Both. The default\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>\u003Ccode>MemScope\u003C\u002Fcode> — who has to agree:\u003C\u002Fp>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Value\u003C\u002Fth>\n\u003Cth>Meaning\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Cta\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>Only programs in this block\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Gpu\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>All programs on this device. The default\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>Sys\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>The whole system, including the host\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\n\u003C\u002Ftable>\n\u003Cp>Passing \u003Ccode>None\u003C\u002Fcode> for both gives \u003Ccode>AcqRel\u003C\u002Fcode> and \u003Ccode>Gpu\u003C\u002Fcode>, which is correct and is what\nevery kernel in this tree does. Weakening them is a real optimisation — a\n\u003Ccode>Relaxed\u003C\u002Fcode> counter is cheaper than an \u003Ccode>AcqRel\u003C\u002Fcode> one — but it is the kind of change\nto make when you are measuring and can explain why it is safe.\u003C\u002Fp>\n\u003Ch2 id=\"what-they-cost\">What they cost\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>Contention.\u003C\u002Fstrong> When many programs hit the same address, the hardware serialises\nthem. A histogram where 90% of values land in one bin runs at roughly the speed\nof one program. The fix is usually to reduce first and then use one atomic per\nprogram rather than one per lane.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Non-determinism.\u003C\u002Fstrong> Floating-point addition is not associative, so atomics\narriving in a different order give a different sum in the last bits. Run the same\nkernel twice on the same input and you may get answers that differ by an ulp.\u003C\u002Fp>\n\u003Cp>That is worth stating clearly because of what it does to your tests. A backward\npass using atomics is not bit-reproducible, so exact-equality assertions will\nfail intermittently. Compare with a tolerance, as the tests in this tree do.\u003C\u002Fp>\n\u003Ch2 id=\"the-alternatives-first\">The alternatives, first\u003C\u002Fh2>\n\u003Cp>Before an atomic, check these:\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Can you change who owns the output?\u003C\u002Fstrong> Often a race exists because the kernel\nis parallelised over inputs. Parallelise over \u003Cem>outputs\u003C\u002Fem> instead — one program\nowning each output element — and the race disappears. This is the single most\ncommon fix.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Can you use two kernels?\u003C\u002Fstrong> Partial results per program, then a second kernel\nto combine them. More memory and another launch, but deterministic and usually\nfaster under contention.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Can you reduce within the block first?\u003C\u002Fstrong> If all 128 lanes are adding to the\nsame place, \u003Ccode>T::sum\u003C\u002Fcode> them and do one atomic instead of 128.\u003C\u002Fp>\n\u003Cp>The reasonable default is: parallelise over outputs, and use atomics only when\nthe output index genuinely comes from the data.\u003C\u002Fp>\n\u003Cp>Next: making one kernel serve several dtypes.\u003C\u002Fp>\n",[12,16,19,22,25,28],{"id":13,"text":14,"level":15},"the-problem","The problem",2,{"id":17,"text":18,"level":15},"where-they-are-genuinely-needed","Where they are genuinely needed",{"id":20,"text":21,"level":15},"the-full-set","The full set",{"id":23,"text":24,"level":15},"ordering-and-scope","Ordering and scope",{"id":26,"text":27,"level":15},"what-they-cost","What they cost",{"id":29,"text":30,"level":15},"the-alternatives-first","The alternatives, first",false,{"title":33,"titleHtml":33,"route":34},"Reductions and Scans","\u002Fkernels\u002Fpatterns\u002Freductions",{"title":36,"titleHtml":36,"route":37},"Compile-Time Parameters and Dtype Dispatch","\u002Fkernels\u002Fpatterns\u002Fspecialisation",1786271829994]