Loading...

Gradient accumulation fakes a big batch on small memory — .grad already sums, so scale each micro-batch by 1/N and step once | AIWedia