The nice thing is if M >= B^2 (i.e., total memory is large enough to fit a square region of the image, where each row/column of the square fits a full block/page of memory), you can transform from row/column order to Hilbert/Z-order without needing to do more I/Os.
So you can't do such a conversion in a streaming fashion, but there is no need to load all data in memory either.
Yeah, those images explain it pretty well. There is one slight difference, though.
Because the Hilbert/Z-order curves are defined recursively, algorithms for converting between the curve and row/column order are "cache-oblivious", in that they don't need to take an explicit block size. You write it once, and it performs well on any system regardless the CPU cache size and/or page size.
So you can't do such a conversion in a streaming fashion, but there is no need to load all data in memory either.