Trying to achieve the best vectorization/locality/prefetch/.. of Matrix Multiplication on x86 Processors