An image is a tensor you can multiply. Channel-last layout, colour as a 3x3 matrix, 2D rotations and affine maps, then the move that makes convolution a matrix multiply: im2col.
Learning Objectives
→Reshape a flat buffer into channel-last (H, W, C) layout
→Convert RGB to luminance as a linear combination of channels
→Apply a 3x3 colour matrix to every pixel
→Rotate 2D points with a 2x2 rotation matrix, and apply a 2x3 affine map
→Zero-pad an image, extract sliding patches with im2col, and compute a valid convolution as one matrix multiply