I know it goes a tiny a bit against the spirit of what y'all are doing, but applying a few layers of pixel-wise local attention (so 1x1 "patch", with 3x3 or 5x5 attention window, basically treating it as a dynamic conv) has worked way better than both linear and conv unpatching in my recent experiments.
Diagram for reference https://x.com/1rreverant/status/2107546198093730287 (In my most recent recent experiment I actually removed the MLP and just used a linear project on the pixel's hidden states)
That sounds like an interesting idea!
Can you confirm I'm understanding correctly?
1) Linear unpatchify as usual to go from hidden states to pixel space
2) Attention within a local window (e.g. 3x3, 5x5) to "blend" pixel space data and come up with a better image (as an alternative to MLP or Convolution)
And follow up questions:
1) How do you handle boundaries between your "attention windows"? Do you move the window just like a convolution does or are the "attention windows" all mutually exclusive from one another?
2) How much faster/slower is this operation vs. a linear layer + MLP?