logoalt Hacker News

E-Reverance • last Tuesday at 10:33 PM • 1 reply • view on HN

I know it goes a tiny a bit against the spirit of what y'all are doing, but applying a few layers of pixel-wise local attention (so 1x1 "patch", with 3x3 or 5x5 attention window, basically treating it as a dynamic conv) has worked way better than both linear and conv unpatching in my recent experiments.

Diagram for reference https://x.com/1rreverant/status/2107546198093730287 (In my most recent recent experiment I actually removed the MLP and just used a linear project on the pixel's hidden states)


Replies

schopra909 • last Wednesday at 2:10 AM

That sounds like an interesting idea!

Can you confirm I'm understanding correctly?

1) Linear unpatchify as usual to go from hidden states to pixel space

2) Attention within a local window (e.g. 3x3, 5x5) to "blend" pixel space data and come up with a better image (as an alternative to MLP or Convolution)

And follow up questions:

1) How do you handle boundaries between your "attention windows"? Do you move the window just like a convolution does or are the "attention windows" all mutually exclusive from one another?

2) How much faster/slower is this operation vs. a linear layer + MLP?

➕ show 1 reply