logoalt Hacker News

Lercyesterday at 10:26 PM1 replyview on HN

It might be beneficial while not being optimal on its own.

The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.

I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.


Replies

dnauticstoday at 2:32 AM

You can probably distribute training more easily too