Good guesses. Crawler spoofing and pulling content that’s only hidden client-side are both in the mix, but you’re right that they aren’t enough for the bigger sites. For what it’s worth, NYT works reliably right now and WSJ is hit or miss, so WSJ is still the hard one.
I’m going to stay a little vague on the rest, because anything I describe in detail here is something a publisher can patch by Monday. The general idea is to treat each site as its own problem: several methods per request, a check on whether the article actually came back (not just a teaser), and on to the next method if it didn’t.
On archive.is: yes, the CAPTCHA is the annoying part. Volume is low enough that it’s manageable, and it’s a last resort rather than the main path.