logoalt Hacker News

philipkglasstoday at 4:33 PM0 repliesview on HN

I need to process a modest amount of imagery (about 25 million images, and growing) for NSFW content and general captioning/description. About 5% of it contains nudity or partial nudity, and about 10% of that 5% contains sexual activity.

In theory, modern vision language models could classify human nudity and sexual activity very thoroughly. But every model I have tried is reluctant to clearly describe what is notable about sexualized/nude images. The models are deliberately under-exposed to nude and sexualized content during training and further RLHF'd away from generating straightforward descriptions of such images.

Models also occasionally hallucinate WTF captions for ordinary adult sexual activity. I recently ran a baseline test with frames extracted from adult videos and about 1/3000 frames was mis-captioned as involving a child according to Gemma 4 12b.