logoalt Hacker News

Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)

13 pointsby popopandatoday at 12:30 PM0 commentsview on HN

Comments

popopandatoday at 12:31 PM

[flagged]