It can learn. When my agents makes mistake they update their memories and will avoid making the same...

charcircuit • today at 5:20 AM • 3 replies • view on HN

It can learn. When my agents makes mistake they update their memories and will avoid making the same mistakes in the future.

>Reinforcement learning, on the other hand, can do that, on a human timescale. But you can't make money quickly from it.

Tools like Claude Code and Codex have used RL to train the model how to use the harness and make a ton of money.

Replies

kelnos • today at 7:16 AM

That's not learning, though. That's just taking new information and stacking it on top of the trained model. And that new information consumes space in the context window. So sure, it can "learn" a limited number of things, but once you wipe context, that new information is gone. You can keep loading that "memory" back in, but before too long you'll have too little context left to do anything useful.

That kind of capability is not going to lead to AGI, not even close.

➕ show 2 replies

Dansvidania • today at 7:13 AM

That’s not learning. That’s carrying over context that you are trusting is correctly summarised over from one conversation to the next.

➕ show 1 reply

otabdeveloper4 • today at 7:09 AM

> they update their memories

Their contexts, not their memories. An LLM context is like 100k tokens. That's a fruit fly, not AGI.

➕ show 1 reply

alt Hacker News

Replies