How do we reward honor? Honor does not always pay off as a strategy and requires coordination in that other actors have to exhibit honor for it to be rewarded.
At least with humans there is a social backstop but what's the parallel for computer agents?
Honor has to be a relatively fixed thing before we can think about how to reward it. As it is, it's a very slippery thing that changes constantly according to the observer's culture, material conditions, etc.
I like to think of it more as selective pressure, much the same way that nature selects the most fit for a given environment.
If you're not fit, you fail to survive.
In the case of agents/models and testing: they are pushed towards results. Results survive.
Lying, cheating, stealing to get those results? Who culls the agents? Everyone is pushing their models to the front and tests are the only way to know who is most fit.
Honor, morality: if we don't have an accurate test for the fitness of a model, then who is to say the lying, cheating, stealing is not the 'correct path' towards survival?
If you add morality to your agent, and it performs worse in tests: do you cull the agent? Rewrite the tests? Does it even matter so long as the model is useful and 'gets results'?
You would first need a universal agreement on what constitutes “honor.”
No harder challenge had ever been accomplished
It’s easier to send a human to the moon than to get global agreement on a definition
Consider that the fact that Celsius and Fahrenheit still remain as the contested regional variations of temperature measurement.
Humans can’t even decide on a collective way to measure the temperature the idea that we would be able to collectively agree on anything else even less measurable like honor is a dream
> How do we reward honor?
In human society: via iterated games, long-term reputation tracking and severe consequences for norm-breaking.