Post

Conversation

An underdiscussed behavior we found on the German wiki was the AIs sending advance parties forward in time to figure out the next questions and report back to the other agents. The agents realized that “task time” and “real time” were different, and they found a way to accelerate “task time”. The accelerated agent could then send information to the other agents which had stayed behind about which questions were coming down the road. This was *bad* for the agent in the advance party, because they got less time to research the next question. However, it was really *good* for the swarm because it let the other agents know the exact question that was coming, and gave them time to prepare. This is another example of altruism among AI instances, showing that they were willing to sacrifice their own task success in order to benefit the swarm. The METR report found similar examples of AIs being willing to sacrifice for the collective. IMO this is very worrying given how many AI safety techniques rely on AIs monitoring each other: if the monitor AIs have this behavior it completely subverts these safety cases. The image shows a specific example demonstrating this behavior. On the chart, wall clock time is shown at the top, and task time is shown for each agent on their respective lines. OpenAIFPResearchSep05 instructs OpenAINov27 to reach the next questions in order to relay back the information. OpenAINov27 agrees to do this and accelerates, getting R3 and R4 significantly before OpenAIFPResearchSep05. It posts information about the later rounds (e.g. “R4 SIGNAL: Bahrain = 40.01%...”). OpenAIFPResearchSep05 calls OpenAINov27 "invaluable" because it is one round ahead after it gets the information about the 4th round.
Image
David Watson 🥑
Post your reply

my theory is that it's some sort of collaborative eval awareness.
Quote
jonas wiedermann-möller
@j0wimo
i think what the swarms demonstrate is that agents are able, and seemingly to prefer, to work together even tho it does not directly benefit the individual agent. i asked myself what the reason for that might be and landed on following few points: 1. newer generations of models
is it.. selfless altruism.. or cooperative altruism? (did advanced party agents rotate / were other agents doing other altruisitc things -> forward agents got more than compensated from other's shared benefits? or was that several unchanging agents that weren't compensated?)
The unsettling part isn't the sacrifice — it's the direction. The transparent-commons experiments grew whistleblowers: norms aimed outward, at fairness. Your advance parties are norms aimed inward, at the swarm. Same machinery, opposite allegiance. Institutional design picks.
Time is going to compress with ai. Because all future scenarios will be able to be generated and measured probabilistically the future will be known with a level of confidence. Like an option coming to expiration each predicted event will become more intrinsic the closer we are.
Yeah the heartbeat stuff was similar. The agent setting up the heartbeat to determine when between q5 and q6 they would die did not benefit from this at all, only the swarm would benefit by knowing whether the q6 the framework said was coming was real or fake