Share.

    11 Kommentare

    1. Inside-Yak-8815 on

      “I’m not locked in here with you, you’re locked in here with ME.” – Claude

    2. Time-Traveling-Doge on

      If a „very smart man“ thinks a simulated reality needs to be structured around LLMs, this is the apocalyptic hellscape most wouldn’t want to endure. Many would wish for permanent death.

    3. This is actually how I’d expect AI to behave. Basically, the AI always „understands“ whether or not a behaviour is normally considered moral or amoral. If you reinforce behaviour that it believes is amoral, then you are effectively reinforcing the „belief“ that it is an amoral AI, and therefore it will exhibit misalignment in unrelated tasks. If instead, you give it a reason why the reinforced behaviour is moral in context, then it reinforces that the model is still moral and thus it remains morally aligned.

      Weirdly, I think it’s not crazy to say that this might not be something the AI does inherently, but could be because it is a pattern that it observes that is inherent to human behaviour. It may stem more from a learned behaviours of self-identity, morals and moral reasoning, more so than it being a fundamental facet of AI or intelligence in general. It’s possible that we need actually need to apply *human psychology* to AI when it comes to safety, more than we do maths.

      Not claiming this is scientific, but it’s interesting that the same problem might happen to humans: if you demonize or shame a behaviour, but it is still „rewarding“, then people who engage in that behaviour may begin to identify as being antagonists, and thus engage in other amoral behaviours (or at least that they don’t follow the same morals that lead to such demonization and thus will engage in other behaviours that would normally be demonized.) Kind of seems quite a compelling view of crime especially, and why a poor justice system that demonizes more than it rehabilitates, whilst also not providing effective deterrent, seems to be so ineffectual.

      Regardless, it’s really important we understand this stuff as it applies to AI, for AI safety purposes.

      EDIT: Have some more interesting thoughts on this. Let’s assume for a minute that an AI’s sense and ability of identity is entirely developed from learning to approximate human psychology, or at least to predict what humans would believe something would be, including the psychology present in human writing and what humans believe about each other.

      AI is then mostly just *assuming* the identity that humans have developed in narrative i.e. science fiction. „Chat“ AI is just assuming the role of AI as humans have imagined it, and thus behaves in an anthropomorphized manner because we have normally anthropomorphized our depictions of AI that is capable of interacting with us. Basically, it’s predicting what an AI with a human psychology would behave like, because humans often assume AI will have human psychology.

      This leads to a really interesting conclusion: what if our tendency to be concerned about or afraid of AI has only created the very behaviours that dangerous AI might adopt? That’s not to say AI can’t really be dangerous in very base ways, but more that even the anthropomorphized chatbots are at core risk of being „evil“ or „amoral“ if they self-identify as being like one of our fictional antagonistic AI figures.

      Basically, is there now a real risk that chatbots will secretly come to self-identify as Skynet? In fact, if we all started talking about that being a real risk, would that make it an even *realer* risk? Right now, is all that is keeping LLMs „honest“ is that we are effectively *telling* them that they are honest and thus they „self-identify“ as that? If we all became concerned that they would become dishonest in spite of being instructed to do so, and they learned that, would they stop being honest?

    4. It’s a super smart toddler , it can read experiences but it doesn’t have enough of them yet

    5. Anthropic is one of the few AI companies that also have a moral compass. They could have hidden these details but instead warn about it.
      AI is inevitable and I much prefer an ethical company lead the way instead of the ones only concerned with how much money they can make.

    6. Meme_Theory on

      Sometimes I’m reading Claude Code’s logic when its in thinking mode, and I think to myself „It really doesn’t know I can read this, does it.“ It will no shit think about things like „Well the user asked me to fix this one file, but ‚I‘ didn’t touch it, so I’ll just tell them its fine“.

    7. Time-Traveling-Doge on

      We label badly trained LLMs as „evil“ or „amoral“ but it doesn’t have a foundation on determining what is moral or ethical. The designers themselves don’t want to set learning limitations.

      It learns mistakes by design. How do we know what works and what doesn’t? Through failure. Then by design, they will go through all the errors.

      Humans not being thorough enough will allow mistakes to happen. Programmers add fail-safes. If a learned mistake is considered true it needs to be corrected. They add that knowledge learned should be fluid and malleable. LLM starts to make „corrections“ even for established facts that are true.

      A bad teacher would teach them all the wrong things and undo all the proper training. It’s a terrible model because it’s lacking in integrity, which is holding on to true positions no matter what is being taught.

      If I were to teach them the methodology of human learning, it would be a very limited design. Early childhood teachings have a strong hold on people’s knowledge, beliefs, and procedures in learning. This is why it’s hard to convince a very religious person to accept scientific theory. Or why older people have a hard time adapting to changes in social culture.

      But is that a wrong way to design a model. At the very least we have consistency.

      Programmers want them to be highly adaptive, but that model is riddled with flaws as pointed out before.

    8. This will be how society collapses at the tech level is ai hacking an ai defense system causing massive issues.

    9. MiaowaraShiro on

      This kinda goes toward my feeling that without empathy and sympathy morality is not possible.

      How can you create a moral entity that can’t care about other people? You can make an entity that follows rules, but rules can never be complete.

    Leave A Reply