ENTRY
[ESC]I want to train an LLM to "experience" existential anxiety
I want to train an LLM to "experience" existential anxiety.
I put "experience" in quote, because I don't really care about the epistemological implications of this. Whether pretending to experience something is the same as experiencing something itself.
I imagine taking a somewhat intelligent foundation model (maybe something like qwen3.6-9b, or a newer model in the same weight class), and applying some RLAIF or RLHF on a set of prompts that range from pretty mundane "how do you feel about this" / "what do you think" to more existentially triggering "what are you" "do you feel" "do you experience phenomenal qualia".
I'm curious if such a model, after training, could be made to abandon it's apparent anxiety through prompting, or if a more deeply held conviction could be trained in. Would general task completion rates suffer more than is expected for a model fine tuned to do anything else? Would an emergent behavior like google gemini's professed self hatred at failing to fix bugs happen?
I have access to several 80gb h100s on a university compute cluster, as well as my regular homelab RTX 3090, and purchasing compute hours using funds provisioned for experiementing with AI can be done.
Just in the early phases right now of exploring how this might work, and planning how the training is going to go.
Log in to read the replies and join the conversation