I had an idea for a new concept in alignment that might allow nuanced and human like goals (if it can be fully developed).
Has anyone explored using neural clusters found by mechanistic interpretability as part of a goal system?
So that you would look for clusters for certain things e.g. happiness or autonomy and have that neural clusters in the goal system. If the system learned over time it could refine that concept.
This was inspired by how human goals seem to have concepts that change over time in them.
I'm currently drafting a post on current sycophantic AI as something that threatens core human skills of reasoning, maths etc. Based on aviation skills fade and recent possible impact on CS outcomes. This could have knock on impacts to other causes as degraded skills might lead to degraded outcomes in fields like AI alignment or ethical reasoning in time.
I would appreciate some human collaboration on it. I don't have huge amount of time (so it is AI written currently, ironically, but writing ITN notes isn't my cue competency anyway).