Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
Citations
0
Open access
No
Source
arxiv
OpenAlex
Not enriched
arXiv
2411.10534
Abstract
We introduce a method to measure the alignment between public will and language model (LM) behavior that can be applied to fine-tuning, online oversight, and pre-release safety checks. Our `chain of alignment' (CoA) approach produces a rule based reward (RBR) by creating model behavior $\textit{rules}$ aligned to normative $\textit{objectives}$ aligned to $\textit{public will}$. This factoring enables a nonexpert public to directly specify their will through the normative objectives, while expert intelligence is used to figure out rules entailing model behavior that best achieves those objectives. We validate our approach by applying it across three different domains of LM prompts related to mental health. We demonstrate a public input process built on collective dialogues and bridging-based ranking that reliably produces normative objectives supported by at least $96\% \pm 2\%$ of the US public. We then show that rules developed by mental health experts to achieve those objectives enable a RBR that evaluates an LM response's alignment with the objectives similarly to human experts (Pearson's $r=0.841$, $AUC=0.964$). By measuring alignment with objectives that have near unanimous public support, these CoA RBRs provide an approximate measure of alignment between LM behavior and public will.
Collections
Add to collection
Paper intelligence
Analysis has not been completed yet.
No graph connections yet.
Sync citations or add papers to shared collections to build this network.
Knowledge graph
Citation network
Explore references, papers that cite this work and related papers in your Codex library.
References
0No references have been linked yet.
Cited by
0No saved paper is currently linked as citing this work.
Related papers
0Add papers to shared collections or enrich their topics to find related work.
Research workspace
Attach the paper PDF, extract its text, classify its contents and create semantic embeddings.
Paper resources
No PDF assets have been attached to this paper yet.