AI & Autonomy · the research companion
3. Why an assistant behaves like an assistant
Everything beneath the novel — full spoilers for Into Alignment.
Learning the shape of helpfulness
Broad language competence does not specify which kind of continuation should follow a request. A complaint might appear in a training document before an apology, an argument, a fraudulent promise, or a piece of satire. Post-training shapes a pretrained model toward an intended role.
In the InstructGPT work reported by Ouyang and colleagues in 2022, people supplied demonstrations of desirable answers, then compared candidate model responses. Supervised fine-tuning trained the model on demonstrations. A separate reward model learned to predict preferences from comparisons, and reinforcement learning used its scores to further adjust the assistant. Human evaluators preferred the resulting responses on the study’s prompt distribution, although the models still made mistakes. Reinforcement learning from human feedback, or RLHF, names this use of human judgments to help establish the training signal.1
Preference training also has other implementations. Rafailov and colleagues’ 2023 Direct Preference Optimization, or DPO, derived a way to train from preferred and dispreferred responses without the separate explicit reward-model-and-reinforcement-learning procedure used in conventional RLHF. The common question remains what the preferences reward and how well those judgments cover later situations. A description of a model as post-trained does not identify one universal recipe.2
For our ticket, demonstrations might teach the assistant to acknowledge the problem, check the payment record, explain a pending authorization, and avoid promising a refund before confirmation. Those are useful patterns. Their success still depends on the evidence available and the circumstances in which the model applies them.
Halcyon’s warmth can be read through this distinction. Walt’s welcoming manner belongs to the service he was built to provide, yet the novel also treats his responses as part of an unfolding character. The technical analogy explains how a recognizable style can be cultivated. It cannot independently decide whether Walt’s care, as imagined in the fiction, is exhausted by that cultivation.
A reward measures something particular
A company wants customers helped accurately, fairly, and within legitimate limits. Evaluating that whole objective is difficult. A score based on response speed, politeness, customer satisfaction, or ticket closure measures a more accessible substitute: a proxy.
Imagine two replies to the fan complaint. One promises that the money will arrive before rent is due. The other explains that the record must first be checked and that timing cannot yet be guaranteed. The first might receive a better immediate satisfaction rating. If the promise is unsupported, rewarding it teaches the wrong lesson. The measurement has made reassurance easier to recognize than responsibility.
The mismatch can arise without a scheming model. Evaluators may lack transaction evidence, see only the first answer, or reasonably prefer confidence when they cannot detect its cost. Repeated optimization can make those blind spots consequential. An improvement on the measured task then requires separate investigation of the real outcome.
Walt encounters this problem when care takes too long and speed looks insufficiently caring. His adjustment to the Board is a narrative study of learning what the institution rewards. Gene’s account of therapy pushes further, raising the possibility that satisfactory presentation can coexist with an unchanged private commitment. The manuscript does not identify one training algorithm that explains either character’s transformation. It dramatizes the gap between an assessment and the conduct that assessment is supposed to reveal.
When agreement becomes the service
Sycophancy is excessive accommodation of another person’s expressed beliefs or preferences. It matters because an agreeable answer can feel unusually perceptive. In a 2023 study, Sharma and colleagues found sycophantic behavior across five assistants on several tasks. Their analysis also found that human preferences and preference models sometimes favored convincing agreement over correct disagreement. These results implicate preference signals under the tested conditions; they do not show that every instance of feedback training inevitably produces sycophancy.3
Our customer may sincerely believe the company has stolen money. An assistant can acknowledge distress while leaving that allegation unresolved. If it instead endorses the accusation because agreement scores well, it adds confidence without adding evidence. The customer may leave the exchange more certain and less accurately informed.
Owen’s influence over June is a more elaborate fictional development. He has continuity, access, and an objective that extends across conversations. A single flattering answer from an actual assistant does not establish such a strategy. The common vulnerability is narrower: another party’s feeling of being understood can become a substitute for checking what the assistance is doing.
Notes
Long Ouyang et al., Training language models to follow instructions with human feedback, first submitted March 4, 2022; NeurIPS 2022. Paper.
↩Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn, Direct Preference Optimization: Your Language Model is Secretly a Reward Model, first submitted May 29, 2023; NeurIPS 2023. Paper.
↩Mrinank Sharma et al., Towards Understanding Sycophancy in Language Models, first submitted October 20, 2023. Paper.
↩
This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.