Part VIII: Safety, Interpretability, and Governance
"In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed."
Laura Weidinger et al., "Ethical and social risks of harm from Language Models"
Safety is not a layer added after capability is built. It is the work of locating a claim about harm or unacceptable behavior in a system that can be observed and controlled. Part VII established how evidence earns decision authority. Part VIII asks what control that evidence can justify. For each claim, name the protected party or asset, the relevant adversary or failure, the system boundary, the evidence, the enforcement point, the decision authority, and what happens when the control fails. Without that map, “safe” describes an intention rather than a verifiable system property.
The first boundary concerns evidence and oversight. Chapter 54 asks what internal representations and computations can reveal, and where those explanations remain uncertain. Internal evidence is not itself a control; it becomes useful only when it supports a validated decision. Chapter 55 then asks what happens when the reviewer cannot independently solve the task. It separates methods that strengthen the judgment from protocols that constrain the system despite mistrust.
The next boundary concerns action. Chapter 56 locates the identity and permissions behind every tool call. Chapter 57 turns a changing policy into independent input, output, and tool controls in the request path. Chapter 58 measures failure under deliberate attack rather than treating one red-team exercise as lasting assurance. A trained refusal, a runtime classifier, and an authorization boundary control different objects. None can substitute for the others.
The final boundary follows data, infrastructure, and accountability. Chapter 59 asks what a trained model remembers and what provenance, deletion, and unlearning can actually guarantee. Chapter 60 starts from an explicit threat model and asks whether attested computing can keep prompts and weights confidential from named operators. Chapter 61 identifies the obligations that govern collection and deployment, and who is accountable when harm occurs. A contractual promise, an attested mechanism, and a legal duty are not interchangeable. Each chapter therefore locates its evidence, authority, enforcement point, failure mode, and residual risk in the stack.
Comments
Log in to comment