Developers train refusals with examples, classifiers, and internal probes, but these safeguards remain statistical rather than certain.
Modern AI systems learn to refuse by studying examples of requests they should reject. Technology Review reports that companies reward models for refusing harmful prompts and penalize them for rejecting harmless ones. They may also use other AI systems to run these exercises.
Developers add smaller models called classifiers around the main system. One classifier can stop a dangerous prompt before it reaches the chatbot; another can block a harmful answer before it reaches the user. These layers catch different problems, but none catches every bad request.
Some newer safeguards, called probes, watch the model’s internal activity instead. Researchers look for patterns of “activations”—changes inside the model when it prepares to refuse. Technology Review says researchers still treat their understanding of those patterns as a hypothesis, not a complete explanation.
That uncertainty matters because refusal is probabilistic. Researchers have used jailbreaks—prompts designed to bypass safeguards—to draw restricted answers from widely used models. Technology Review reports that some researchers even succeeded by putting requests into poetic verse.
The opposite failure is over-refusal. A broad safety margin can block legitimate work along with dangerous requests. Technology Review describes a chatbot that redirected a harmless question about the difference between sake and makgeolli, apparently after safeguards were tightened around biological and cyber risks.
These protections also carry a practical cost. Anthropic said one type of classifier added 24% to chatbot compute costs, according to Technology Review.
For teams building or buying AI systems, the useful test is not whether a model refuses. Test whether it resists harmful requests while still handling legitimate security, medical, and research questions. Refusal should be treated as one safety layer—not a guarantee.
