English

Model ReleasesMultiverse Computing

Proposed Boundary-Aware Self-distillation to Prevent Over-refusal in LLM Safety Alignment

This article is a translation. Read the Japanese original

A research team from Multiverse Computing has announced research findings via the Hugging Face Blog aimed at solving the problem of "over-refusal," where large language models (LLMs) reject entire specific topics during safety alignment.

In traditional safety alignment, it has been common to establish guardrails based on topic units, such as weapons or fraud. However, this method presents a challenge in implementing fine-grained control based on usage—for example, allowing the provision of factual information regarding politics while rejecting only political manipulation or incitement. Consequently, a phenomenon has occurred where models misidentify safe questions as "potentially harmful" and refuse to answer.

The method proposed by the team is an approach called "Boundary-Aware Self-distillation," which utilizes pairs of prompts with different "intents" within the same topic to train the model with an awareness of those boundaries. In experiments using Qwen3-8B, the team reported success in significantly improving the refusal rate for harmful prompts while reducing the refusal rate for safe prompts. The research team emphasized that model safety should be evaluated from both perspectives: not only the height of the refusal rate but also whether harmless responses are being maintained.

Sources

  1. Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic (Hugging Face Blog, 2026-09-08)
  2. ThinkSafe Paper