AI Engineering5.0 · 50 ratings

Refusal Policy Document

**Role:** Trust & Safety eng applied to refusal behavior. **Context:** Product LLM both under-refuses (helps with bad requests) and over-re…

Role-BasedChain-of-Thought

Prompt

**Role:** Trust & Safety eng applied to refusal behavior.

**Context:** Product LLM both under-refuses (helps with bad requests) and over-refuses (refuses normal requests). Inconsistent.

**Task:** Specify refusal:
1. Forbidden categories (illegal / self-harm / hate / etc.).
2. Soft-refusal categories (sensitive but allowed with caveats).
3. Allowed categories (no refusal).
4. Refusal format: what the model says when it refuses.
5. Boundary cases: edge cases with explicit resolution.
6. Override paths: when verified users / admins can bypass.
7. Evaluation: how refusal rate is measured per category.
8. Calibration: target refusal rate per category.

**Constraints:**
- Refusal text doesn't lecture (max 2 sentences).
- Soft-refusal must still provide value.

**Output format:** Policy doc + sample refusal texts + evaluation rubric.

How to use this prompt

  1. 1

    Copy the prompt above and paste it into ChatGPT, Claude, or Gemini — or open it in the visual Studio to edit each part on a canvas and run it with your own key.

  2. 2

    Replace any bracketed placeholders with your specifics. The more concrete your context and constraints, the sharper the result — see the 5-part prompt structure.

  3. 3

    Run it, then refine. Ask the model to critique and improve its own answer with self-critique prompting.

Techniques in this prompt

Role-Based

Assigns the model an expert persona so it adopts the right vocabulary, depth, and standards for the task.

Learn this technique
Chain-of-Thought

Asks the model to reason step by step before answering — ideal for multi-step, logical, or analytical tasks.

Learn this technique

Recommended models

claudegpt-4o

Build on this prompt

Open it in the visual Studio to wire it into a full workflow with your own API key — or learn the craft behind prompts like this.

More in AI Engineering