Every decision the app makes returns a set of outputs. You choose which ones in the Outputs section of your policy, and you can add your own custom outputs for anything else you want the model to tell you about each piece of content.
Built-in outputs
Two outputs are always on:
- Label: the most appropriate label for the content, from your policy's categories.
- Assessment: the safety assessment: clear vs. flagged.
The rest are optional; turn on the ones your workflow uses:

- Severity Score: how bad the violation is, on a 0–3 scale (0 = clear, 3 = most severe).
- Reason: a short reason for the decision (10 words or fewer).
- Analysis: a step-by-step walkthrough of the decision. Most useful while iterating on a policy; most teams turn it off in production.
- Risk Score: a 0–10 certainty scale (more below). If the risk score is over 5, the assessment should be flagged.
- Policy Quotes: the verbatim policy clauses the decision rests on, violations and exemptions each with a short reason. Powers the annotated policy pane in eval review.
- Image Description, Image Text, Audio Description: for media content: a short description of the image or audio, and any text present in the image.
How to think about the risk score
The risk score is the model's certainty that the content violates your policy, from 0 to 10:
- 0–1: very certain safe. Obviously safe; you'd expect every moderator to agree there's no violation.
- 2–3: certain safe. Most moderators would agree there's no violation.
- 4–6: toss-up. Subjective or borderline; you'd expect moderators to disagree depending on their perspective.
- 7–8: certain violation. A clear violation most moderators would agree on.
- 9–10: very certain violation. Obvious; you'd expect every moderator to agree.