Skip to content

Report triage ​

Triage of a user report about a post: your platform sends the classifier score and any known-content match with the report, events count the author's strikes and reports, and uncertain reports go to a moderator. Pack: examples/policies/report-triage/.

A pattern for expressing a trust & safety policy, not compliance guidance: the thresholds, lists and fields are placeholders chosen to make the example run; a real policy comes from the business's own rules. It calls no app, so it runs in an empty workspace as it is.

The decision it automates ​

Whether to leave the content up (approve), send it to a moderator, or remove it (reject), when your platform starts the run.

Data expected ​

  • The author as a subject (kind user) whose lastReport (classifierScore, knownContentMatch) your platform writes before it starts the run.
  • report.created and strike.issued events on the author, posted through POST /events.

Steps and rules ​

StepTypeWhy
rulesevaluate_rulesThe report-triage rule set under max_severity; nothing is collected.
routebranchrules.outcome == manual_review opens a case.
reviewcreate_casemoderation, priority high, 1-hour SLA.
decideemit_decisionFrom the rules, or the reviewer's decision.

The rule set report-triage (useCase: monitoring) uses max_severity: a failed block rule rejects, a failed warn rule sends the run to review, and the weights add up to the score.

RuleTypeSeverity, weightFails when
report_known_contentcomparisonblock, 100lastReport.knownContentMatch is not false
report_classifier_scorescore thresholdwarn, 40lastReport.classifierScore is above 0.8
report_repeat_offendervelocitywarn, 30more than two strike.issued events in 90 days
report_burstvelocitywarn, 20more than ten report.created events in 24 hours

What a reviewer sees ​

A moderation case with the rule results and the author's timeline of reports and strikes. The moderator's reject removes the content; decision.created tells your platform.

How to adapt it ​

Call your moderation classifier and image matching from the workflow as apps you build instead of sending their answers on the subject, add an appeal workflow reviewed by a different moderator, and export the audit trail for transparency reporting.

Run it ​

bash
pnpm policy:import examples/policies/report-triage --publish

The fixture reports a post with a classifier score of 0.86 three times in the last 15 minutes, by an author with one strike: the score rule fails, a moderation case opens, the check removes the content (reject) and the run completes with a manual reject.

Released under the Apache-2.0 License.