The arithmetic forces automation
Moderation starts from a constraint, not a preference. The volume uploaded to a large platform is far beyond what any number of humans could individually review. There is no budget, and no labour pool, that makes human-first moderation possible.
So the pipeline is always the same shape:
- Classifiers score everything. Every item, automatically, in milliseconds.
- The confident majority is auto-resolved. Clearly fine: published. Clearly violating: removed.
- The uncertain middle escalates to humans. A tiny fraction of the total, which is the only reason human review is affordable at all.
Humans are not the first line. They are the appeals court for the ambiguous band, and the classifier decides who gets there.
This matters for a reason people miss: when a decision feels obviously wrong to you, it usually was not a bad judgement by a person. It was a threshold, applied automatically, to an item that never reached a human at all.

