Viva Republica (Toss) · Mar 2026 – Jun 2026
Ad Review Agent
+25 ppaverage accuracy gainad-copy typo policy pilot · mean gain of the refinement and held-out sets (36 ads)
01Background & Goals
- Ad policies are often a line or two, but an AI reviewer needs a guideline with criteria, boundaries and examples, and people had to write one for every policy.
- When the guideline was stricter or looser than real reviewers, it rejected good ads or missed violations, and finding where it went wrong was hard.
02Key Challenges
- C1 Turning a short policy into workable review criteria
- C2 Closing the gap between the guideline and reviewers with data
- C3 Avoiding overfitting to the refinement data
03Contributions
Automatic guideline generationC1
- Found the basis in the policy source text and turned a short policy into a guideline with decision criteria, violation/OK boundaries and policy references
Data-driven refinementC2C3
- Had AI compare the guideline's decisions with reviewers', find the patterns behind the disagreements and revise the guideline (e.g. marking compounds, amount units and ad-style phrasing common in ads as OK)
- Split refinement and held-out data to confirm the improvement holds on unseen ads
- People only checked the revisions; nobody edited the guideline by hand
04Tech Stack
- Framework / Platform
- LLM, Prompt engineering
- Methodology
- Data-driven prompt refinement, Holdout validation
05Results
- Raised average accuracy by 25 pp on an ad-copy typo policy pilot (36 ads)
- Accuracy also rose on held-out ads never used for refinement, so the gain holds on new ads
- This approach became the prompt-refinement cycle of the Inspection Automation Platform