07 Experiment → System · 2026
AI quality judges
Automated reviewers that grade what the AI said.
Every six hours, each automatically sent reply is scored good, risky or bad against a rubric. A second judge looks at drafts a colleague rewrote and decides whether the AI was wrong, incomplete or just differently worded. A shadow judge asks whether a reply would have been safe to send automatically. A collector replays the knowledge-base searches the agent really made and records how well retrieval matched.
S07 · A batch of graded replies: good / risky / bad · [supply]
- Years
- 2026
- Kind
- Experiment → System
- Status
- Running
- Cluster
- 2. Guests
- Relations (as written)
- grades (6)
Connected to
- 6. AI guest service graded by this
Built during: Chief Technology Officer