Lukas Insam CTO, MUSE.holiday · Val Gardena / South Tyrol

07 Experiment → System · 2026

AI quality judges

Automated reviewers that grade what the AI said.

Every six hours, each automatically sent reply is scored good, risky or bad against a rubric. A second judge looks at drafts a colleague rewrote and decides whether the AI was wrong, incomplete or just differently worded. A shadow judge asks whether a reply would have been safe to send automatically. A collector replays the knowledge-base searches the agent really made and records how well retrieval matched.

Years
2026
Kind
Experiment → System
Status
Running
Cluster
2. Guests
Relations (as written)
grades (6)

Connected to

Built during: Chief Technology Officer