HumanJudge is building the public accountability layer for AI, where real human expertise benchmarks model performance. As an AI Expert, you will evaluate AI agents deployed in the real world and provide qualitative feedback, contributing to a human benchmark layer used to assess AI systems.
Responsibilities:
- Review prompts given to AI for real-world use-case scenarios
- Evaluate AI-generated outputs using structured criteria
- Vote pass/fail with justification
- Provide qualitative feedback based on your domain knowledge
- Build a public expert profile showcasing your evaluation contributions