RAGnarok Human Evaluation

About this study →

Can we trust AI-based evaluation of RAG systems?

RAG systems (retrieval-augmented generation) answer questions using information retrieved from technical documentation. Increasingly, automated evaluators — including AI judges — are used to assess whether those answers are relevant, faithful to their sources, and complete. This open-source research study measures how well those automated evaluations agree with human judgment: yours.

Your role is simple: review a small set of anonymized RAG answers and judge them against the documentation excerpts provided. No coding required.

The studyA reproducible, openly published study on human vs. automated RAG evaluation.
Your roleReview 10–15 anonymized cases and answer four yes/no questions about each.
Your timeAbout 30–45 minutes total. Progress is saved after each case — you can stop and resume later.

No account or email is required. No personal information is collected. The full methodology is public — see About this study.

Which technology are you most comfortable evaluating?

How familiar are you with it? (optional)

By continuing, you agree that your anonymous annotations may be included in the publicly released research dataset and study results. No email, name, account, or other directly identifying information is required. Please read the annotator guide linked from where you received this study.