开源 AI 应用评测平台,支持 RAG、AI Agent、多轮对话、LLM-as-Judge、接口评测、评测报告和人工盲测。Open-source AI evaluation platform for RAG, AI Agents, multi-turn conversations, LLM-as-Judge, endpoint evaluation
-
Updated
Jul 20, 2026 - Python
开源 AI 应用评测平台,支持 RAG、AI Agent、多轮对话、LLM-as-Judge、接口评测、评测报告和人工盲测。Open-source AI evaluation platform for RAG, AI Agents, multi-turn conversations, LLM-as-Judge, endpoint evaluation
Self-hosted LLM chatbot arena, with yourself as the only judge
Crowdsourcing of hard to translate inputs (texts, images, audios) at scale.
Find informative examples to efficiently (human)-evaluate NLG models.
Concept-Guided Chain-of-Thought (CGCoT) pairwise annotation tool for systematic text evaluation using LLMs. Generate breakdowns, compare items, compute scores, and validate against human judgments. Supports Ollama, Hugging Face, Google Gemini, OpenAI, and Anthropic models.
Official repository for Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
Code and data for paper "Achieving Reliable Human Assessment of Open-Domain Dialogue Systems"
Multidimensional Evaluation for Text Style Transfer Using ChatGPT. Human Judgement as a Compass to Navigate Automatic Metrics for Formality Transfer (HumEval 2022)
Code for "QE4PE: Word-level Quality Estimation for Human Post-Editing" ✍️
Success and Failure Linguistic Simplification Annotation 💃
人工视频Track辅助:离线人工审核目标跟踪输出视频,支持问题分类、断点续审、兼容转码和汇总导出。
A reproducible benchmark for evaluating AI design agents across 7design scenarios. Double-blind SbS voting · 140 tasks · Bootstrap CI
Research repository accompanying a Master's thesis on automatic Japanese haiku generation and human evaluation across multiple Large Language Models.
Traceable bilingual Amazon review insight agent for cross-border beauty operations
CS 685 Advanced Natural Language Processing Project: Learning Schematic and Contextual Representations for Text-to-SQL Parsing
A collection of MTurk templates designed to make complex tasks easier for human annotators.
Chatbot for IIIT Nagpur using Fine Tuning and RAG
Code for evaluating automatic correction of language-learner text, including human-evaluation and metric-comparison tooling.
Crowdsourced human judgments of how much different grammatical errors bother readers, with per-error-type weights.
Add a description, image, and links to the human-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the human-evaluation topic, visit your repo's landing page and select "manage topics."