# The Eval Index > The living leaderboard of LLM and agent evaluation & benchmark tooling, ranked > daily by momentum (stars, push-recency, rising-newness) from live GitHub signals. Updated: 2026-07-21T10:52:00.152769+00:00 Tools indexed: 274 ## Top eval tools by momentum - [langfuse/langfuse](https://github.com/langfuse/langfuse) โ€” momentum 87, โญ31562 โ€” Observability โ€” ๐Ÿชข Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playgro - [mlflow/mlflow](https://github.com/mlflow/mlflow) โ€” momentum 86, โญ27134 โ€” Observability โ€” The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all - [promptfoo/promptfoo](https://github.com/promptfoo/promptfoo) โ€” momentum 85, โญ23456 โ€” Red Teaming & Safety โ€” Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare p - [comet-ml/opik](https://github.com/comet-ml/opik) โ€” momentum 85, โญ20743 โ€” Observability โ€” Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehe - [openobserve/openobserve](https://github.com/openobserve/openobserve) โ€” momentum 85, โญ20321 โ€” Observability โ€” Open source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM - [Tencent/WeKnora](https://github.com/Tencent/WeKnora) โ€” momentum 84, โญ18668 โ€” RAG Eval โ€” Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning - [confident-ai/deepeval](https://github.com/confident-ai/deepeval) โ€” momentum 84, โญ16996 โ€” Eval Frameworks โ€” The LLM Evaluation Framework - [dataelement/bisheng](https://github.com/dataelement/bisheng) โ€” momentum 82, โญ11725 โ€” Observability โ€” BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and - [Sahir619/fable-method](https://github.com/Sahir619/fable-method) โ€” momentum 82, โญ1765 โ€” Agent Eval โ€” The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eva - [Arize-ai/phoenix](https://github.com/Arize-ai/phoenix) โ€” momentum 81, โญ10650 โ€” Observability โ€” AI Observability & Evaluation - [oumi-ai/oumi](https://github.com/oumi-ai/oumi) โ€” momentum 80, โญ9363 โ€” Eval Frameworks โ€” Easily fine-tune, evaluate and deploy Gemma 4, Qwen3.5, Qwen3.6, gpt-oss, DeepSeek-R1, or any open s - [NVIDIA/garak](https://github.com/NVIDIA/garak) โ€” momentum 79, โญ8516 โ€” Eval Frameworks โ€” the LLM vulnerability scanner - [open-compass/opencompass](https://github.com/open-compass/opencompass) โ€” momentum 79, โญ7219 โ€” Benchmarks โ€” OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, Inter - [katanemo/plano](https://github.com/katanemo/plano) โ€” momentum 79, โญ6882 โ€” Observability โ€” Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability - [traceloop/openllmetry](https://github.com/traceloop/openllmetry) โ€” momentum 78, โญ7316 โ€” Observability โ€” Open-source observability for your GenAI or LLM application, based on OpenTelemetry - [jeinlee1991/chinese-llm-benchmark](https://github.com/jeinlee1991/chinese-llm-benchmark) โ€” momentum 78, โญ6301 โ€” Agent Eval โ€” ้ž็บฟๆ™บ่ƒฝ NoneLinear - ReLE่ฏ„ๆต‹๏ผšไธญๆ–‡AIๅคงๆจกๅž‹่ƒฝๅŠ›่ฏ„ๆต‹๏ผˆๆŒ็ปญๆ›ดๆ–ฐ๏ผ‰๏ผš็›ฎๅ‰ๅทฒๅ›Šๆ‹ฌ374ไธชๅคงๆจกๅž‹๏ผŒ่ฆ†็›–chatgptใ€gpt-5.4ใ€่ฐทๆญŒgemini-3.1-proใ€Claude-4. - [Giskard-AI/giskard-oss](https://github.com/Giskard-AI/giskard-oss) โ€” momentum 78, โญ5662 โ€” Red Teaming & Safety โ€” ๐Ÿข Open-Source Evaluation & Testing library for LLM Agents - [coze-dev/coze-loop](https://github.com/coze-dev/coze-loop) โ€” momentum 78, โญ5628 โ€” Observability โ€” Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent developmen - [GoogleCloudPlatform/agent-starter-pack](https://github.com/GoogleCloudPlatform/agent-starter-pack) โ€” momentum 77, โญ6520 โ€” Observability โ€” Ship AI Agents to Google Cloud in minutes, not months. Production-ready templates with built-in CI/C - [Kiln-AI/Kiln](https://github.com/Kiln-AI/Kiln) โ€” momentum 77, โญ4968 โ€” RAG Eval โ€” Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data g - [Marker-Inc-Korea/AutoRAG](https://github.com/Marker-Inc-Korea/AutoRAG) โ€” momentum 77, โญ4934 โ€” RAG Eval โ€” AutoRAG: Now your agent can find anything in your computer. It gets smarter if you are using it freq - [Andyyyy64/whichllm](https://github.com/Andyyyy64/whichllm) โ€” momentum 76, โญ5916 โ€” Benchmarks โ€” Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aw - [pydantic/logfire](https://github.com/pydantic/logfire) โ€” momentum 76, โญ4381 โ€” Observability โ€” AI observability platform for production LLM and agent systems. - [Agenta-AI/agenta](https://github.com/Agenta-AI/agenta) โ€” momentum 76, โญ4319 โ€” Observability โ€” The open-source workspace for building and running AI agents. Build agents through chat, share them - [open-compass/VLMEvalKit](https://github.com/open-compass/VLMEvalKit) โ€” momentum 76, โญ4294 โ€” Benchmarks โ€” Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchma - [Tencent/AI-Infra-Guard](https://github.com/Tencent/AI-Infra-Guard) โ€” momentum 76, โญ4163 โ€” Red Teaming & Safety โ€” A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, - [tensorzero/tensorzero](https://github.com/tensorzero/tensorzero) โ€” momentum 75, โญ11693 โ€” Observability โ€” TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, - [Helicone/helicone](https://github.com/Helicone/helicone) โ€” momentum 75, โญ5975 โ€” Observability โ€” ๐ŸงŠ Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC - [EvolvingLMMs-Lab/lmms-eval](https://github.com/EvolvingLMMs-Lab/lmms-eval) โ€” momentum 75, โญ4324 โ€” Benchmarks โ€” One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks - [truera/trulens](https://github.com/truera/trulens) โ€” momentum 75, โญ3450 โ€” Observability โ€” Evaluation and Tracking for LLM Experiments and AI Agents - [langwatch/langwatch](https://github.com/langwatch/langwatch) โ€” momentum 75, โญ3411 โ€” Observability โ€” The platform for LLM evaluations and AI agent testing - [harbor-framework/harbor](https://github.com/harbor-framework/harbor) โ€” momentum 75, โญ3342 โ€” Agent Eval โ€” Framework for evaluating and improving agents - [modelscope/evalscope](https://github.com/modelscope/evalscope) โ€” momentum 75, โญ3115 โ€” RAG Eval โ€” A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and p - [lmnr-ai/lmnr](https://github.com/lmnr-ai/lmnr) โ€” momentum 75, โญ3103 โ€” Observability โ€” Laminar - open-source observability platform purpose-built for AI agents. YC S24. - [FreedomIntelligence/Awesome-AI4Med](https://github.com/FreedomIntelligence/Awesome-AI4Med) โ€” momentum 74, โญ2841 โ€” Benchmarks โ€” A curated list of medical LLMs, multimodal systems, datasets, benchmarks, and more. ๐Ÿฅ - [openlit/openlit](https://github.com/openlit/openlit) โ€” momentum 74, โญ2628 โ€” Observability โ€” Open source platform for AI Engineering: OpenTelemetry-native LLM Observability, GPU Monitoring, Gua - [ifixai-ai/iFixAi](https://github.com/ifixai-ai/iFixAi) โ€” momentum 74, โญ1586 โ€” Red Teaming & Safety โ€” Catch your AI's mistakes and blind spots before your customers or regulators do. iFixAi runs 45 insp - [future-agi/future-agi](https://github.com/future-agi/future-agi) โ€” momentum 74, โญ1467 โ€” Observability โ€” Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applicati - [benchflow-ai/awesome-evals](https://github.com/benchflow-ai/awesome-evals) โ€” momentum 74, โญ743 โ€” Agent Eval โ€” A curated, non-BS library of the best resources for building and evaluating AI agents โ€” papers, blog - [AgentOps-AI/agentops](https://github.com/AgentOps-AI/agentops) โ€” momentum 73, โญ5719 โ€” Observability โ€” Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most