{"repo":"confident-ai/deepeval","free":true,"listed":false,"github":"https://github.com/confident-ai/deepeval","clone":"git clone https://github.com/confident-ai/deepeval.git","description":"The LLM Evaluation Framework","language":"Python","stars":17680,"topics":["evaluation-metrics","evaluation-framework","llm-evaluation","llm-evaluation-framework","llm-evaluation-metrics","python"],"license":"Apache-2.0","category":"ai-agents","readme_excerpt":"The LLM Evaluation Framework Documentation Metrics and Features Getting Started Integrations Confident AI Deutsch Español français 日本語 한국어 Português Русский 中文 DeepEval is a simple-to-use, open-source LLM evaluation framework, for evaluating large-language model systems. It is similar to Pytest but specialized for unit testing LLM apps. DeepEval incorporates the latest research to run evals via metrics such as G-Eval, task completion, answer relevancy, hallucination, etc., which uses LLM-as-a-judge and other NLP models that run locally on your machine . Whether you're building AI agents, RAG pipelines, or chatbots, implemented via LangChain or OpenAI, DeepEval has you covered. With it, you can easily evaluate: - LLM apps end-to-end as black boxes - Complete agent trajectories across every decision and action - Individual agent steps such as LLM calls, tool use, retrieval, and sub-agent handoffs Use these evaluations to determine the optimal models, prompts, and architecture to improve your AI quality, prevent prompt drifting, or even transition from OpenAI to Claude with confidence. [!IMPORTANT] Want to compare iterations, share evaluation reports, and monitor your AI in production? Sign up for Confident AI, the enterprise AI evals and observability platform. Want to talk LLM evaluation, need help picking metrics, or just to say hi? Come join our discord. 🔥 Metrics and Features - 📐 Large variety of ready-to-use LLM eval metrics (all with explanations) powered by ANY LLM of ","default_branch":null,"files":null,"tree":[],"storefront":"/r/confident-ai","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/confident-ai/deepeval/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}