Skip to content
Home

DeepEval

Provides a framework for unit testing and benchmarking LLM applications with research-backed metrics, synthetic data generation, and tracing.

deepeval.comLLM ToolsMay 2026Confident AI Inc.

DeepEval is an open-source framework for evaluating and testing LLM applications, solving the problem of unreliable AI outputs by enabling teams to build evaluation pipelines. It targets developers and data teams at companies of all sizes, from startups to Fortune 500s. It is positioned as a pytest-native unit testing framework for LLMs that runs in CI/CD or as Python scripts, with research-backed metrics and transparent scoring.

Key features

  • Pytest-native evals for CI/CD
  • 50+ research-backed metrics
  • Hallucination, faithfulness, relevancy metrics
  • Native conversational evals
  • Multi-modal support (text, images, audio)
  • G-Eval criteria-based scoring
  • DAG metrics for multi-step scoring
  • QAG for reference-grounded scoring
  • Trace and grade agent steps
  • Synthetic golden data generation
  • Simulate conversations across personas
  • X0.1/day
GTM channels
  • Blog
  • Marketplace
  • Community
  • Docs
  • Changelog
ICP
  • Software developers
  • Engineering teams
  • Data analytics teams