Multimodal LLM-as-a-Judge
Provides simulation infrastructure and evaluation platform for training and testing AI agents and LLMs, addressing reliability and hallucination detection
Provides simulation infrastructure and evaluation platform for training and testing AI agents and large language models, addressing reliability, hallucination detection, and long-horizon task performance. Sells to developers, data and engineering teams building AI applications in enterprises and research organizations. Delivered as an infrastructure API and platform with evaluators, experiments, datasets, and RL environments for simulation-based training and testing.
Key features
- Digital World Models simulation
- RL environments for agents
- Patronus Evaluators scoring
- Patronus Experiments optimization
- Patronus Datasets adversarial testing
- Lynx hallucination detection model
- Glider rubric-based judge model
- FinanceBench financial Q&A benchmark
- SimpleSafetyTests safety diagnostic suite
- EnterprisePII sensitive data detection
- Deep research reasoning evaluation
- Multi-turn dialogue evaluation
- Long horizon task planning
- Agentic memory evaluation
- X0.4/day
- LinkedIn0.2/day
GTM channels
- Blog
- Newsletter
- Partner program
- Marketplace
- API
- Docs
ICP
- Software developers
- Data analytics teams
- Engineering teams