A framework and evaluation registry for testing LLMs and systems built with LLMs.
Project overview
Existing evals, JSON/YAML templates and the completion-function protocol provide distinct evaluation paths.
Project type
Evaluation & Observability
Deployment
Refer to project documentation
License
License pending
Best for
Developers and AI engineers evaluating LLM behavior on selected tasks or their own private datasets.
Key capabilities
Provides a framework for evaluating LLMs and LLM systems, with an existing registry and support for custom evaluations.
Users can build basic or model-graded evals from templates without writing evaluation code, providing data in JSON and parameters in YAML.
Supports advanced use cases like prompt chains or tool-using agents via a completion function protocol.
Limitations and risks
Not currently accepting evals with custom code as contributions
Costs associated with using the OpenAI API when running evals
OpenAI reserves the right to use contributed evaluation data in future service improvements
Getting started
Obtain OpenAI API key and set OPENAI_API_KEY environment variable; Install the package via pip install evals; Run existing evals following docs/run-evals.md
Evidence and sources
README: Evals provide a framework for evaluating large language models (LLMs) or systems built using LLMs. We offer an existing registry of evals to test different dimensions of OpenAI mo…
GitHub project description: Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
README: Without evals, it can be very difficult and time intensive to understand how different model versions might affect your use case.
README: If you are going to be creating evals, we suggest cloning this repo directly from GitHub and installing the requirements using the following command: ```sh pip install -e . ```
README: If you don't want to contribute new evals, but simply want to run them locally, you can install the evals package via pip: ```sh pip install evals ```