Inspect AI
UK AI Security Institute framework for large language model evaluations and benchmarks.
About
Inspect is an open source framework for large language model evaluations created by the UK AI Security Institute. Evaluations are composed from datasets, solvers, and scorers in Python, with built-in components for prompt engineering, multi-turn dialog, tool usage, and model-graded scoring, so tasks range from simple question-answering benchmarks to agentic evaluations in which models operate tools. A companion collection provides more than 200 prebuilt evaluations that run against supported model providers without modification, and the framework is extensible through ordinary Python packages that add new elicitation and scoring techniques. Tooling includes a log viewer for examining transcripts of evaluation runs. Released under the MIT license with documentation at inspect.aisi.org.uk, Inspect is used by safety researchers, government evaluators, and engineering teams that need reproducible, auditable measurements of model capability and behavior.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Observability & Evaluation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- MIT
- Added
- May 7, 2026
Related Tools
ML experiment tracking, visualization, and collaboration
Open source LLM engineering platform for tracing and analytics
Open-source library for evaluating and tracking LLM applications.
Open-source AI metadata tracker for logging and comparing ML experiments.
Open-source AI observability platform for tracing, evaluation, and experimentation.
Python framework for unit testing and evaluating LLM applications with metrics like G-Eval.