Inspect AI

UK AI Security Institute framework for large language model evaluations and benchmarks.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

Inspect is an open source framework for large language model evaluations created by the UK AI Security Institute. Evaluations are composed from datasets, solvers, and scorers in Python, with built-in components for prompt engineering, multi-turn dialog, tool usage, and model-graded scoring, so tasks range from simple question-answering benchmarks to agentic evaluations in which models operate tools. A companion collection provides more than 200 prebuilt evaluations that run against supported model providers without modification, and the framework is extensible through ordinary Python packages that add new elicitation and scoring techniques. Tooling includes a log viewer for examining transcripts of evaluation runs. Released under the MIT license with documentation at inspect.aisi.org.uk, Inspect is used by safety researchers, government evaluators, and engineering teams that need reproducible, auditable measurements of model capability and behavior.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Easy (2/5)
License
MIT
Added
May 7, 2026

Related Tools

Featured

ML experiment tracking, visualization, and collaboration

Open Source
Easy
0.0 (0)
Featured

Open source LLM engineering platform for tracing and analytics

Open SourceSelf Hosted
Easy
0.0 (0)

Open-source library for evaluating and tracking LLM applications.

Open SourceSelf Hosted
Easy
0.0 (0)

Open-source AI metadata tracker for logging and comparing ML experiments.

Open SourceSelf HostedOffline
Easy
0.0 (0)

Open-source AI observability platform for tracing, evaluation, and experimentation.

Open SourceSelf Hosted
Easy
0.0 (0)

Python framework for unit testing and evaluating LLM applications with metrics like G-Eval.

Open SourceSelf HostedOffline
Easy
0.0 (0)
Browse all AI Observability & Evaluation tools