aicoolies logo
Inspect AI parent UK AISI mark
Inspect AI parent UK AISI mark

Inspect AI

UK AI Security Institute framework for LLM safety evaluations

freeopen sourceverified Aug 24, 2026

Inspect AI is an MIT-licensed framework from the UK AI Security Institute for running large language model evaluations, including tool use, multi-turn dialogue, model-graded scoring, and reusable evaluation tasks.

Read our Inspect AI review

A detailed review by the aicoolies team — click to read

Inspect AI is an open-source evaluation framework created by the UK AI Security Institute. It structures evaluations around tasks, datasets, solvers, tools, model interfaces, and scorers so researchers and engineering teams can run repeatable tests instead of relying on ad hoc prompt checks.

The framework supports prompt-engineering workflows, tool-using agents, multi-turn dialogue, model-graded evaluation, and reusable evaluation tasks. These capabilities make it relevant to safety research, red-team programs, governance evidence, and CI gates that compare behavior across model, prompt, or agent changes.

Inspect AI does not turn an evaluation score into a production-safety guarantee. Teams still need to choose representative datasets, define appropriate scorers, review model-judge limitations, manage provider credentials and costs, and connect evaluation findings to broader monitoring, security, and release decisions.

Pricing

100% free and open-source (MIT License) developed by the UK AI Safety Institute (UK AISI). $0 software cost with unlimited self-hosted evaluation tasks, local web viewer (`inspect view`), and built-in sandboxing (Docker, Podman, Kubernetes). Users only pay underlying model API token costs directly to LLM providers or local GPU compute infrastructure.

full pricing breakdown →

Platforms

Python framework and CLI for defining and running LLM evaluation tasks across hosted and local model providers.

Categories

Tags

Use Cases

Related Tools

computed discovery: shared active categories · kept separate from editor-verified Alternatives

OpenLIT logo

OpenLIT

OpenTelemetry-native observability for LLM applications with evals and GPU monitoring

OpenLIT is an open-source AI engineering platform that provides OpenTelemetry-native observability for LLM applications. It combines distributed tracing, evaluation, prompt management, a secrets vault, and GPU telemetry in a single self-hostable stack. With 50+ integrations across LLM providers and frameworks, it lets teams monitor AI applications using their existing observability backends like Grafana, Datadog, or Jaeger.

Open Source
Cilium logo

Cilium

eBPF-based networking, security, and observability for Kubernetes

Cilium is a CNCF Graduated, Apache-2.0 project for Kubernetes networking, security, and observability using eBPF. It can replace kube-proxy, enforce identity-aware L3-L7 network policies, and add Hubble flow observability plus Tetragon runtime-security signals. Current source checks support GKE Dataplane V2 using Cilium/eBPF and Azure CNI Powered by Cilium for AKS.

Open Source
Playwright logo

Playwright

Reliable end-to-end testing

Cross-browser E2E testing framework by Microsoft supporting Chromium, Firefox, and WebKit with one API. Features auto-waiting, tracing with timeline/screenshots/DOM snapshots, codegen for recording tests, and parallel execution. Component testing for React, Vue, Svelte. Built-in API testing, network mocking, and mobile emulation. Known for reliability and speed vs Selenium/Cypress. 70K+ GitHub stars, rapidly becoming the E2E standard.

Open Source
Pangolin logo

Pangolin

Identity-aware VPN and reverse proxy for zero-trust remote access

Identity-based remote access platform built on WireGuard that combines reverse proxy and VPN capabilities. Pangolin supports clientless browser access for web apps and client-based private-resource access across macOS, iOS, Windows, Linux, and Android, with zero-trust rules, peer-to-peer tunnels, automatic SSL, SSO/OIDC options, and cloud or self-hosted deployment.

freemium
DeepEval logo

DeepEval

Apache-2.0 Python framework for repeatable LLM, RAG, agent, MCP, and safety evaluation workflows.

DeepEval is an Apache-2.0 Python framework for evaluating LLM apps, RAG systems, agents, MCP workflows, and safety behavior with repeatable test cases. It works locally and in CI/CD, then connects to Confident AI for hosted reports, observability, red teaming, and governance when teams need shared evidence instead of ad-hoc prompt reviews and manual QA.

freemiumOpen Source
reviewdog logo

reviewdog

Automated code review for any linter on CI

reviewdog is an open-source automated code review tool that integrates any linter or static analysis tool with GitHub, GitLab, Bitbucket, and Gitea pull requests. Parses output in errorformat, Checkstyle XML, SARIF, and JSON formats to post inline review comments on changed lines only. Works with GitHub Actions, Travis CI, CircleCI, GitLab CI, and Jenkins. Supports 40+ languages through universal linter adapter architecture.

Open Source

Used in Stacks

Comparisons

Promptfoo vs Inspect AI: Product CI or Frontier-Model Evaluation?

Promptfoo and Inspect AI are both open-source evaluation frameworks, but their operating models differ sharply. Promptfoo is designed for application teams that want config-driven prompt, model, agent, and security tests in everyday CI. Inspect AI, developed by the UK AI Security Institute and Meridian Labs, is designed for rigorous model evaluations built from datasets, solvers, scorers, tools, agents, and sandboxes. **Promptfoo is the better default for most product engineering teams** because it reaches a release gate faster and combines regression testing with red teaming. Inspect AI is the specialist choice for benchmark authors, safety researchers, and teams evaluating frontier-model capabilities or autonomous behavior.

PromptfooInspect AI

FAQ

What is Inspect AI?

Inspect AI is an MIT-licensed framework from the UK AI Security Institute for running large language model evaluations, including tool use, multi-turn dialogue, model-graded scoring, and reusable evaluation tasks.

Is Inspect AI free?

Yes — Inspect AI is free to use. 100% free and open-source (MIT License) developed by the UK AI Safety Institute (UK AISI). $0 software cost with unlimited self-hosted evaluation tasks, local web viewer (`inspect view`), and built-in sandboxing (Docker, Podman, Kubernetes). Users only pay underlying model API token costs directly to LLM providers or local GPU compute infrastructure.

Is Inspect AI open source?

Yes — Inspect AI is open source.

Is Inspect AI still maintained?

Yes — Inspect AI is active. Its listing was last verified on August 24, 2026.

How does Inspect AI score in our review?

Our hands-on review scores Inspect AI 85/100 overall, based on speed, privacy, and developer-experience testing.