Tool overview
Inspect AI is listed under AI Governance Safety & Evaluation AI tools.
What is Inspect AI?
Inspect AI is an MIT-licensed framework from the UK AI Security Institute and Meridian Labs for evaluating language models and agents. It supports datasets, solvers, scorers, tools, model providers, sandboxes, logs, and extensible evaluation packages.
Best for
Research and safety teams running reproducible model evaluations
Who is it for?
Decision note
Accepted for Preview after independent V411 official-source research. Apply only after failures = 0, warnings = 0, unmapped = 0, and the proposed changes are reviewed.
Key features
Task, solver, and scorer framework
Broad model-provider support
Agent tools and sandbox execution
Python API, CLI, logs, and viewer
Use cases
Evaluating model capabilities
Testing AI agents safely
Running benchmark suites
Building custom evaluation pipelines
Pros
- MIT-licensed official repository
- Extensible evaluation architecture
- Strong sandbox and logging support
Limitations
Inspect is an evaluation framework, not a hosted model. Teams must supply model access, datasets, infrastructure, permissions, and methodologically sound tasks and scorers.
Pricing details
Billing options
Free self-hosted
Pricing note
Inspect AI is free under the MIT license. Model API calls, cloud compute, sandboxes, storage, and benchmark execution create separate operational costs.
Supported languages
- English
Integrations
OpenAI
Anthropic
AWS Bedrock
Azure AI
Hugging Face
Please log in to join the discussion.