aicoolies logo

Testing & QA

Test runners, visual regression, E2E testing, and QA automation

108 tools

last updated August 16, 2026

showing 48 of 108 tools

Argos CI logo

Argos CI

Visual regression testing for CI/CD pipelines

Argos CI is a visual regression testing platform that automatically catches unintended UI changes in CI/CD pipelines. It integrates with Playwright, Cypress, Storybook, and Puppeteer, featuring a stabilization engine that filters flaky pixel differences from genuine regressions. Used by teams at Meta and MUI for frontend quality gates.

Open Source
BrowserMCP logo

BrowserMCP

Automate local Chrome browser via MCP

BrowserMCP is an MCP server that enables AI agents to automate a local Chrome browser — navigating pages, clicking elements, filling forms, extracting content, and taking screenshots. It gives coding agents the ability to interact with web applications the way a human would, directly from Claude Desktop, Cursor, or any MCP client.

Open Source
Bugster logo

Bugster

Autonomous E2E test agent that explores your app

Bugster is an autonomous testing agent that explores applications independently to write and run end-to-end tests without requiring manual script creation. It identifies broken flows by navigating the UI like a real user, automatically discovering clickable elements, form inputs, and navigation paths, then generating reproducible test scripts for the issues it finds.

paidOpen Source
Currents logo

Currents

Parallel test orchestration for Playwright and Cypress

Currents is a test orchestration platform that parallelizes Playwright and Cypress test suites across multiple CI machines for faster feedback. It provides intelligent test splitting, re-run strategies for flaky tests, detailed analytics dashboards, and native Playwright chunking. Achieves up to 50% test suite speed improvements through optimized distribution and parallel execution management.

api-usage-based
DeepTeam logo

DeepTeam

Open-source LLM red-teaming framework with 40+ attack types

DeepTeam is an open-source red-teaming framework for systematically testing LLM applications against 40+ adversarial attack types. It covers OWASP Top 10 for LLMs including jailbreaks, prompt injection, PII leakage, and hallucination attacks. Built as the sister project of DeepEval for security testing alongside evaluation. Apache-2.0 licensed.

Open Source
DogQ logo

DogQ

AI test writing with self-healing and visual editor

DogQ is a testing tool that writes, heals, and suggests tests automatically through a visual interface for managing the entire testing lifecycle. It provides a low-code alternative for teams without dedicated SDETs (Software Development Engineers in Test), automating test creation from user interactions and maintaining test health as the application evolves with AI-driven self-healing.

freemiumOpen Source
Fairlearn logo

Fairlearn

Python toolkit for assessing and mitigating ML model fairness issues

Fairlearn is a Microsoft-backed open-source Python toolkit that helps developers assess and improve the fairness of machine learning models. It provides metrics for measuring disparity across groups defined by sensitive features, mitigation algorithms that reduce unfairness while maintaining model performance, and an interactive visualization dashboard for exploring fairness-accuracy trade-offs. Integrated with scikit-learn and Azure ML's Responsible AI dashboard.

Open Source
FinalRun logo

FinalRun

AI QA agent specialized for mobile apps

FinalRun is a specialized AI QA agent for mobile applications that automates testing of complex mobile gestures and flows on both iOS and Android platforms. It addresses the specific mobile QA gap that generic web-testing tools miss, handling touch interactions, swipe gestures, device rotation, and platform-specific behavior that require mobile-native understanding.

paidOpen Source
Floci logo

Floci

Free open-source local AWS emulator as a drop-in LocalStack replacement

Floci is a free open-source AWS emulator designed as a lightweight drop-in replacement for LocalStack Community Edition. It runs on port 4566 with the same endpoint conventions, supporting S3, SQS, DynamoDB, RDS, ElastiCache, API Gateway, Cognito, IAM, and twenty-plus other services. The Docker image is ninety megabytes versus LocalStack's one gigabyte and starts in twenty-four milliseconds.

Open Source
Foundry logo

Foundry

Blazing fast Solidity development toolkit

Foundry is a blazing-fast portable toolkit for Ethereum smart contract development written in Rust. It provides Forge for testing and deploying Solidity contracts with native Solidity tests that run orders of magnitude faster than JavaScript-based alternatives, Cast for interacting with EVM chains from the command line, Anvil as a local testnet node, and Chisel as a Solidity REPL. Foundry has become the standard toolchain for serious Solidity development.

Open Source
GPT Driver logo

GPT Driver

AI-powered resilient mobile test flows

GPT Driver is a mobile app QA tool that uses generative AI to create resilient test flows with self-healing capabilities and broad framework support for cross-platform applications. It leverages AI to handle the non-deterministic nature of mobile UI where elements shift between releases, generating tests that adapt to layout changes rather than breaking on every app update.

paid
Ghost Inspector logo

Ghost Inspector

Codeless browser testing with visual test recorder and scheduling

Ghost Inspector provides codeless browser testing through a visual recorder that captures user interactions and converts them into automated test suites. Tests run on managed infrastructure with scheduled execution, CI/CD integration, and Slack notifications. Features visual comparison for UI regression detection, API testing, and test organization with folders and tags for managing large test suites.

freemium

Git Bayesect

Bayesian git bisection for finding commits that caused flaky tests

Git Bayesect applies Bayesian inference to git bisection, solving the problem of finding commits that introduced non-deterministic bugs like flaky tests. Unlike standard git bisect which requires binary pass-fail results, Git Bayesect handles probabilistic outcomes where a test might pass sometimes and fail sometimes, using entropy minimization to efficiently narrow down the culprit commit.

Open Source
Great Expectations logo

Great Expectations

Data quality validation framework for Python

Great Expectations is an open-source Python framework for validating, documenting, and profiling data quality. Teams define expectations as expressive unit tests for their data using an intuitive API, then validate datasets against those rules in CI/CD pipelines or production workflows. It connects to pandas, Spark, and SQL sources, generates data documentation automatically, and integrates with orchestrators like Airflow and Prefect for continuous data quality monitoring.

freemiumOpen Source
Hurl logo

Hurl

CLI tool for running and testing HTTP requests from plain text files

Hurl is a command-line tool that runs HTTP requests defined in simple plain text files with built-in assertions for testing. It supports GET, POST, PUT, GraphQL, multipart, cookies, authentication, and response validation with JSONPath, XPath, and regex. Chain multiple requests with variable capture between steps. 17,000+ GitHub stars, Apache 2.0 license, written in Rust for speed. Ideal for API testing in CI/CD pipelines as a lightweight alternative to Postman collections.

Open Source
JetBrains logo

JetBrains

Professional IDEs with Junie AI coding agent

JetBrains is a software development company that produces a family of language-specific IDEs built on the IntelliJ platform — IntelliJ IDEA (Java/Kotlin), PyCharm (Python), WebStorm (JS/TS), PhpStorm, RubyMine, GoLand, Rider (.NET), and CLion (C/C++). Each IDE provides deep, language-aware tooling: smart refactoring, static analysis, test runners, debuggers, and first-class AI Assistant and Junie coding agent integration.

paid
Knip logo

Knip

Find unused files, dependencies, and exports in JavaScript and TypeScript projects

Knip is an open-source CLI tool that detects unused files, dependencies, devDependencies, and exports in JavaScript and TypeScript codebases. It analyzes the full dependency graph to identify dead code that accumulates over time — especially relevant for AI-generated codebases where unused artifacts pile up faster than manual cleanup can handle. With over 10,800 GitHub stars, it has become a standard code hygiene tool in the JS/TS ecosystem.

Open Source
Krkn logo

Krkn

CNCF Sandbox chaos engineering framework for Kubernetes resilience

Krkn is a CNCF Sandbox chaos engineering tool that tests Kubernetes cluster resilience by injecting controlled failures. It simulates pod kills, node failures, network partitions, CPU/memory pressure, and zone outages. Krkn-AI adds AI-powered scenario generation that suggests chaos experiments based on cluster topology. Supports CI/CD integration for automated resilience testing in deployment pipelines.

Open Source
LambdaTest KaneAI logo

LambdaTest KaneAI

GenAI-powered test agent with natural language test authoring

KaneAI is LambdaTest's GenAI-powered test automation agent that creates, evolves, and debugs tests from natural language descriptions. It generates test scripts in multiple frameworks including Selenium, Playwright, and Cypress from plain English instructions. Features intelligent test maintenance that automatically updates tests when application UI changes and two-way editing between natural language and code.

paid
Lighthouse logo

Lighthouse

Web page quality auditing

Google's open-source automated tool for auditing web page performance, accessibility, SEO, best practices, and Progressive Web App compliance. Generates detailed reports with actionable improvement suggestions and scores from 0-100 for each category. Built into Chrome DevTools, available as a CLI tool, Node.js module, and via PageSpeed Insights web interface. Essential for web performance optimization and meeting Core Web Vitals thresholds that affect Google search rankings.

Open Source
Lost Pixel logo

Lost Pixel

Open-source visual regression testing tool

Lost Pixel is an open-source visual regression testing tool that serves as an alternative to Percy and Chromatic. It captures and compares screenshots of UI components and application pages across Storybook, Ladle, Histoire, and custom screenshot sources like Cypress or Playwright. Integrated directly into GitHub Actions pipelines, it detects unintended visual changes before they reach production, with a free SaaS tier available for open-source projects.

freemiumOpen Source
Lychee logo

Lychee

Fast async link checker written in Rust

Lychee is a fast, asynchronous link checker written in Rust that finds broken URLs and email addresses in Markdown, HTML, reStructuredText, and websites. Available as a CLI tool, Rust library, and GitHub Action, it validates links with configurable concurrency, rate limiting, and retry logic. Supports GitHub token authentication for API rate limit avoidance and can check both internal file links and external HTTP endpoints across entire repositories or websites.

Open Source
MCP Inspector logo

MCP Inspector

Official debugger for MCP server development

MCP Inspector is the official interactive developer tool from the Model Context Protocol team for testing, debugging, and validating MCP servers. It provides a visual interface to inspect available tools, test transport configurations, export configs for different clients, and verify protocol compliance during MCP server development.

Open Source
MCP for Unity logo

MCP for Unity

Open-source MCP bridge between AI assistants and the Unity Editor

MCP for Unity is CoplayDev’s MIT-licensed bridge between MCP-compatible AI assistants and the Unity Editor. It exposes tools for assets, scenes, GameObjects, scripts, tests, profiling, and build-oriented workflows. The community project supports Unity 2021.3 LTS through 6.x and is explicitly not affiliated with Unity Technologies.

Open Source
MCPJam logo

MCPJam Inspector

Test and debug MCP servers before they ship

Open-source platform for inspecting, debugging and regression-testing MCP servers, MCP Apps and ChatGPT apps, with OAuth and protocol conformance for local and CI workflows.

freemiumOpen SourceTelemetry
MSW logo

MSW

API mocking for browser and Node.js

Mock Service Worker intercepts network requests at the service worker level, enabling seamless API mocking for tests and development without changing application code. With 16k+ GitHub stars, MSW is the go-to solution for frontend developers who need realistic API simulation during development and testing without spinning up backend servers.

Open Source
Mabl logo

Mabl

Low-code AI test automation for modern teams

Mabl is a low-code AI test automation platform for end-to-end testing of web apps, APIs, mobile, and accessibility. Uses machine learning for auto-healing tests that adapt to application changes, reducing flaky test maintenance. Features a no-code visual builder, parallel cross-browser execution, performance testing, and native CI/CD integration. Provides unified reporting with insights into test coverage and quality trends. Integrates with Jira, Slack, GitHub, and major CI/CD tools.

paid
Mago logo

Mago

Blazing-fast PHP linter, formatter, and static analyzer in Rust

Mago is a comprehensive PHP toolchain written in Rust that unifies linting, formatting, and static analysis into a single binary. It enforces PER-CS formatting standards, catches code smells with 100+ lint rules, and performs deep type inference for semantic analysis. Inspired by Clippy and OXC from the Rust ecosystem, it delivers performance orders of magnitude faster than PHPStan and Psalm while requiring no PHP runtime to execute.

Open Source
Meticulous logo

Meticulous

AI-powered frontend testing with zero flakiness

AI-powered tool that automatically generates and maintains E2E tests by recording user sessions in production or staging. Eliminates flaky tests by replaying user flows deterministically without writing a single line of test code. Detects visual and functional regressions by comparing recorded flows against new deployments, catching bugs that traditional test suites often miss.

paid
Midscene.js logo

Midscene.js

AI-powered vision-driven UI automation for web, Android, and iOS

Midscene.js is an open-source UI automation framework from ByteDance's Web Infra team that uses vision-based AI models to understand and interact with interfaces. It replaces fragile CSS selectors with natural language descriptions, supporting web browsers via Playwright and Puppeteer, Android via ADB, and iOS via WebDriverAgent from a unified JavaScript SDK.

Open Source
MiniStack logo

MiniStack

Free MIT-licensed drop-in replacement for LocalStack

MiniStack is a free, MIT-licensed drop-in replacement for LocalStack that emulates 33 AWS services using real infrastructure — actual Postgres for RDS, real Redis for ElastiCache, real Docker for ECS — rather than faking API responses. Born from LocalStack's surprise paywall in March 2026, it starts in 2 seconds, idles at 30MB RAM versus LocalStack's ~500MB, and runs real services for accurate local AWS development.

Open Source
Mobile MCP logo

Mobile MCP

MCP server for mobile device automation and testing

Mobile MCP is an open-source MCP server that enables AI agents to automate Android and iOS devices — navigating apps, tapping elements, extracting screen content, and running tests on simulators, emulators, and physical devices. It brings agentic mobile engineering to any MCP-compatible AI assistant.

Open Source
Mocha logo

Mocha

Flexible JavaScript test framework

Feature-rich JavaScript test framework for Node.js with BDD/TDD interfaces, flexible assertion library choice, and extensive plugin support. The veteran test runner trusted by thousands of projects, offering async test support, configurable reporters, and a mature ecosystem that makes it reliable for both unit and integration testing in Node.js applications.

Open Source
Octomind logo

Octomind

AI-powered E2E test generation and maintenance platform

Octomind is an AI-powered testing platform that automatically generates, runs, and maintains end-to-end Playwright tests for web applications. It observes user flows, creates test cases from natural language descriptions, and self-heals tests when UI changes would break traditional selectors. Backed by $4.8M seed funding from Paua Ventures with enterprise production deployments.

freemium
LangChain logo

OpenEvals

Lightweight eval library for LLM applications

OpenEvals is a lightweight evaluation library from the LangChain team for testing LLM application quality using LLM-as-judge patterns. It provides pre-built prompt sets and evaluation functions that score model outputs against criteria like accuracy, relevance, coherence, and safety without requiring complex infrastructure. Available as both Python and JavaScript packages, OpenEvals complements OpenAI Evals with a simpler, framework-agnostic approach to quality measurement in agentic workflows.

Open Source
Oxlint logo

Oxlint

Rust-powered JavaScript linter that is 50-100x faster than ESLint

Oxlint is an extremely fast JavaScript and TypeScript linter built as part of the OXC (Oxidation Compiler) toolchain written in Rust. It runs 50-100x faster than ESLint by parsing and analyzing code in a single optimized pass without requiring any plugins or configurations. Oxlint ships with 520+ built-in rules covering correctness, performance, and style checks, and is designed to run alongside ESLint during migration. Part of Evan You's VoidZero initiative, OXC has over 20,000 GitHub stars.

Open Source
Pact logo

Pact

Consumer-driven contract testing for APIs and microservices

Pact is the de-facto contract testing framework for HTTP APIs and event-driven systems across 12+ languages. It verifies that services can communicate correctly by testing API contracts from the consumer's perspective before deployment, catching breaking changes in microservice architectures. Maintained by Pact Foundation with 7,000+ combined GitHub stars across implementations. PactFlow offers a managed SaaS broker. Used widely in enterprise microservice teams to prevent integration failures.

Open Source
Percy logo

Percy

Visual testing and review platform

BrowserStack-owned visual testing platform for automated screenshot comparison across browsers and screen sizes. Integrates with CI pipelines to catch visual regressions before they reach production. Renders pages in real browsers and highlights pixel-level differences, helping frontend teams maintain visual consistency across every supported browser and viewport combination.

freemium
Playwright logo

Playwright MCP

Microsoft's MCP server for structured browser automation by AI agents

Playwright MCP is Microsoft's Model Context Protocol server that enables AI agents to automate web browsers through structured tool calls. It exposes Playwright's browser automation capabilities as MCP tools for navigation, clicks, forms, extraction, and screenshots. The Microsoft-maintained repo has 30K+ GitHub stars and is a durable default for structured browser interaction in agent workflows.

Open Source
Puppeteer logo

Puppeteer

Headless Chrome Node.js API

Node.js library by Google that provides a high-level API for controlling headless or full Chrome and Chromium browsers programmatically. Used extensively for web scraping, automated testing, PDF generation, screenshot capture, form submission, and performance monitoring. Supports page navigation, DOM manipulation, network interception, and cookie management. Works with Chrome DevTools Protocol directly. The most widely-used browser automation tool in the Node.js ecosystem with 89K+ GitHub stars.

Open Source
QA Sphere logo

QA Sphere

Modern test management with AI organization

QA Sphere is a modern, clutter-free test management system that uses AI to organize test cases and track coverage across complex projects. It streamlines the management side of QA that is often handled through messy spreadsheets, providing structured test case repositories, execution tracking, and coverage analytics for technical project managers and QA leads.

freemium
QA Wolf logo

QA Wolf

AI-generated E2E tests, managed QA service

Fully managed QA service that uses AI to generate and maintain Playwright-based E2E tests. QA Wolf engineers write, run, and maintain tests on your behalf, targeting 80% coverage with custom pricing based on test volume. A unique hybrid approach combining AI test generation with human QA expertise, eliminating the burden of test maintenance that slows most engineering teams.

paid
RagaAI Catalyst logo

RagaAI Catalyst

AI testing and evaluation for agents and LLM apps

RagaAI Catalyst is a comprehensive Python SDK for observability, monitoring, and evaluation of LLM and agentic applications. Provides agent tracing with execution graph visualization, self-hosted dashboard with analytics, synthetic data generation, multi-metric evaluation framework, and guardrail management. Built for teams running production RAG systems and AI agents who need systematic testing, debugging, and performance optimization workflows.

Open Source
Relicx logo

Relicx

Test generation from real user behavior sessions

Relicx uses generative AI to conduct visual and functional testing based on real user behavior captured from production sessions. It prioritizes the user experience in testing rather than just code coverage, generating regression suites from actual user journeys to ensure that the flows real people use most frequently are always covered by automated tests.

paid
Ruff logo

Ruff

Extremely fast Python linter and formatter written in Rust

Ruff is a Python linter and code formatter written in Rust by Astral that runs 10-100x faster than existing tools like flake8, isort, and Black. It implements over 800 lint rules from dozens of popular plugins in a single binary, handles auto-fixing for most violations, and includes a built-in formatter compatible with Black. Adopted by FastAPI, Hugging Face, Pandas, and Apache Airflow, Ruff has over 38,000 GitHub stars and processes entire codebases in milliseconds.

Open Source
SWE-bench logo

SWE-bench

Benchmark for evaluating AI coding agents on real GitHub issues

SWE-bench is a benchmark from Princeton NLP that evaluates AI coding agents by testing their ability to resolve real GitHub issues from popular open-source projects. Each task provides an issue description and repository state, and the agent must produce a working patch that passes the project's test suite. With 4,600+ GitHub stars, it has become the standard yardstick for comparing autonomous coding tools like Devin, Claude Code, and OpenHands.

Open Source
Safari MCP Server parent Safari mark

Safari MCP Server

Apple's Safari-native MCP server for web debugging agents

Safari MCP Server is Apple's safaridriver-based MCP server in Safari Technology Preview, giving compatible coding agents local access to Safari page content, console logs, network requests, screenshots, JavaScript evaluation, interactions, viewport controls, and accessibility/performance checks.

freeTelemetry
Storybook logo

Storybook

UI component workshop

Storybook is an open-source workshop for building, documenting, and testing UI components in isolation. Render every visual state of React, Vue, Angular, Svelte, or Web Components, write interaction tests, visual regression tests, and accessibility checks, and publish a living design reference. Trusted by thousands of teams including GitHub, Airbnb, Shopify, and Microsoft as their single source of UI truth.

Open Source