Skip to content
aicoolies logo

Explore / The tool atlas

Categories

Explore developer tools by the work they support, from coding and design to data, security, and deployment. Start with a broad category, then follow its subcategories to narrow the list and compare tools that address a similar part of your workflow.

18 categories · 20 subcategories

Find your next building block.

These 81 tools handle the data side of a machine-learning system: getting data in, versioning it, checking it, turning it into vectors, training on it and tracking what happened. It is a category for the person who owns the pipeline rather than the model, and it leans heavily open: 71 of the 81 entries (87.7%) ship under an open licence. That ratio changes how you should read the page. In most categories the free option is the compromise; here it is usually the default, and the paid product is a hosted layer on top. Kubecost-style splits are everywhere: Cleanlab is a free library with a $500/mo Studio SaaS above it; Metabase is free open source with Starter at $100/mo and Enterprise from $20k/year; lakeFS and DVC are Apache-2.0 with an enterprise tier attached. Ray (92/100 in Ray) and LLaMA-Factory (91/100, LLaMA-Factory) are both free and both carry the highest scores in the category, which is not something you can say about the coding-tool shelves. The practical axis of choice is which stage of the pipeline hurts. For versioning, DVC brings Git semantics to datasets and pipelines while lakeFS does the same for object storage and data lakes — different blast radius, same idea. For processing, Polars is a Rust DataFrame library aimed squarely at pandas-shaped pain. For fine-tuning, LLaMA-Factory is the broad framework and Unsloth (88/100, Unsloth, Apache 2.0) is the option for constrained hardware. For retrieval, Marqo does tensor search and Hugging Face's Text Embeddings Inference serves embeddings, rerankers and classifiers. For data quality, Cleanlab (79/100) finds label errors. For tracking, Weights & Biases (85/100, free plan, Pro from $60/mo) is the commercial default; it is one of only 10 non-open entries here. One honest caveat: review coverage is thin relative to the size of the shelf — 7 of the 81 tools have a scored review. Everything else on the page is indexed and verified but not yet graded, so treat the absence of a score as "not reviewed yet", not "not recommended". Where a score does not exist, the licence line and the stack-membership count are the most useful signals in the card. DVC, for example, carries no score but sits in 3 stacks, which tells you it survived somebody's real setup. Nothing in this category is in the graveyard. Start from the stage that is failing, check the licence, and only then look at what the hosted tier costs.

Two different kinds of failure share this shelf, and conflating them is the most common mistake teams make when they shop here. One kind is the ordinary sort: a service is down, latency spiked, an exception was thrown. The other is specific to LLM applications — nothing crashed, every request returned 200, and the answers got worse. The tooling for those two problems overlaps far less than the category name suggests. For the second kind you need trace-level visibility into prompts, retrieved context, model responses and evaluation results. Langfuse (87, Langfuse) is the centre of gravity: it appears in 13 of the 589 published comparisons, the highest comparison count of any tool in this batch, and in 9 stacks. It is open source and self-hostable, with Hobby free, Core from $29/mo and Pro from $199/mo. Arize Phoenix (87, Arize Phoenix) is the OpenTelemetry-native alternative, free and open source, and worth a look specifically if you already emit OTel and would rather not run a second tracing convention. LangSmith (80, LangSmith) is the narrowest of the three — its own verdict tells teams outside LangChain and LangGraph to compare Langfuse and Phoenix first. Helicone (84, Helicone) solves a smaller problem with almost no integration cost, since changing a base URL is the whole setup; Hobby covers 10,000 requests, Pro is $79/mo. For the first kind, the incumbents are here too and they score well. Grafana (90, Grafana) is the highest-scoring tool in the category, AGPL v3 self-hosted and free. Prometheus (85) is Apache 2.0 with no commercial version at all. Sentry (86, Sentry) sits in 6 stacks and is free self-hosted, or $26/mo on Team. Datadog (88) buys breadth — free for 5 hosts, then from $15/host/mo on Pro. The consolidation to keep in mind: two entries here are graveyard records, not options. Humanloop's own migration guide states the platform was sunset on September 8th, 2025 following its acquisition (humanloop.com, accessed 2026-08-18); WhyLabs is likewise flagged. If a comparison article still lists either as a live choice, it predates that. This is the best-covered category in this batch — 46 of 69 tools carry a scored review, roughly two-thirds — so the verdict lines are unusually reliable here as a filter. 45 of 69 (65.2%) are open source. Practical sequence: pick one infrastructure tool and one LLM-trace tool rather than hunting for a single product that does both well, and check whether your framework already emits OpenTelemetry before you commit to a proprietary SDK.

This category holds the 103 tools that find security problems in code, in dependencies, in pipelines and — a newer job — in the models and agents a team now ships. If you are responsible for what a scan reports on Monday morning, this is your shelf. The page renders 48 of the 103, so it helps to know which of five jobs you are hiring for. Secrets. Gitleaks (84/100, Gitleaks, MIT and free) and TruffleHog (86/100, TruffleHog) both scan git history for credentials; TruffleHog's distinguishing move is verifying whether a found secret is still live, which is what separates a real incident from noise. Code. Semgrep (87/100, Semgrep) and SonarQube (87/100, SonarQube) sit in CI and enforce rules; SonarQube's Community Build is free, Semgrep's free tier stops at 10 repos and 10 contributors. Consolidated platforms. Snyk (86/100, Snyk, free tier, Team from $25/mo) and Aikido Security (86/100, free for 2 users, Basic $300/mo) cover several attack surfaces at once, which is a procurement decision as much as a technical one. Code health as a gate. DeepSource (84/100) and CodeAnt AI (82/100) blend static analysis with AI review and reporting. The fifth job is the one that changed in 2026: securing the AI itself. Of the six most recently touched entries here, five belong to that wave, all verified on 2026-08-16 — MCP-Scan, which scans MCP servers for tool poisoning and prompt injection; DeepTeam, an open-source LLM red-teaming framework covering 40+ adversarial attack types; Giskard, for bias, drift and model vulnerability testing; Microsandbox, which gives agents hardware-isolated microVMs to run code in; and Inspect AI, the UK AI Security Institute's MIT-licensed evaluation framework. NVIDIA's garak (verified 2026-04-21) sits in the same group and already appears in 3 stacks. None of these replace a SAST tool. They answer a question a SAST tool was never asked. Two caveats before you browse. First, review coverage here is thin: 30 of the 103 tools carry a scored review, so absence of a score is not a verdict — it usually means the tool has not reached the review queue yet. Second, nothing in this category is currently in the graveyard, which is unusual on this site and worth reading as a sign that the field is still adding rather than consolidating. If you are starting from zero, the cheapest useful sequence is a secret scanner in pre-commit, a rules engine in CI, and only then a platform contract — in that order, because the first two are free and catch the failures that actually get exploited.

A Model Context Protocol server is a small process that hands one capability to an agent — repository access, live documentation, a browser, a database — over a standard interface, so the agent does not need bespoke glue per integration. If you run Claude Code, Cursor or Copilot and you have ever copy-pasted a file into a chat to give it context, this is the shelf that removes that step. The useful split on this shelf is between servers that do something and infrastructure that manages servers, and most readers arrive shopping for the first while the second is what eventually bites them. Capability servers are the obvious half. Context7 (88 in Context7) serves version-pinned documentation, which targets the failure mode where a model writes an API call that was correct two releases ago. GitHub MCP Server (89) exposes the repository, issue and pull-request surface directly. chrome-devtools-mcp (88, Apache-2.0) lets an agent profile and debug a real page instead of reasoning about it. Serena (86) attaches a language server so the agent gets symbol-level understanding rather than grep. Firecrawl MCP Server (84) covers the open web. The management half matters as soon as you are running more than two or three. Every server you install is another process with your credentials, and the honest read on directories like Smithery (78) is that convenience and trust are traded against each other — you are running code someone else published. That is exactly the gap Docker MCP Gateway (87, MIT) addresses with a catalog, image verification, container limits, secret scanning and tool allowlists. FastMCP (87, Apache-2.0) is the other side of the same problem: if the server is yours, you control what it can reach, and authentication and logging become your responsibility rather than a vendor's. Recent additions to the catalog track that shift: a gateway, an official registry, and a vulnerability scanner are plumbing rather than capability. When a scanner and a registry land in the same batch, the ecosystem has moved past the "install everything and see" phase, and your install policy should move with it. Most tools here are open source, which for this category is more consequential than usual: an MCP server you cannot read is an unaudited process holding your tokens. Most do not yet carry a scored review. And verify licenses individually — Directus is free to self-host under MSCL-1.0-GPL but is not flagged open source in our catalog, and Vanna AI's public repository is archived and read-only even though the commercial product continues. Pick one capability server, run it for a week, and decide on a gateway before you install the fourth.

Everything on this page is a bill, not a binary. These 18 entries are the plans, API tiers and credit bundles that sit underneath the tools in the rest of the catalogue — the thing you are actually paying for when a coding agent says it is free. Nothing here is open source, 0 of 18, which is unusual in this catalogue. If you want to avoid a subscription entirely, the route is a bring-your-own-key client from the CLI agents category paired with an API tier here, not a free plan on this page. The entries split into two kinds. General assistants are what a whole team uses: Claude (93/100 in Claude, the highest of the seven scored entries here — Free, Pro $20 monthly or $17/mo annual, Max at $100 or $200/mo), ChatGPT (92/100, Plus $20/mo, Pro $200/mo), Gemini (86/100, Google AI Pro $19.99/mo, Ultra $249.99/mo), Mistral AI (88/100, Pro $14.99/mo, Team $24.99/user/mo), Perplexity (87/100, Pro $20/mo) and Grok (80/100, with API pricing quoted per model — Grok 4.3 at $1.25/$2.50 per million input/output tokens). Demand is concentrated: Claude, ChatGPT and Codex each appear in 12 published comparisons, more than double any other entry on this page. Coding-credit plans are the newer and faster-moving half, and they are where the price spread lives. Z.AI's plan runs $18/mo for Lite ($12.60 annually, 10,000 credits a week) up to $168/mo for Max. Minimax starts at $10/mo, Alibaba's plan at roughly $10/mo for Lite, OpenCode Go at $5 for the first month then $10/mo. Kilo Pass sells zero-markup credits from $19/mo. That is an order of magnitude below the flagship tiers, and the trade is model choice and capacity rather than raw access. Capacity is the 2026 constraint worth planning around. Cerebras Code lists Pro at $50/mo for 24 million tokens a day and Max at $200/mo for 120 million — and both are recorded as sold out, with API usage-based pricing as the fallback. A plan you cannot buy is not a plan. One removed entry: Windsurf Plans is a graveyard record. Legacy Windsurf pricing was folded into Devin Desktop on 2 June 2026 and windsurf.com/pricing now redirects to Devin pricing, so any budget still carrying that line needs rechecking against current Devin plans. Six of these records were re-verified on 2026-08-17, the freshest set of any category in this batch — which is appropriate, because pricing is the field that decays fastest. Confirm against the vendor's own page before you commit a budget.

An agent framework is the library you build an LLM application with — it owns control flow, state, tool dispatch and retries so you are not hand-rolling a while-loop around a chat completion. You need one the moment your feature has to make several model calls in sequence, remember what happened between them, and recover when step four throws. Most of these entries are not direct competitors. Almost all of the choice collapses into two questions, and neither is answered by a feature table. The first is what shape your control flow actually has. If it branches, pauses for a human, and has to survive a process restart, you want an explicit graph with checkpointing — LangGraph, which scores 86 in LangGraph, is built around exactly that model. If the work decomposes into roles that hand off to each other, CrewAI's role abstraction (81, CrewAI) is a shorter path. If what you need is a typed function that returns validated structured output, Pydantic AI (85) is closer to ordinary Python than either. The second question is which layer you are missing, because a third of this category does not compete with the orchestrators at all. Mem0 (88) is a memory store. E2B (87) is a Firecracker sandbox for running model-written code. Browser Use (85, MIT) and Stagehand (85) give an agent a browser. LlamaIndex (87) is retrieval and document parsing first, orchestration second. Adding one of these to a framework you already have is usually the right move; replacing your framework to get one is not. Two things are worth carrying into the decision. Most tools here are open source, so by default you can read the control loop before you depend on it. And most do not yet carry a scored review, so a card without a score has not been assessed yet; it has not scored badly. The lifecycle event that defines this category in 2026 is the end of the managed-thread era. OpenAI's Assistants API shut down on 26 August 2026, and OpenAI's own migration guidance points to the Responses API and the Agents SDK. Retired tools are not shown in the list below. Start from the shape of your control flow, then check whether you need a framework at all or just a layer.

This is the window you sit in for eight hours, so switching costs are real and the decision is worth more care than a tool you invoke once a day. The category indexes 55 environments — desktop editors, AI-native forks, JetBrains agents, cloud workspaces and notebooks — and the grid shows 48 of them. Three questions settle most of it. Where does the code run? Locally with VS Code (90/100 in VS Code, free), Cursor (91/100, Cursor, Hobby free and Pro $20/mo) or Zed (85/100, Zed, free and open source); or in a browser with Replit (81/100, Starter free, Core $25/mo), StackBlitz (85/100, Starter free, Teams $19/user/mo) or Gitpod (82/100, 50 free hours a month). Is the AI a fork or a guest? Cursor, Trae (78/100, free including all AI features) and Google Antigravity (88/100, free preview then $20/mo Pro) rebuild the editor around the agent; VS Code and the JetBrains IDEs keep the editor and let the agent in as an extension, which is what Junie (86/100, Junie) does — its argument is that JetBrains already understands project structure, inspections and refactoring, so the agent inherits that. How much process do you want? Kiro (80/100, free 50 credits/mo, Pro $20/mo) trades generation speed for a spec-driven workflow; Tessl, verified 2026-07-15, takes the same bet from outside the editor. This is also the category with the heaviest attrition in this set: five of its members are in the graveyard, and three of those closed inside eighteen months. JetBrains discontinued Fleet on 22 December 2025, refocusing on JetBrains Air. Windsurf became Devin Desktop on 2 June 2026 under Cognition, and its old URLs now redirect — the aicoolies entry still carries an 88/100 review score, but it is an archived record, not a live option. AWS Cloud9's standalone era ended on 25 July 2024, Codeanywhere on 1 September 2024, and Atom was archived on 15 December 2022. If you are choosing an editor to standardise a team on, that list is the most useful thing on this page. Practical advice: if you already have a working setup, the honest question is whether an agent-native fork buys you more than an extension in the editor you know. If you are starting fresh, pick on where the code runs first — local versus browser is the choice you cannot cheaply reverse. 17 of the 55 tools here carry a scored review, so the graded set is small enough to read end to end before you commit.

This shelf holds two things that are not substitutes for each other: the terminal you type into, and the agents that run inside it. Ghostty and Claude Code both live here and there is no version of the question where you pick one over the other. Read the category as an emulator layer plus an agent layer, and the 105 entries stop looking like a ranking. A scope note before you start: 27 of the 48 cards below also appear under AI CLI Agents, which indexes 40 tools and covers the agent half on its own. Use this page when you want both layers in one view. For the agent layer, the axis that actually decides the purchase is who pays for inference and on what terms. Claude Code (92 in Claude Code, present in 33 of the 589 published comparisons and 20 of the 131 stacks) is bundled into a Claude Pro or Max subscription or billed through the API — you buy the model and the agent together, and you cannot swap the model. Aider (83) and Goose (84) invert that: the client is free, you bring your own key, and the bill lands with whichever provider you chose. Amp (85) sits in between with a free credit grant up to $10/day and pay-as-you-go credits at zero markup for non-enterprise users. Kilo Code (82) is free and BYOK in its open-source tier but sells Teams seats at $15/user/mo. None of these is cheaper in the abstract; the answer depends entirely on how many tokens a working day costs you. The second thing worth checking before you commit is what the licence says rather than what the README implies. 90 of the 105 tools here (85.7%) are flagged open source, but the exceptions are not uniform. Crush (84) ships under FSL-1.1-MIT — source-available now, MIT later — which is a different risk profile from OpenCode's plain MIT. Kiro (80) and Codex (80) are proprietary clients over vendor plans. If your organisation has a policy on copyleft or on source-available licences, that policy will decide this shortlist faster than any benchmark. Two dated changes from 2026 that a stale bookmark will not tell you. Google moved unpaid and Google One users of Gemini CLI to Antigravity CLI on 18 June 2026; the Standard, Enterprise and Google Cloud paths continue, so "Gemini CLI is free" is now true only for some readers. And Charm archived the Mods repository on 30 June 2026, with its last observed push on 2026-03-09. Mods, exa and fig are all graveyard records here and none of them appears in the grid below. Only 32 of the 105 tools carry a scored review, so treat an unscored card as unassessed rather than weak. If you are choosing an emulator, start with Ghostty and stop; if you are choosing an agent, start from your token budget and your licence policy, then read two reviews.

An AI coding assistant writes, edits or reviews code alongside you. This shelf tags 100 tools directly — which is exactly why the label is close to useless on its own: it covers products that share almost nothing beyond the adjective. The parent hub list also includes tools from child shelves (code-review-ai, chat-based-coding, inline-completion), so the grid below can show more than those 100 direct tags. Start by deciding which kind of assistant you are shopping for. The division that matters is not which model sits underneath — most of these now run several — but where the tool lives and how much it may do before it asks. Four shapes cover almost everything here. Inline completion inside an editor you already use: GitHub Copilot, 88/100 in GitHub Copilot, free for 2,000 completions a month and $10/mo for Pro. An AI-native editor that owns the whole window: Cursor, 91/100 in Cursor, which appears in 42 of the 589 published comparisons — more than any other tool on this page. A terminal-resident agent that edits files and runs commands on its own: Claude Code at 92/100 (Claude Code), Aider at 83/100 and free if you bring your own API key. And an asynchronous reviewer that never touches your editor but comments on pull requests: CodeRabbit, 88/100, free for public repos and $24/user/mo after; Greptile, 85/100, $30/seat/mo with 50 reviews included. Those four shapes stack without redundancy — an authoring tool plus a review bot is a coherent setup. Two tools of the same shape is where the budget leaks. Quality and security gates such as Semgrep (87/100), SonarQube (87/100) and Snyk (86/100) may still appear in this parent list via child shelves when they ship AI-assisted analysis — they are not coding-authoring assistants tagged on this shelf itself. Consolidation was the story of 2026. Six tools from this category sit in the graveyard, three of them retired this year: Continue was acquired by Cursor and its standalone product discontinued on 18 June 2026, the Roo Code extension shut down on 15 May 2026 with its GitHub repository archived, and Codeball's PR-triage product went dark on 12 June 2026. Kite (2022), GitHub Copilot X (2024) and Phind fill out the list. Read that as a warning about picking on novelty: the survivors are mostly the tools with a business model you can name. Coverage on the tools tagged here is unusually deep — 65 of the 100 have a scored review — so the practical route is to pick your shape first, open the review of the two or three candidates in it, and only then look at price. Where two tools are close, the head-to-head comparison pages will usually be more decisive than either review alone.

These are the tools that turn a prompt, a Figma file or a screenshot into frontend code. The question that actually matters is not how good the first screen looks — all of them clear that bar now — but what happens on the second iteration, when a designer moves a component and you have to decide whether to regenerate or edit by hand. The 61 entries answer that question in three different ways, and confusing them is the main way readers waste a week. Prompt-to-app builders own the whole loop: you describe the app, they generate, host and iterate. Bolt.new (84 in Bolt.new) runs the stack in WebContainers for instant feedback and prices at 1M tokens a month free, Pro at $25/mo. Lovable (82) is in 7 published comparisons, more than any other entry in the ranked set below, and generates React with Supabase behind it. v0 (87) scores highest of the three, and is the one whose React and Tailwind output is closest to code you would have written. Replit (81) adds the browser IDE and hosting around the same idea. Design-source bridges do less and survive longer. Rather than generate an app, they feed real design context to an agent you already run. Figma MCP Server (82) is the first-party route — free during the beta according to Figma's own docs, with Figma saying the capability will eventually become usage-based. Figma Context MCP (84) is the open-source alternative and scores higher; Talk to Figma MCP (81, MIT) goes further and writes back, which its review flags as too permissive for read-only governance needs. Then there is the substrate the output lands in, which is the part nobody shops for and everybody needs. Panda CSS (86, MIT) gives you zero-runtime styling. Storybook appears in 4 of the 131 stacks — more stack membership than most scored tools ranked here — and has no published review yet. Figma itself sits in this category too, at 90, the highest score among the entries ranked here; it is the source file, not a code generator, so read that number as a statement about Figma and not about design-to-code. What dies in this category is instructive: not the generators, but the targets. React deprecated Create React App on 14 February 2025 for want of active maintainers, pointing new apps at Vite, Next.js, React Router or Expo. Galileo AI was acquired by Google in May 2025 and now exists as Google Stitch. Both are graveyard records and neither appears in the grid below. If your generator emits a scaffold, check what that scaffold is before you inherit its maintenance. 38 of the 61 tools here (62.3%) are open source and only 12 carry a scored review — just under 20% coverage, so an unscored card means unassessed, not rejected. Decide first whether you want generation or context, then read one review from that half.

DevOps & Deployment covers everything between a merged pull request and a service that answers requests: CI runners, hosting platforms, container orchestration, infrastructure-as-code, and the quality and security gates that sit in the pipeline. It is the largest category indexed here — 321 tools, of which the page renders 48. Worth saying plainly: those 48 are a cut-off, not a shortlist. The groupings below exist to give the cut-off an editorial reason. The axis that actually separates these tools is not their feature lists, it is who operates the machine. At one end, managed platforms absorb the operational burden. Vercel (89, Vercel) appears in 18 of the 131 published stacks — more than any other deployment target in this category — and the judgement its review lands on is about the invoice rather than the capability. At the other end are platforms you run yourself: Coolify (84, Coolify) is free self-hosted with cloud from $5/mo, and is framed in its review as Vercel-like convenience without the lock-in. Between the two sit tools that describe infrastructure instead of hosting it. Terraform (85, Terraform) is the obvious example, and note that the catalogue does not record it as open source; its review names the BSL licence change as a legitimate concern for some organisations. The second thing the demand ranking shows is a category quietly changing shape. The most-compared tool on this shelf is not a deploy target. Ollama (88, Ollama) is free, open source, and appears in 11 of the 589 published comparisons and 11 stacks — the highest comparison centrality of anything in this category. Open WebUI (88), AnythingLLM (86) and Llamafile (79) sit near the top of the same list. Model-serving has become part of the deployment problem, and the six most recently verified entries here confirm the drift: DronaHQ, Superblocks, LibreChat, Unsloth, Coroot and K8sGPT, all verified 2026-08-16, are internal-tool builders, chat front-ends, a fine-tuning stack and Kubernetes diagnostics. Five entries are flagged as graveyard records rather than live options — Travis CI OSS, the Heroku free tier, Lerna, Create React App and Rome. Treat them as history, not candidates. Reading order that works: settle your operating position first — managed, self-hosted, or declared — because it eliminates most of the 321 immediately. Then pick a CI runner, then the gates. 89 of the 321 tools carry a scored review; where a tool has one, the verdict line is the fastest way to see who it is not for.

Three separate decisions are stacked on this one page: where your data is stored, how your application code talks to it, and — increasingly — whether retrieval belongs in the same engine as the rest of your data. 98 tools are indexed and 48 render. Knowing which of the three questions you are answering removes most of them straight away. Start with storage, because it constrains the rest. Supabase (90, Supabase) appears in 24 of the 131 published stacks — more than any other tool anywhere in this batch — and its review's argument is that you get a real Postgres database rather than a proprietary data store. Neon (90, Neon) scores the same and takes a different angle: serverless Postgres with branching and scale-to-zero, free for 100 projects at 100 compute-hours per project per month, with usage-based Launch pricing around $15/mo. Firebase (84, Firebase) is the non-Postgres option, fast to production and tagged in the catalogue with telemetry concerns. The access-layer decision is narrower than the shelf makes it look. Drizzle ORM (88, Drizzle ORM) sits in 7 stacks and its review's core claim is that it respects SQL instead of hiding it. Prisma (85, Prisma) trades that for developer experience and pays in bundle size and runtime overhead, per its own verdict. Both are free and open source; the choice is a taste question about how much SQL you want in your codebase, not a capability gap. Then there is the thing this category has actually become. Eight of the twelve highest-demand tools here are vector stores or search engines — Qdrant (88), Pinecone (87), pgvector (86), Weaviate (85), Chroma (84), Milvus (84), plus Supabase and Neon carrying vector workloads on Postgres. Every one of the six most recently verified entries (Upstash Vector, LanceDB, OpenSearch, Vespa, Cloudflare Vectorize and Ragie, all checked 2026-08-16) is a retrieval product. If retrieval is your actual question, the dedicated vector-database page is the better shelf; what belongs here is the narrower decision of whether pgvector on your existing Postgres is enough, given it is free, open source, and appears in 7 of the 589 published comparisons. Two caveats before you shortlist. 71 of the 98 tools (72.4%) are recorded as open source, but licence terms vary sharply within that — Weaviate is BSD 3-Clause, Chroma and Qdrant are Apache 2.0, and Directus (85) is free self-hosted under MSCL-1.0-GPL, which the catalogue does not count as open source. And only 30 of 98 tools carry a scored review, so most cards on this page have demand data but no verdict yet. No entries here are flagged as graveyard records.

Two different jobs share this shelf, and confusing them wastes months. One is checking that code does what it was written to do — unit runners, end-to-end browser suites, CI gates. The other is checking that a model does what you hoped it would, which is a scoring problem, not an assertion problem. A expect(x).toBe(y) has no useful analogue when the output is a paragraph. Read the ranking here and the split is obvious. Among the twelve highest-demand members, Playwright carries the highest review score at 91/100 (Playwright) and sits in 15 published stacks. Vitest follows with 85/100 and 12 stacks. Both are free, both are deterministic, both answer "did it break". Directly alongside them sit DeepEval (88/100), Langfuse (87/100), Promptfoo (86/100), TruLens (83/100) and RAGAS (79/100), none of which will tell you a test failed — they tell you a score moved, and you decide whether that matters. If you are choosing browser automation, the useful axis is how much determinism you are willing to give up for resilience against DOM churn. Playwright is fully deterministic and gives you the Trace Viewer when something fails at 3am. Stagehand (85/100) keeps Playwright underneath and adds act/extract/observe primitives on top. Skyvern (80/100, AGPL-3.0 self-hosted) goes furthest, driving pages by vision; our review records an 85.85% success rate on the WebVoyager benchmark, which is genuinely high and still means roughly one run in seven does not complete. Price that in before putting it in CI. On the LLM-evaluation side the split is ownership, not capability. Promptfoo and DeepEval run in your repo as libraries. LangSmith (80/100) is hosted, free to 5,000 traces a month then $39/seat/mo, and our review is explicit that teams not already on LangChain or LangGraph should compare Langfuse and Arize Phoenix first. One removal worth knowing: Google's Web Vitals Chrome extension reached end of life in January 2025 after its functionality shipped into the Chrome DevTools Performance panel. It is kept here as a graveyard record, not a recommendation, and is excluded from the 108 tools the page counts. 75 of the 109 catalogued members (68.8%) are open source, and 42 carry a scored review. Start from a picked tool below, then read the review before you adopt — the verdicts are where the trade-offs live.

Four jobs are filed together here, and they have almost nothing in common beyond producing text. Publishing an API reference, running a docs-as-code site, storing team knowledge, and giving a coding agent something machine-readable to load are four different purchases. 54 tools are indexed and 48 render; picking the right job first is what makes the shelf usable. For API reference, Scalar (89, Scalar) is the highest-scoring tool in the category — open source, free at $0, with Pro at $72/month — and its review credits the combination of readable design, multi-language examples and built-in API testing. Swagger/OpenAPI remains the underlying specification work, free as open-source tooling with SwaggerHub from $95/mo. For docs-as-code sites, Nextra and Fumadocs are both free Next.js-based generators; neither carries a scored review yet. For hosted documentation with the least setup, Mintlify (82, Mintlify) appears in 5 stacks, the most of anything on this page, at Free on Starter and $150/mo on Growth. Team knowledge is a different axis again, and it is the local-versus-hosted question: Obsidian (88, Obsidian) keeps Markdown on your disk, Notion (86, Notion) keeps everything in a workspace you rent, free per member with Plus at $12/member/mo. The fourth job is new. Documentation is now something coding agents read, and the catalogue's recent verification pass reflects that — AGENTS.md, GitMCP and FastGPT all appear alongside the conventional tools. Codebase Memory MCP (84, Codebase Memory MCP) is an MIT-licensed local MCP server whose verdict is refreshingly narrow: use it when agents struggle with large or unfamiliar repositories, skip it when your projects are small. Two honesty notes before you rely on this page. First, review coverage here is the thinnest in this batch: 5 of 54 tools have a scored review, under one in ten. Seven of the twelve highest-demand entries — Payload CMS, Strapi, GitMCP, Sanity, Nextra, Fumadocs and Swagger/OpenAPI — have no score at all. They rank on demand signal, not on assessment. Second, several of those unscored entries also carry older verification dates: Payload CMS, Strapi, GitMCP, Sanity and Nextra were last checked on 2026-04-21 and Fumadocs on 2026-05-22, roughly four and three months before this draft. Pricing and licence terms move faster than that, so confirm anything cost-sensitive against the vendor before committing. At 28 of 54 (51.9%) open source, this is the least open-source-weighted category in this batch. That is a fair reflection of the field: hosted documentation platforms have no self-hosted equivalent that matches them on polish, and the headless CMS options are where the open-source strength actually sits. No entries are flagged as graveyard records.

Thirty tools are indexed here and 29 render as cards, but they answer two quite different needs, and knowing which one you have will save you scrolling. The first is the browser as your own workspace: things that make GitHub, JSON and page debugging less tedious. Requestly (83/100 in Requestly) intercepts and rewrites HTTP traffic, mocks endpoints and replays sessions — free plan at $0, Pro at $12/user/month monthly or $9 annually, with an AGPLv3/proprietary source posture rather than an unrestricted self-hosting promise. React Scan (MIT, no paid tier) highlights unnecessary React re-renders and installs via npm, CLI, script tag or extension. JSON Crack (Apache 2.0) turns JSON into a navigable graph. Octotree adds a file-tree sidebar to GitHub repositories, Refined GitHub applies hundreds of UI fixes to the same site, and JSON Formatter pretty-prints JSON in place. The second need is the browser as infrastructure for something else — a machine driving a page. Hyperbrowser (80/100, Hyperbrowser) sells that as hosted capacity, priced per unit: $0.10 per browser-hour, $10/GB of proxy data, $0.02 per AI-agent step. Notte follows the same model at $20/month with $20 of credits, sessions at $0.05/hour and LLM tokens passed through at provider price with no markup. The open alternatives run on your own machine instead: Agent Browser (Apache 2.0, headless), Page Agent (MIT), BrowserOS (open source, local-first and privacy-focused, BYOK), Webwright, and Midscene.js (MIT, aimed at end-to-end testing). Apple's Safari MCP Server is a free developer preview in Safari Technology Preview and needs remote automation enabled on a supported macOS release. Be aware of how thin the review coverage is here: 2 of the 30 tools carry a scored review. That is not a quality signal about the other 28, but it does mean you should weigh the licence line and the comparison and stack counts on each card more heavily than usual — those are the fields backed by data rather than by an unwritten review. One removal worth knowing about, because plenty of guides still recommend it: Google's Web Vitals extension was discontinued on 7 January 2025 with Chrome 132, and its real-time Core Web Vitals overlay moved natively into the Chrome DevTools Performance panel. Its aicoolies record was refreshed on 2026-08-16 but is kept as a graveyard entry — it does not appear in the grid below, and the replacement is already in your browser.

If you are here, you are probably solving one of four unrelated problems: tracking work, taking notes, poking at an API, or wiring two systems together without writing a service. This category holds all four — 129 tools, 48 of them shown — which makes it the least coherent shelf on the site and the one most in need of a map. The question that cuts across all four is where your working state lives. Obsidian (88, Obsidian) keeps notes as Markdown files on your disk; its review's argument is longevity — the notes outlast the app. Bruno (78, Bruno) makes the same bet for API collections, storing them as plain text in your repository so they diff and merge like code, and it earns 6 of the 131 published stacks despite the lowest score among the top-demand tools here. Against that, Linear (89, Linear), Notion (86) and Zapier (88) hold your state in their database, and that is the trade you are actually making. Linear is free on Starter and $8/user/mo on Plus; Notion is free per member with Plus at $12; Zapier's free tier is 100 tasks/month with paid plans from $19.99/mo. None of those numbers is the deciding factor. Whether you can walk away with your data usually is. The change worth flagging is that this category absorbed agent operations during 2026 without anyone renaming it. Taskmaster AI (85, Taskmaster AI) turns a PRD into a task graph for coding agents. Vibe Kanban (86, Vibe Kanban) runs several coding agents in parallel across isolated Git worktrees. Agent Orchestrator (86, Agent Orchestrator) is a control plane for the same problem. Three of the twelve highest-demand tools on a productivity page are now agent schedulers, and none of them existed as a category five years ago. One honest signal from the data: Zapier and Make (85, Make) both appear in zero of the 131 published stacks, despite scoring 88 and 85. High individual scores, no evidence of stack adoption on this site. Read that as "the reviews rate them well, the stack authors have not reached for them", not as a defect. Coverage here is thin — 29 of 129 tools carry a scored review, so most cards on this page have no verdict behind them yet. One entry, Insomnia's open-source edition, is flagged as a graveyard record. Start from the state question, then narrow by the specific job; the four jobs on this shelf rarely share a winner.

Agent skills and prompts are portable instructions you add to a coding agent: SKILL.md skill folders, skill packs for Claude Code, Cursor, and Codex, prompt pattern libraries, and the registries that distribute them. This shelf covers shareable instruction packs and worked examples, not the agents, MCP servers, or frameworks that run them. Skill packs are the core of the page, from Anthropic's Agent Skills and Hugging Face Skills to community collections such as Awesome Agent Skills. Rules files and repository instruction formats (AGENTS.md, Cursor Rules, and spec formats such as Spec Kit and OpenSpec) belong to the system prompts and rules subcategory and are listed here as well. Reusable prompt pattern libraries such as Fabric round out the shelf. When choosing, prefer a pack built for the agent you actually use over a generic prompt collection, check when the repository was last updated, and use the open SKILL.md or AGENTS.md conventions when the same instructions need to work across several tools.

Where do the weights actually run? That single question sorts this page, and it is worth answering before you compare anything else, because it determines your cost curve, your latency floor and your data-residency story all at once. The catalogue reflects the split almost exactly: 31 of the 61 catalogued entries (50.8%) are open source. Roughly half of this shelf is a vendor you send tokens to; the other half is software you install. Few categories divide that evenly, and the even division is not an accident — it is the state of the market in 2026. On the rented side, published per-token pricing is now precise enough to model before you commit. The Anthropic API (90/100, Anthropic API) lists Haiku 4.5 at $1/$5, Sonnet 4.6 at $3/$15 and Opus 4.7 at $5/$25 per million input/output tokens as of 16 August 2026 — a fivefold spread between the cheapest and most expensive model from a single vendor, which usually moves a bill more than switching vendors does. On the consumer side Claude scores 93/100 (Claude, verified 17 August 2026) and ChatGPT 92/100 (ChatGPT); both appear in 12 of the 589 published comparisons, so the head-to-head questions are already answered on the comparison pages rather than here. Gemini (86/100) lists a free tier alongside AI Pro at $19.99/mo and AI Ultra at $249.99/mo. Note that both Gemini and DeepSeek (90/100) carry a telemetry-concerns tag in our catalogue — a flag on data handling, not on quality. On the owned side, the choice is throughput versus convenience. vLLM (91/100, free and open source) is the serving layer for production GPU workloads where utilisation is the constraint; our review is explicit that it is infrastructure, not an application framework. Ollama (88/100) is the opposite trade — free, in 11 published stacks and 11 comparisons, and it exposes an OpenAI-compatible endpoint so most client code does not change. LM Studio (84/100) is the same idea with a desktop GUI instead of a terminal. Together AI (89/100) sits between the two camps: someone else's H100s at $6.49/hr, running open weights you chose. One historical entry to ignore as an option: Google Bard was discontinued in February 2024 and is kept only as a graveyard record. It is already excluded from the 60 tools this page counts. 20 of the 61 entries carry a scored review. Start with the shortlist, then read the pricing summary on the tool page — provider pricing in this category moves faster than anything else on the site.

Latest community tier lists

Explore Tiers →