A private, local-first coding workflow that runs Ollama on the developer's hardware, uses Cline for editor-guided implementation and Aider for Git-native terminal repair, reproduces dependencies and tests in Docker, and coordinates every lane in tmux. Human review remains the merge boundary.
curated / stacks
Stacks
Curated, opinionated tool combinations for specific use cases, roles, and budgets.
131 stacks published
showing 48 of 131 stacks
A production-focused SaaS workflow that moves AI-assisted code from Cursor through isolated database changes, identity, CI, preview deployments, and release monitoring. Supabase is the production backend, Neon handles disposable migration branches, Clerk owns identity, GitHub Actions gates changes, Vercel delivers them, and Sentry closes the feedback loop.
Build agent-authored pull requests behind an evidence-first quality gate. Claude Code proposes the change, Playwright and reviewdog turn tests and diagnostics into reviewable signals, Qodo provides an independent PR review, GitHub Actions enforces the required checks, and Sentry connects post-merge regressions to the release—while human reviewers retain merge authority.
This stack turns scarce or privacy-restricted data into a fine-tuned open model without ever training on raw records. You generate synthetic training data with Gretel and SDV, validate and clean it with Cleanlab, fine-tune an open LLM with LLaMA-Factory, and version every dataset and model with DVC so the entire run is reproducible. The result is a modular, auditable path from synthetic data to a domain-tuned model with full data lineage.
Delegate bounded GitHub issues to autonomous coding agents while keeping acceptance tests, review evidence, CI, and human merge authority explicit.
Run Claude Code and Codex in isolated parallel lanes, then turn their output into reviewable pull requests through deterministic review and CI gates.
Build a durable context pipeline for AI coding with repository memory, current documentation, portable context packs, and CI-enforced rules.
A continuous evaluation workflow that combines Promptfoo matrices, DeepEval assertions, RAGAS retrieval metrics, Langfuse production traces, and garak security probes across pull requests and releases.
A shift-left red-team pipeline using garak for broad probes, PyRIT for adaptive attacks, DeepTeam for risk-driven scenarios, Agentic Security for CI integration, and iFixAi for continuous agent-safety scanning.
A reliability workflow for tool-using agents that combines AgentOps sessions, DeepEval behavioral checks, Langfuse trace-linked scores, Inspect AI benchmarks, and Sentrial runtime detection.
A vendor-neutral judge pipeline using DeepEval for rubrics, Opik for datasets and experiments, Langfuse for production scoring, Arize Phoenix for analysis, and LangSmith as an optional managed operations layer.
A production-minded RAG evaluation workflow that combines RAGAS and DeepEval metrics with Langfuse traces, Arize Phoenix retrieval analysis, and GitHub Actions release gates.
A code-first data path from source connectors to governed RAG and training datasets with Meltano, Polars, Deep Lake, Great Expectations, and DVC.
A staged human-feedback and training-data workflow using Label Studio, Argilla, Cleanlab, Snorkel AI, and Labelbox for annotation, preference curation, quality review, weak supervision, and governed delivery.
A reproducible ML data lifecycle combining DVC, lakeFS, MinIO, Great Expectations, and Weights & Biases for versioned objects, quality gates, pipelines, and experiment lineage.
A privacy-aware document ingestion workflow that turns PDFs and multimodal files into governed, searchable AI data with OpenDataLoader PDF, Dolphin, Pixeltable, Deep Lake, and Presidio.
Coordinate OpenHands, SWE-Agent, Aider, Ollama, and tmux for an issue-to-patch workflow with isolated sandboxes, deterministic tests, human approval, and explicit escalation.
Compare Pi, Crush, Amp, and OpenCode in a tmux workflow that separates open-source and hosted lanes, provider credentials, implementation, and review.
Combine Goose, OpenCode, Ollama, DesktopCommanderMCP, and tmux for a local-first coding workflow with explicit model, filesystem, process, and network boundaries.
Run Claude Code, Codex, Gemini CLI, and OpenCode in isolated tmux panes, then compare their patches through one explicit worktree, test, and review contract.
A layered, documentation-backed MCP operating model for teams that need controlled discovery, federation, security scanning, protocol validation, custom server development, and containerized runtime isolation before agents receive production access.
A worktree-first command center for running multiple coding agents on one repository: Conductor.build isolates workspaces, Vibe Kanban tracks lanes, Claude Code implements, Checkpoints enables rollback, Cursor reviews diffs, and GitHub Actions decides what can merge.
A governed AI coding delivery loop that starts with Taskmaster planning, refreshes implementation context through Context7, moves repository work through GitHub MCP, codes with Claude Code, validates browser paths through Playwright MCP, and audits the session with Spotlight before merge.
A self-hostable agent evaluation and observability stack for teams replacing ad hoc LangSmith-style dashboards with open dev loops: Judgeval scores behavior, Laminar traces and tests workflows, TraceRoot debugs failures, Prompt Flow organizes eval runs, and OpenAI Agents SDK provides a runnable agent surface.
A production RAG retrieval stack for teams that need more than a demo chatbot: LlamaIndex coordinates indexing and retrieval, RAGFlow and RAG-Anything handle difficult multimodal documents, and Judgeval keeps retrieval quality measurable before and after launch.
A production LLM evaluation stack should catch regressions before release, probe security failures, and close the loop with real traces and user feedback. This stack combines Promptfoo for CI gates, DeepEval/OpenAI Evals for metric-heavy test suites, and Langfuse or Helicone for observability and production datasets.
A stack for teams adopting multi-agent autonomous development. Covers daemon-based issue processing with Symphony, parallel agent fleet management with Agent Orchestrator, Claude-native swarm intelligence with ruflo, and lightweight multi-engine looping with ralphy. All tools are open source and composable for different team sizes and workflows.
A modern self-hosted infrastructure stack for teams and homelabbers who want secure remote access, container management, SSH workflows, and lightweight authentication without enterprise complexity or subscription costs. All tools are open source and deployable on minimal hardware.
A production stack for giving AI agents the ability to browse, interact with, and extract data from the web. Combines Browserless for headless browser infrastructure, Page Agent for in-page AI interaction, and Nango for API integrations — enabling AI agents to operate across both web interfaces and APIs.
A complete open-source stack for feature flags, A/B testing, and progressive rollouts using GrowthBook for experimentation, Flagsmith for simple feature toggles, and supporting infrastructure. All tools are self-hostable and replace paid platforms like LaunchDarkly and Statsig at zero software cost.
A complete open-source toolkit for working with Chinese AI models covering fine-tuning, inference serving, agent development, and voice synthesis from the leading Chinese AI research labs.
A complete toolkit for deploying and managing Model Context Protocol servers at production scale, from browser automation through server management to gateway routing and agent orchestration.
A complete API documentation toolkit combining modern reference rendering, developer portal generation, and property-based API testing for teams that treat their API docs as a product.
A curated frontend toolkit combining headless primitives, zero-runtime styling, animated components, and number transitions for building polished React applications with full design control.
A comprehensive security testing toolkit that combines AI-powered vulnerability discovery, LLM security assessment, API fuzzing, and supply chain analysis to protect modern applications across their entire attack surface.
A self-hosted identity infrastructure built entirely on open-source components for organizations that need complete control over authentication data, protocols, and user management without vendor lock-in or per-user pricing.
A comprehensive FinOps toolkit combining Kubernetes-specific cost allocation with multi-cloud visibility, autonomous optimization, and infrastructure governance to control cloud spending across the organization.
A complete open-source observability toolkit for Kubernetes that combines eBPF-powered monitoring, network flow visibility, alert enrichment, and ChatOps integration without requiring application instrumentation changes.
A fully open-source toolkit for AI-assisted software development, combining autonomous coding agents with code review automation and intelligent code completion across VS Code and JetBrains IDEs.
A production-ready toolkit for fine-tuning large language models from data preparation through deployment, combining comprehensive training orchestration with speed optimization and distributed compute.
A Postgres-native stack that adds Elasticsearch-quality search, RAG-ready vector retrieval, and self-hosted documentation without introducing external infrastructure. Everything runs within or alongside your existing PostgreSQL deployment.
A complete stack for running multiple AI coding agents in parallel with structured workflows, workspace isolation, session traceability, and cross-model code review. Covers the full cycle from task planning through code delivery and audit.
A stack for teams that need reproducible AI training pipelines with full dataset version control. Combines Dolt's Git-for-data SQL database with OpenBB for financial data ingestion and SWE-bench for agent evaluation, providing branching, diffing, and audit trails across the entire data lifecycle.
A complete local AI serving and development stack optimized for AMD Ryzen AI hardware. Combines Lemonade's NPU-accelerated inference with Open WebUI's interface and Ollama as a CPU fallback, covering text, image, and speech modalities entirely on-device with zero cloud dependencies.
Give your AI agents full web interaction: Firecrawl for web data extraction, Browser Use for autonomous browsing, Stagehand for structured browser automation, and Hyperbrowser for cloud browser infrastructure.
Build AI-powered applications entirely in TypeScript: Mastra for agent framework, Vercel AI SDK for streaming UI, Prisma for type-safe database access, and Supabase for auth, storage, and real-time backend.
LLM Observability Stack
variesMonitor, trace, and optimize your LLM applications: Langfuse for deep tracing and evaluation, Helicone for request logging and analytics, Portkey for AI gateway routing, and Sentry for error tracking across your full stack.
Open-Source DevOps Stack
variesBuild production infrastructure with open-source tools: Terraform for infrastructure provisioning, Kubernetes for container orchestration, Grafana for dashboards, and SigNoz for full-stack observability.