aicoolies logo
Khoj logo
Khoj logo

Khoj

Open-source AI second brain with deep research and RAG

freemiumopen sourceupdated Aug 16, 2026

Khoj is an open-source personal AI app that serves as a self-hostable second brain. It connects to your documents — PDFs, Markdown, Notion, Word — and uses RAG to answer questions grounded in your knowledge base. Supports any local or cloud LLM including Llama, Claude, GPT, and Gemini. Features custom agents, scheduled automations, deep research mode, semantic search, and Obsidian, Emacs, and WhatsApp integrations. Over 33,000 GitHub stars, YC-backed.

Khoj occupies a unique position in the AI tools landscape as a personal knowledge companion that bridges your private documents with the reasoning capabilities of large language models. Unlike generic chat interfaces that rely solely on training data, Khoj indexes your files using semantic embeddings and retrieves relevant context before generating answers, dramatically reducing hallucinations when working with your own content. The system supports a wide range of document formats including PDFs, Markdown files, Notion pages, Word documents, and org-mode files, making it compatible with most knowledge workers' existing workflows without requiring format conversion.

The platform's agent system allows users to create specialized AI assistants with custom knowledge bases, personas, and tool access. A research agent might have access to academic papers and web search, while a project assistant might draw only from specific repository documentation. Khoj's scheduled automation feature enables recurring research tasks that deliver results as personal newsletters or notifications. The deep research mode performs multi-step investigation across both local documents and web sources, synthesizing findings into comprehensive reports.

Khoj scales from a fully offline, on-device deployment using local models through Ollama to a cloud-hosted enterprise installation with team management and SSO. The self-hosted option runs via Docker or pip with complete data sovereignty, while the hosted version at app.khoj.dev offers a free tier for individual users. With over 33,000 GitHub stars, Y Combinator backing, native Obsidian and Emacs plugins, and support for image generation and voice interaction, Khoj has established itself as the leading open-source alternative to commercial AI assistants for knowledge-intensive work.

Pricing

Free self-hosted; cloud free tier available; paid plans for teams

Platforms

Web, Desktop, Obsidian, Emacs, WhatsApp — Docker or pip self-host

Categories

Tags

Use Cases

Open WebUI logo

Open WebUI

Self-hosted AI platform with ChatGPT-like interface for local and cloud LLMs.

Extensible, self-hosted AI platform with 290M+ Docker pulls and 124K+ GitHub stars. Supports Ollama, OpenAI-compatible APIs, and any Chat Completions backend. Features built-in RAG, multi-user RBAC, voice/video calls, Python function workspace, model builder, and web browsing. Runs entirely offline with enterprise features including SSO and audit logging.

free
AnythingLLM logo

AnythingLLM

All-in-one self-hosted AI app with RAG, agents, and multi-user support

AnythingLLM is an open-source, privacy-first AI application that turns any document into an interactive knowledge base. It bundles document ingestion, vector storage (built-in LanceDB), RAG pipelines, AI agents, and multi-user access into a single deployable package. Supports 30+ LLM providers including OpenAI, Anthropic, Ollama, and local models. With 62K+ GitHub stars and MIT license, it runs as a desktop app or Docker container with zero configuration required out of the box.

freemiumOpen Source
LibreChat logo

LibreChat

Self-hosted multi-model AI chat platform

LibreChat is an open-source ChatGPT-like interface with 35K+ GitHub stars supporting multiple AI providers in a single self-hosted platform. Connect OpenAI, Anthropic, Google, Mistral, local models via Ollama, and custom endpoints simultaneously. Features conversation branching, file uploads, code interpreter, plugins, presets, multi-user support with RBAC, and LDAP/SSO authentication. Privacy-focused alternative to commercial AI chat services with full data ownership.

Open Source
Ollama logo

Ollama

Run LLMs locally with one command

Tool for running large language models locally on your machine with a simple CLI interface. Download and run Llama 3, Mistral, Gemma, Phi, Code Llama, and dozens of other open-source models with a single command. Features model management, GPU acceleration (NVIDIA/AMD/Apple Silicon), OpenAI-compatible API server, Modelfile for customization, and multi-model switching. Ideal for offline AI development, privacy-sensitive use cases, and local testing. 120K+ GitHub stars.

Open Source
Hyprnote logo

Hyprnote

Local-first AI notepad for meetings and voice notes

Hyprnote is a local-first AI notepad designed for capturing and processing meeting notes and voice recordings. It runs entirely on-device for privacy, transcribes audio using local models, and generates structured summaries, action items, and follow-ups. Built with Rust and Tauri for native desktop performance. Over 8,000 GitHub stars with strong privacy-focused community adoption.

Open Source

Related Tools

computed discovery: shared active categories · kept separate from editor-verified Alternatives

KTransformers parent kvcache-ai logo

KTransformers

Heterogeneous CPU-GPU inference and SFT for large MoE models

Open-source framework for running and fine-tuning large Mixture-of-Experts models with heterogeneous CPU-GPU execution, optimized kernels, limited VRAM and SGLang or LLaMA-Factory integrations.

Open Source
vLLM Production Stack parent vLLM logo

vLLM Production Stack

Official Kubernetes and Helm reference stack built on the vLLM inference engine

Official vLLM reference implementation for scaling the existing inference engine on Kubernetes with Helm, request routing, KV-cache offload, autoscaling and Prometheus/Grafana observability.

Open Source
Dynamo logo

NVIDIA Dynamo

Distributed inference orchestration above vLLM, SGLang and TensorRT-LLM

Open-source, datacenter-scale orchestration layer that coordinates vLLM, SGLang and TensorRT-LLM across nodes with disaggregated serving, KV-aware routing, multi-tier cache management and automatic scaling.

Open Source
GPUStack logo

GPUStack

Open-source GPU control plane for scalable AI model serving

Open-source GPU cluster manager that configures vLLM, SGLang, TensorRT-LLM or custom engines, serves models through compatible APIs, and provisions SSH-accessible GPU instances across on-premises, Kubernetes and cloud environments.

Open Source
Mooncake logo

Mooncake

Disaggregated KV cache storage and transfer for LLM serving

Open-source infrastructure for disaggregated LLM serving that pools KV caches across prefill and decode workers, with high-performance transfer, distributed storage and integrations for vLLM and SGLang.

Open Source
LMCache logo

LMCache

Reusable KV cache infrastructure for scalable LLM inference

Open-source KV cache management layer that persists, offloads and reuses model key-value caches across requests and serving engines to reduce repeated prefill work and improve inference throughput.

Open Source

FAQ

What is Khoj?

Khoj is an open-source personal AI app that serves as a self-hostable second brain. It connects to your documents — PDFs, Markdown, Notion, Word — and uses RAG to answer questions grounded in your knowledge base. Supports any local or cloud LLM including Llama, Claude, GPT, and Gemini. Features custom agents, scheduled automations, deep research mode, semantic search, and Obsidian, Emacs, and WhatsApp integrations. Over 33,000 GitHub stars, YC-backed.

Is Khoj free?

Khoj offers a free tier alongside paid plans. Free self-hosted; cloud free tier available; paid plans for teams

Is Khoj open source?

Yes — Khoj is open source.

What are the best Khoj alternatives?

The top editor-verified Khoj alternatives are Open WebUI, AnythingLLM, LibreChat, and more.