Skip to content
aicoolies logo

ParadeDB vs pg_textsearch — Feature-Rich Postgres Search or Fast BM25 at Scale

ParadeDB and pg_textsearch both keep search inside Postgres. ParadeDB is broader for facets, phrase queries, joins, and analytics; Timescale’s pg_textsearch is a custom bm25 index access method, not GIN/tsvector, and its March 2026 benchmarks beat ParadeDB on MS MARCO query latency and throughput.

analyzed by Raşit Akyol April 3, 2026 updated September 5, 2026

Verdict

ParadeDB wins by modernizing search in PostgreSQL, embedding Tantivy's high-performance Lucene-style BM25 search engine directly within the database engine. While native PostgreSQL full-text search (tsvector/tsquery) serves simple keyword matching, it struggles with complex relevance ranking, large-scale tokenization, and multi-facet filtering. ParadeDB eliminates the need to run separate Elasticsearch clusters while seamlessly combining full-text search with pgvector for cutting-edge hybrid retrieval. Our pick: ParadeDB.


Quick Comparison

ParadeDBwinner

Pricing
ParadeDB is free and open-source under AGPL-3.0 as a PostgreSQL extension for search and analytics. Commercial deployments with high availability and proprietary licensing are offered under custom Enterprise Bring Your Own Cloud (BYOC) contracts.
Pricing Model
Open Source
Platforms
PostgreSQL extension on any Postgres platform
Open Source
Yes
Telemetry
Clean
Status
Active
Editorial Pick
—
Last Verified
Aug 26, 2026
Description
ParadeDB brings Elasticsearch-quality full-text search, BM25 ranking, and hybrid vector-keyword search directly into PostgreSQL as native extensions. Backed by a 12 million dollar Series A with over 500,000 Docker deployments, it eliminates the overhead of running separate search infrastructure. Teams get powerful search within their existing Postgres stack without managing additional clusters.

pg_textsearch

Pricing
100% free and open source under the PostgreSQL License ($0 software cost). pg_textsearch is Timescale's high-performance BM25 full-text and hybrid search extension for PostgreSQL with zero software licensing fees.
Pricing Model
Open Source
Platforms
PostgreSQL extension. Works on any platform where PostgreSQL runs. Compatible with pgvector for hybrid search.
Open Source
Yes
Telemetry
Clean
Status
Active
Editorial Pick
—
Last Verified
Sep 6, 2026
Description
pg_textsearch is a PostgreSQL extension from Timescale that adds BM25 relevance-ranked full-text search directly inside Postgres. Using the same ranking algorithm as Elasticsearch and Lucene, it provides search-engine quality results without requiring a separate search cluster — particularly valuable for developers building RAG pipelines on PostgreSQL who want semantic-quality ranking alongside pgvector.

What Sets Them Apart

Both ParadeDB and pg_textsearch keep search inside PostgreSQL, but the original article incorrectly framed pg_textsearch as PostgreSQL’s built-in full-text search layered on GIN and tsvector. Timescale’s README shows a separate extension loaded through shared_preload_libraries, and its SQL install script creates a custom bm25 index access method. The correct split is not native GIN versus Rust search; it is ParadeDB’s broader search feature surface versus pg_textsearch’s purpose-built BM25 index, Block-Max WAND query path, and benchmarked latency focus.

ParadeDB and pg_textsearch at a Glance

BM25 ranking is no longer a ParadeDB-only advantage. pg_textsearch’s public README describes configurable BM25 parameters, the ORDER BY content <@> 'search terms' query form, and CREATE INDEX ... USING bm25 rather than CREATE INDEX ... USING gin. ParadeDB’s pg_search also targets Elasticsearch-style relevance, but the comparison should treat both products as search-grade Postgres extensions instead of contrasting full BM25 against PostgreSQL ts_rank.

The feature surface still favors ParadeDB when an application needs phrase queries, highlighting, tokenizers and token filters, filters, facets, aggregates, joins, or the columnar/analytics layer described in ParadeDB’s own README. Its Tantivy-backed approach stores term positions by default, which is why the Timescale benchmark notes ParadeDB can support phrase queries such as “quick brown fox.” That is a real capability gap, not a performance claim.

pg_textsearch is narrower, but it is not zero-install built-in PostgreSQL search. It requires installing the extension, preloading pg_textsearch, creating the extension in the database, and building a bm25 index. What it buys for that operational step is a Postgres-native extension workflow, expression indexes over JSONB or transformed text, partial indexes for scoped search, partition support, and parallel index builds for large tables.

Feature Breadth, BM25 Ranking, and Hybrid Claims

Hybrid/vector search should be described carefully. ParadeDB’s current README marks vector search and hybrid search as “coming soon,” while its search stack is already strong for BM25, phrase search, facets, filtering, joins, and analytics. pg_textsearch does not provide vector search itself; teams that need semantic retrieval typically pair it with pgvector, pgvectorscale, pgai, or a separate vector system rather than expecting pg_textsearch to solve hybrid ranking alone.

Index maintenance follows different trade-offs than the original text claimed. pg_textsearch is not a GIN index with fastupdate behavior; the extension defines a bm25 index access method and operator class. Timescale’s 1.3.1 SQL install file explicitly creates CREATE ACCESS METHOD bm25 TYPE INDEX HANDLER tp_handler, while ParadeDB uses its own Tantivy-backed index structures. Write amplification, segment merging, phrase support, and index size therefore have to be evaluated from each extension’s own index design, not from PostgreSQL GIN defaults.

Language analysis is one place where pg_textsearch intentionally reuses PostgreSQL strengths. The README says it works with PostgreSQL text search configurations such as English, French, and German, while also supporting expression indexes and multi-column search. ParadeDB provides a richer tokenizer and filter configuration surface for teams that want search-engine-style analysis controls. The right choice depends on whether you prefer Postgres text configuration compatibility or deeper search analyzer tuning.

Indexing, Language Analysis, and Benchmark Evidence

The benchmark story should be reversed from the original article. Timescale’s public pg_textsearch vs ParadeDB comparison, attributed to the pg_textsearch benchmark dashboard and commonly circulated by Todd J. Green, reports pg_textsearch 3.1x faster overall query throughput on the 8.8M-passage MS MARCO v1 run, with p50 latency faster across all 1-token through 8+ token buckets. ParadeDB still built the index faster in that run, 140.1 seconds versus 233.5 seconds, and retained phrase-query and broader feature advantages.

At larger scale, the same source reports pg_textsearch ahead on query performance rather than degrading behind ParadeDB. On the 138M-passage MS MARCO v2 experiment, pg_textsearch shows 2.3x faster weighted p50 query latency and 4.7x higher concurrent throughput with 16 clients, while ParadeDB builds the index 1.9x faster and keeps better p95 latency on some longer-query buckets. That is a nuanced trade-off: pg_textsearch appears stronger for BM25 query latency and concurrent throughput at scale; ParadeDB remains stronger for index build speed and richer search features.

The Bottom Line


FAQ

What architectural differences make ParadeDB's BM25 search significantly faster and more accurate than native PostgreSQL tsvector with GIN indexes?

ParadeDB integrates Tantivy (a Rust search engine) directly into PostgreSQL via its pg_search extension, implementing true Okapi BM25 scoring with term frequency saturation and document length normalization. Native PostgreSQL tsvector with GIN indexes uses a simplified ts_rank scoring function without probabilistic saturation or corpus-wide IDF normalization. Architecturally, Tantivy maintains columnar postings lists and block-compressed inverted indexes, outperforming GIN indexes by 5x–20x on multi-million row text datasets.

How do write amplification, VACUUM overhead, and indexing concurrency differ between ParadeDB's Tantivy segments and PostgreSQL GIN indexes?

PostgreSQL GIN indexes suffer from high write amplification and lock contention during frequent INSERT and UPDATE workloads because posting lists must be locked and rewritten within the PostgreSQL buffer pool, generating substantial WAL volume and requiring aggressive VACUUM worker cycles. ParadeDB's pg_search uses Tantivy's segment-based LSM architecture: incoming writes are buffered in memory and flushed as immutable index segments that are merged asynchronously in background threads, isolating text indexing from table lock contention.

How does ParadeDB handle hybrid search (BM25 + vector similarity) compared to combining tsvector with pgvector via custom SQL Reciprocal Rank Fusion (RRF)?

In native PostgreSQL, hybrid search requires running two separate index scans—GIN for tsvector and HNSW for pgvector—followed by a custom SQL CTE implementing Reciprocal Rank Fusion (RRF). ParadeDB provides native hybrid search primitives directly within SQL via pg_search, executing Tantivy BM25 sparse retrieval and dense vector indexing in a single fused execution plan with built-in RRF at the engine level.

What operational trade-offs and deployment constraints should be considered when choosing ParadeDB over native PostgreSQL full-text search?

Native PostgreSQL full-text search (tsvector) is built into core PostgreSQL and available across all managed cloud database services (AWS RDS, Aurora, Cloud SQL, Supabase). ParadeDB requires compiling or loading external C/Rust shared libraries (pg_search), limiting deployment to self-hosted PostgreSQL instances, ParadeDB Cloud, or custom Kubernetes containers. For modest text volumes (<500K records), native tsvector is simpler; for enterprise-scale search (>1M documents, faceting, BM25 relevance), ParadeDB is superior.

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.