Skip to content
aicoolies logo
Synthetic Data Vault logo

Synthetic Data Vault

Open-source library for generating synthetic tabular data

Synthetic Data Vault (SDV) is an MIT-backed open-source Python library for generating synthetic tabular, relational, and time-series data. It learns statistical patterns from real datasets and produces synthetic versions that preserve distributions, correlations, and referential integrity. Supports single-table, multi-table, and sequential data with built-in privacy and quality metrics.

About Synthetic Data Vault

The Synthetic Data Vault is a Python library originating from MIT research that generates synthetic datasets preserving the statistical properties of real data. SDV learns joint distributions, correlations, and constraints using generative models including Gaussian copulas, CTGAN, and TVAE, producing synthetic samples that maintain these properties while containing no real records. The library handles single tables, multi-table relational databases with foreign key relationships, and sequential time-series data.

For relational data, SDV's multi-table synthesizers preserve referential integrity across related tables, generating consistent synthetic databases where parent-child relationships and cardinality distributions match the original structure. Built-in quality metrics compare synthetic data against real data on column distributions, pairwise correlations, and boundary adherence. Privacy metrics evaluate disclosure risk to ensure generated data cannot re-identify individuals from the source dataset.

SDV is open-source under MIT license and backed by DataCebo, which offers commercial extensions. The library integrates naturally into Python ML workflows, producing pandas DataFrames for model training, testing, and analysis. For teams needing realistic test data, privacy-safe datasets for sharing, or augmented training data for ML models, SDV provides the most accessible open-source entry point to synthetic data generation.

Pricing & Platform Specs

Pricing Summary

Free self-hosted Community edition under Business Source License (BSL 1.1) for local modeling ($0); DataCebo Enterprise starts at $500/user/mo plus optional add-on feature bundles from $250/mo.

full pricing breakdown →

Supported Platforms

Python library — any environment with pandas

Explore categories, tags & use cases

Categories

Synthetic data generation platform for privacy and ML

Gretel is a synthetic data platform that generates realistic, privacy-preserving datasets for ML training, testing, and data sharing. It supports tabular, text, and time-series data with configurable privacy guarantees including differential privacy. Features include data augmentation for imbalanced datasets, PII detection and anonymization, and API/SDK access for pipeline integration with BigQuery, Snowflake, and Databricks.

freemium

Entity-based synthetic data generation for enterprise

K2view is an enterprise data platform that generates synthetic data using an entity-based micro-database architecture. It ensures referential integrity across complex multi-relational datasets by treating each business entity as a self-contained unit. Used for privacy-compliant test data generation, data masking, and AI training data creation in financial services, telecom, and healthcare industries.

paid

Side-by-Side Comparisons

Gretel logo
Gretel
vs
Synthetic Data Vault logo
Synthetic Data Vault

Gretel vs Synthetic Data Vault — Cloud Synthetic Data Platform or Local Python Library

Gretel and Synthetic Data Vault both generate synthetic data, but they fit different teams. Gretel is a commercial platform for privacy-preserving data generation, API workflows, and enterprise data operations. Synthetic Data Vault is an open-source Python library for local, reproducible synthetic tabular data generation. Choose Gretel for managed workflows and governance; choose SDV when developers need an open, scriptable library they can run and inspect themselves.

GretelSynthetic Data Vault

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is Synthetic Data Vault?

Synthetic Data Vault (SDV) is an MIT-backed open-source Python library for generating synthetic tabular, relational, and time-series data. It learns statistical patterns from real datasets and produces synthetic versions that preserve distributions, correlations, and referential integrity. Supports single-table, multi-table, and sequential data with built-in privacy and quality metrics.

Is Synthetic Data Vault free?

Synthetic Data Vault offers a free tier alongside paid plans. Free self-hosted Community edition under Business Source License (BSL 1.1) for local modeling ($0); DataCebo Enterprise starts at $500/user/mo plus optional add-on feature bundles from $250/mo.

Is Synthetic Data Vault open source?

Yes — Synthetic Data Vault is open source.

Is Synthetic Data Vault still maintained?

Yes — Synthetic Data Vault is active. Its listing was last verified on September 6, 2026.

What are the best Synthetic Data Vault alternatives?

The first editor-selected Synthetic Data Vault alternatives are Gretel, K2view.