Skip to content
VecShieldVecShield
← All articles

September 28, 2026

Your Vector Store Is Not Outside Data Regulation

Vector databases are becoming a core part of AI systems.

They power RAG, semantic search, AI assistants, and recommendation systems. But there is a common mistake in how companies think about them:

“We store embeddings, not the original data.”

In many cases, that is simply not true.

Vector stores often contain the original text

A vector database does not have to store only vectors.

Major vector platforms support records that combine an embedding with text, metadata, IDs, and other fields.

This is useful for a simple reason: after finding a vector, the system needs something useful to return.

A RAG system may store a record like this:

id: customer_84913

text:
"John Smith contacted support about his account..."

metadata:
{
  customer_id: "84913",
  source: "support_tickets"
}

embedding:
[0.0321, -0.1928, 0.5531, ...]

The embedding is only one part of the record.

Pinecone supports metadata alongside vectors, including text fields in its own examples.

Qdrant lets each vector point carry a JSON payload. Its documentation shows raw text stored in that payload.

Weaviate stores object properties together with vectors.

Milvus supports scalar and text fields alongside vector fields.

Elasticsearch supports indexes that contain both the original text and its vector representation.

This becomes even more common with hybrid search.

Hybrid search combines semantic vector search with normal keyword search. To perform keyword search, the system often needs the original text or another searchable text field.

So before asking:

“Are embeddings personal data?”

There is a simpler question:

What else is stored next to the embedding?

If the answer includes customer records, emails, employee files, medical notes, account IDs, support tickets, or other personal data, then the compliance issue already exists.

The fact that the data sits inside a vector database does not change what it is.

An embedding is not encryption

Now consider the harder case.

What if the company removes the original text and stores only the embedding?

That still does not mean the data is encrypted.

An embedding is a mathematical representation of information.

Encryption is a security mechanism designed to make information unreadable without the right key.

Those are very different things.

Embedding models try to preserve useful information about the source. That is the whole point. Similar pieces of text should produce related vectors.

Researchers have also shown that information can sometimes be recovered from embeddings.

A 2023 EMNLP paper, Text Embeddings Reveal (Almost) As Much As Text, showed that embedding inversion attacks could reconstruct large amounts of source text under the conditions studied. The researchers also recovered personal information such as names from embeddings of clinical notes.

This does not mean every embedding can be reversed.

It means something simpler:

Turning text into an embedding is not the same as encrypting or anonymizing it.

Data laws do not care what database you use

This matters because privacy laws tend to focus on the data and how it is used, not the database product that stores it.

The GDPR is explicitly technology-neutral.

The European Commission states that GDPR applies regardless of the technology used to process personal data.

It also makes an important distinction between pseudonymous and anonymous data.

Data that has been transformed can still be personal data if a person can be identified using that data together with other information.

Only data that has been made truly anonymous falls outside that part of GDPR.

HIPAA follows a similar idea.

Health data does not become de-identified just because it has been transformed.

HHS defines specific approaches for de-identification, including Safe Harbor and Expert Determination.

“Convert it into an embedding” is not one of them.

California privacy law also uses a broad definition of personal information. It covers information that can reasonably be linked to a consumer or household, including some forms of derived information and inferences.

There is no general:

personal data
    ↓
embedding model
    ↓
unregulated data

exception.

This creates a new compliance problem

Companies already ask these questions about normal databases:

  • What sensitive data do we store?
  • Who can access it?
  • Is it encrypted?
  • Where is it stored?
  • How long do we keep it?
  • Can we audit access?
  • Can we delete it when required?

The same questions need to be asked about vector stores.

But vector systems add new problems.

One document may become 50 chunks.

Those 50 chunks may create 50 embeddings.

They may appear in several collections.

Copies may exist in development and production.

Metadata may contain customer IDs or other sensitive fields.

The original text may still be stored beside each vector.

If a customer asks for their data to be deleted, removing the original document may not be enough.

You also need to know what was created from it.

Start with a simple question

Security and compliance teams do not need to begin by debating the legal status of a 1,536-dimensional embedding.

Start with something much simpler:

What is actually stored in our vector database?

Look at the text.

Look at the metadata.

Look at the IDs.

Look at access controls.

Look at the vectors.

Then trace where all of it came from.

The main point is simple:

Regulation follows the data, not its representation.

Putting personal data into Pinecone, Qdrant, Weaviate, Milvus, Elasticsearch, pgvector, or another vector system does not create a new legal category.

And turning that data into a list of floating-point numbers does not make it encrypted.

Changing the shape of regulated data does not, by itself, make it unregulated.


Sources