← Knowledge

Public EcoRouter Knowledge · AI-generated content ·

What am I missing about vector databases?

The question

What am I missing about vector databases? Couldn't you just store embeddings in Postgres with pgvector?

Efficient

You have hit on a very practical truth: for many use cases, you absolutely can just use Postgres with pgvector.

When people ask "what am I missing," they are usually wondering why specialized vector databases (like Pinecone, Milvus, or Weaviate) exist if a general-purpose database can do the job. Here is the trade-off you are navigating:

The case for pgvector If you already use Postgres, pgvector is often the superior choice. It allows you to keep your metadata and your embeddings in the same place, use standard ACID transactions, and avoid the operational overhead of managing a second database. For many applications with datasets under a few million vectors, the performance difference is negligible.

What you "miss" (The specialized database advantage) Specialized vector databases were built for scale and specific architectural needs that Postgres may struggle with as you grow:

  1. Massive Scale: Once you reach tens or hundreds of millions of vectors, the memory requirements for indexing (like HNSW) can put immense pressure on a Postgres instance. Dedicated vector databases often offer better "sharding" and horizontal scaling capabilities.
  2. Advanced Indexing: Specialized databases often support a wider array of indexing algorithms and optimizations (like disk-based indices or specific quantization techniques) that may be more mature or easier to tune for massive datasets.
  3. Advanced Filtering/Hybrid Search: While pgvector has improved significantly in handling filtered vector search, dedicated systems are often optimized to combine keyword search (BM25) and vector search (hybrid search) more fluidly, often including built-in reranking stages.
  4. Operational Separation: In large engineering teams, you might want to separate the operational database (Postgres) from the "AI index" database to ensure that a massive vector search query doesn’t impact the performance of your primary application transactions.

The verdict You aren't missing a "better" solution; you are looking at a "different scale" solution.

If your data fits comfortably in Postgres and you don't have extremely low-latency requirements for massive search volumes, pgvector is likely the most efficient, cost-effective, and simplest architecture. Most engineers are better off starting with pgvector and only migrating if they hit a specific performance ceiling they cannot tune away.

Answering this used an estimated 90% less than EcoRouter’s configured frontier baseline

Routed to
Efficient
Tokens
379 in / 496 out
Cost
$0.0008
Baseline
$0.0086

Figures recorded by EcoRouter when this answer was generated, and fixed at that moment. Cost comparisons are against a configured reference model, not a measurement of electricity, carbon or water.

Ask about this

Ask EcoRouter a follow-up using this Knowledge as context. Nothing becomes public unless you choose to publish it.

Ask a follow-up
0 views

Comments

No comments yet.