Workstate's search stack: Postgres, pgvector and Voyage, with no search cluster

· #workstate #postgres #aws #rag #architecture · 12 min read

Workstate is a shared memory and decision record for people and their AI agents. It indexes a team’s code, documents and tickets, and keeps a ledger of decisions and incidents that agents search before they act and write to when they finish. Agents reach it over MCP (Model Context Protocol), an open standard for connecting AI tools to outside services.

Our first prototype ran on a single consumer-grade GPU. This post is about what Workstate will run on in production, and why we chose one Postgres database over a search cluster. It’s a design: we haven’t built the production version yet.

What a search has to do

A search takes a question from an agent and returns the few passages most likely to answer it, from that team’s data and nobody else’s.

A question goes two ways: it is embedded and matched against stored vectors, and it is run as a keyword search. The two result lists are merged, reranked, and the top eight passages go back to the agent. Every step is limited to one namespace.
One search, from question to answer. Every step stays inside one namespace.

Four stages do the work:

The reranker can only reorder what the first stages found. A passage they miss never reaches the agent, and the agent never knows it was there. So the embedding model sets the ceiling.

Everything is scoped to a namespace, the unit that separates one team’s data from another’s.

The prototype

Our first prototype ran open models on one consumer-grade GPU: Qwen3-Embedding (8 billion parameters) for embeddings and bge-reranker-v2-m3 for reranking. Vectors lived in Qdrant, a dedicated vector database, and accounts, sources and the ledger lived in Postgres. It proved the idea, but it couldn’t carry over to production:

What we weighed

Price was one of four questions, and not the one that decided it:

Most of the search quality comes from the models, and they’re the same in every option: Voyage’s, called directly or, in Atlas, through MongoDB. What changes is where the vectors live. We sized each option against a corpus like our prototype’s test data: about 193,000 chunks, roughly 100 million tokens, stored as 1,024-number vectors in about 1.5 GB. Prices are list prices as of October 2026, and the table says which figures are estimates.

Option Per month Customer data stored Fit with our tools
Qdrant Cloud ~$30–115 for a 2 GB node (outside estimates: Qdrant doesn’t publish its rates) At Qdrant A second database next to Postgres
MongoDB Atlas, which owns Voyage ~$57 (M10) to ~$146 (M20), with embedding and reranking built in At MongoDB A second database, and Postgres stays for the app
Qdrant on an EC2 instance (an AWS virtual machine) ~$25–35 (approximate) In our AWS account A second database that we patch, back up and fail over ourselves
Amazon OpenSearch Service ~$90+ for instances, ~$175+ for Serverless (approximate) In our AWS account A search cluster to run beside Postgres
S3 Vectors, vector storage in Amazon S3 Under $1 In our AWS account A second store, with no transactions and about 100 ms per query
Postgres + pgvector on RDS $0–50 on top of the Postgres we need anyway In our AWS account The same database as the app and the ledger

The models cost the same whichever store holds the vectors. Embedding 100 million tokens costs $12 with voyage-4-large and $2 with voyage-4-lite. That’s once per customer, plus small top-ups as their content changes, and Voyage gives each account its first 200 million tokens free for each model. Each search costs about a tenth of a cent, almost all of it reranking: 50 candidates come to about 20,000 tokens. On cost alone, S3 Vectors would have won. Enterprise adoption and cohesion pointed at Postgres.

The decision

Against the four questions:

  1. Enterprise adoption. Customer data stays in our AWS account, in a database security teams already know how to assess. The one company the design adds to a security review is Voyage, and What we gave up covers that.
  2. Cohesion. One database holds the app’s data, the ledger and the vectors: one backup, one set of credentials, one access model. A ledger entry and its vectors commit in the same transaction, so they can’t drift apart, and the reconcile job goes away. Keyword search comes from the same database: embeddings are bad at exact names, such as a class name, a config key or a line from a stack trace, and Postgres full-text search is good at them.
  3. Operations. RDS handles patching and backups, and Multi-AZ (a standby in a second availability zone) handles failover.
  4. Cost. Only S3 Vectors is cheaper, and it would bring back a second store, with no transactions and about 100 ms a query.

Why not Amazon’s own models?

Amazon Bedrock has its own embedding model, Titan Text Embeddings V2, at $0.02 per million tokens: the same price as voyage-4-lite. It would keep the models inside AWS too, which would make enterprise reviews simpler. We chose Voyage anyway, for quality and cohesion. Titan has no separate mode for questions and documents, which Voyage’s models use to tune each embedding for its side of the search. And Titan’s vectors can’t share an index with any other model, while Voyage’s 4 series puts its large, lite and nano models in one embedding space. Free plans, paid plans and an outage fallback can all share one index. Reranking isn’t cheaper on Bedrock either: Cohere Rerank 3.5 costs about $0.002 a search, twice Voyage’s.

Making pgvector work for many teams

We haven’t built this storage layer yet. The schema below is the plan:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE chunks (
  id          uuid PRIMARY KEY,          -- stable for (namespace, path, chunk_index)
  namespace   text NOT NULL,
  source_id   uuid,
  repo        text NOT NULL,
  path        text NOT NULL,
  chunk_index int  NOT NULL,
  body        text NOT NULL,
  embed_model text NOT NULL,             -- which model made the vector
  embedding   halfvec(1024) NOT NULL,    -- 16-bit floats: half the memory
  search      tsvector GENERATED ALWAYS AS (to_tsvector('simple', body)) STORED
);

CREATE INDEX chunks_embedding ON chunks USING hnsw (embedding halfvec_cosine_ops);
CREATE INDEX chunks_search    ON chunks USING gin (search);
CREATE INDEX chunks_scope     ON chunks (namespace, repo);

And the vector half of a search:

SET hnsw.ef_search = 100;                 -- candidates per pass; the default, 40, is fewer than we want
SET hnsw.iterative_scan = relaxed_order;  -- keep scanning until enough rows pass the filter

SELECT id, path, body
FROM chunks
WHERE namespace = $1
ORDER BY embedding <=> $2::halfvec        -- cosine distance to the question's vector
LIMIT 50;                                 -- possibly a little out of order: the reranker sorts them

What makes this work for many teams in one table:

Free and paid plans on different models

Vectors are a model’s output. There’s no converting voyage-4-lite vectors into voyage-4-large ones: the only way to get large-model vectors is to embed the text again with the large model. What makes plans cheap anyway is that the Voyage 4 models share one embedding space. In Voyage’s words, “All embeddings created with the 4 series are compatible with each other.” So one namespace can hold a mix of both.

A free namespace has all its chunks embedded with voyage-4-lite. On upgrade, a background job re-embeds each chunk with voyage-4-large. For a while the namespace holds a mix of both and search keeps working; when the job finishes, every chunk uses voyage-4-large.
Upgrading a namespace in place. Search keeps working throughout.

The embed_model column is what lets the job find the chunks still on lite. Reranking can follow the plan too, with rerank-3-lite for free and rerank-3 for paid, and that needs no re-indexing at all.

Two rules keep this safe. Every model writes vectors of the same size, 1,024 numbers. And models only ever mix within the Voyage 4 family. Vectors from different families, say Voyage and Titan, don’t cause an error when mixed: search just returns worse results, with nothing to notice. Our current code already guards against this, by refusing writes from any model but the one that built the index. The new store keeps that guard, widened to the Voyage 4 family at 1,024 dimensions and no further.

How much worse is lite? Across the 29 datasets in Voyage’s general retrieval benchmark, large beats lite by 4.8% on average, scored by NDCG@10 (normalised discounted cumulative gain at 10, a standard score for how well the top 10 results are ordered). Reranking should narrow the gap, since it reorders whatever the first stage finds, but we’ll measure that on our own data rather than assume it.

When Voyage is down

Every search embeds its question through Voyage’s API, so if the API is down, search is down.

The fix is voyage-4-nano, which Voyage released as open weights under the Apache 2.0 licence. It has about 340 million parameters and shares the larger models’ embedding space. It can run inside our own search service and embed questions when the API fails. Or it can embed every question, which saves a network round trip on each search: Voyage’s own advice is to embed documents once with the large model and queries with a smaller one. How fast it runs without a GPU is still to be measured.

If reranking fails, search will fall back to vector order and say so in the response: worse results, but still results.

What we gave up

What’s next

  1. Build the pgvector storage layer.
  2. Run our own retrieval test: the same questions through the same reranker, with voyage-4-large, voyage-4-lite, voyage-code-4, voyage-4-nano and Amazon Titan. It measures how often the right file is among the 50 candidates, and how often it’s in the top 8 after reranking. The test already exists: on the prototype, search put the right file in the top 10 for 71% of 28 test questions, against 57% for grep patterns written by someone who already knew the answer. One result would change the plan: Voyage doesn’t say whether voyage-code-4 shares the 4 series’ embedding space, so if it wins on code, free and paid plans might not be able to share an index.
  3. Measure the real token counts on the first production index. Voyage reports them with every request.

Then a follow-up post with the numbers.

Sources