Before an answer can be good, its evidence has to be.
Finding sources is easy. Keeping the right data, resolving duplicates, and retrieving relevant evidence across domains is the engineering problem.
Personal project
A question worth building for
“Which sources are about the same thing—and which actually answer this question?”
Source qualityEntity resolutionRelevant retrieval
Illustrative scenario
The spark
I needed a knowledge layer I could inspect and improve.
An embedding index cannot repair poor source selection or the wrong definition of a duplicate. I built Cortex around ingestion, domain rules, reconciliation, and retrieval as separate decisions.
Deduplication
Venue record AVenue record B
One canonical entity
Two descriptions of the same place.
Story clustering
Article AArticle B
One event, distinct evidence
Related perspectives that should stay distinct.
Under the surface
Follow the evidence.
Each stage protects the quality of what reaches the answer.
01 / Ingest
Make the domain part of the contract.
Source selection, normalization, retention, and ranking depend on what the data represents. Domain registries make those policies explicit instead of relying on one generic ingestion rule.
Decision: share infrastructure while preserving domain meaning.
02 / Reconcile
Define what “the same” means.
Two venue records may be duplicates. Two news articles about one event may be useful perspectives. Cortex distinguishes deduplication clusters from story clusters so their membership follows different rules.
Decision: separate identity from relatedness.
03 / Embed
Keep the index rebuildable.
PostgreSQL holds the source records. Embeddings are a retrieval layer with a consistent model and version per corpus. This makes index changes possible without treating vectors as the only copy of the knowledge.
Decision: preserve source data independently of its representation.
04 / Retrieve
Test the confusing cases.
Retrieval and reconciliation need examples that look similar but should not match. Golden evaluation cases and hard negatives help expose those errors while domain metadata informs retrieval.
Decision: evaluate the boundary cases, not just easy matches.
Select a stage to explore the engineering.
The hard part
One generic rule erased 24 news clusters.
A reconciliation worker treated news story groups like duplicate entities. It rejected distinct articles, cleared their memberships, and deleted the empty groups on repeated passes. The failure exposed a mismatch in the data model.
I introduced an explicit cluster role: DEDUP, STORY, or NONE. Generic reconciliation handles deduplication; the news domain owns story clustering. The fix moved the distinction into the contract so ownership is enforceable.
In practice
300K+ knowledge items · Six domains
A growing corpus with domain-specific ingestion and retrieval policies.