← All projects

Cortex / Knowledge infrastructure

Before an answer can be good, its evidence has to be.

Finding sources is easy. Keeping the right data, resolving duplicates, and retrieving relevant evidence across domains is the engineering problem.

Personal project

A question worth building for
“Which sources are about the same thing—and which actually answer this question?”
Source qualityEntity resolutionRelevant retrieval

Illustrative scenario

The spark

I needed a knowledge layer I could inspect and improve.

An embedding index cannot repair poor source selection or the wrong definition of a duplicate. I built Cortex around ingestion, domain rules, reconciliation, and retrieval as separate decisions.

Deduplication
Venue record AVenue record B

One canonical entity

Two descriptions of the same place.

Story clustering
Article AArticle B

One event, distinct evidence

Related perspectives that should stay distinct.

Under the surface

Follow the evidence.

Each stage protects the quality of what reaches the answer.

01 / Ingest

Make the domain part of the contract.

Source selection, normalization, retention, and ranking depend on what the data represents. Domain registries make those policies explicit instead of relying on one generic ingestion rule.

Decision: share infrastructure while preserving domain meaning.

02 / Reconcile

Define what “the same” means.

Two venue records may be duplicates. Two news articles about one event may be useful perspectives. Cortex distinguishes deduplication clusters from story clusters so their membership follows different rules.

Decision: separate identity from relatedness.

03 / Embed

Keep the index rebuildable.

PostgreSQL holds the source records. Embeddings are a retrieval layer with a consistent model and version per corpus. This makes index changes possible without treating vectors as the only copy of the knowledge.

Decision: preserve source data independently of its representation.

04 / Retrieve

Test the confusing cases.

Retrieval and reconciliation need examples that look similar but should not match. Golden evaluation cases and hard negatives help expose those errors while domain metadata informs retrieval.

Decision: evaluate the boundary cases, not just easy matches.

The hard part

One generic rule erased 24 news clusters.

A reconciliation worker treated news story groups like duplicate entities. It rejected distinct articles, cleared their memberships, and deleted the empty groups on repeated passes. The failure exposed a mismatch in the data model.

I introduced an explicit cluster role: DEDUP, STORY, or NONE. Generic reconciliation handles deduplication; the news domain owns story clustering. The fix moved the distinction into the contract so ownership is enforceable.

In practice

300K+ knowledge items · Six domains

A growing corpus with domain-specific ingestion and retrieval policies.

Python · FastAPI · PostgreSQL · Vector search · Workers