← All projects

Edith / Personal AI

The conversation should move forward. Your context should come with it.

A WhatsApp assistant built around a difficult question: how do you stay useful across many conversations, projects, and changing priorities?

Private beta

A question worth building for
“What did I decide about Atlas last time—and what should I do next?”
Remember the right detailChoose the right workflowReply in one voice

Illustrative scenario

The spark

I kept repeating context that an assistant should already know.

Remembering a fact is only the start. The useful part is finding it later, connecting it to the right request, and letting newer information update the picture without mixing unrelated work.

A reply in the moment

RecallRelevant memories

RouteFast reply or research

RespondOne consistent voice

Continuity over time

ExtractFacts and actions

ReconcileDuplicates and updates

Retrieve laterContext when relevant

Under the surface

Follow a message through Edith.

A branching workflow, with memory maintained alongside the conversation.

01 / Recall

Retrieve what matters to this request.

Semantic search retrieves relevant memories within the user’s namespace, with rules loaded separately. Recent context complements retrieved matches, so useful older information has a path back into the conversation.

Decision: retrieve relevant context instead of replaying the entire history.

02 / Route

Give each request the depth it needs.

The orchestrator can move from recall to a reply, use web search, or take a deeper research path. Short requests can stay short. Tool use follows the work required rather than one mandatory sequence.

Decision: spend latency and model calls selectively.

03 / Remember

Extract facts without blocking every reply.

Background or post-response extraction identifies memories, rules, and structured records such as actions. Central writes attach source, confidence, scope, and entity metadata. Similarity checks skip near-duplicates and identify likely updates.

Decision: maintain memory as data with provenance.

04 / Respond

Separate reasoning from the final voice.

Reasoning and tool results feed the response layer, which handles Edith’s voice before delivery through WhatsApp. Transport, reasoning, and presentation have separate responsibilities.

Decision: keep a consistent experience across different workflows.

The hard part

Memory needs an update policy, not just a vector search.

A user correction should carry more authority than an inferred fact. I centralized memory writes so inferred content cannot overwrite user-sourced information, added semantic deduplication, and invalidated cached context after changes.

Similarity is a useful signal, but it is not proof that two facts contradict each other. I also benchmark routing with the actual runtime prompt: the selected baseline passed 18 of 21 repeated cases. A faster candidate reached 15 of 21, so I kept the stronger routing baseline and targeted other layers for cost savings.

In practice

18/21 routing cases passed

2.18s median orchestration latency · April 4, 2026
Repeated benchmark cases using the runtime prompt; measures routing, not complete reply latency.

WhatsApp Cloud API · Agents · Vector retrieval · DynamoDB