--:--
--
← projects

Mint - Memory Intelligence

Building an offline AI memory that manages itself

Developers lose context constantly: why a decision was made, what broke last week, what that TODO meant. It lives in commits, terminals and heads, and it's gone when you need it. Mint's engine is already a standalone library, so the next step is Mint for developers: a mint CLI that sits in your workflow and remembers for you. Git hooks capture every commit with a local summary of what changed and why.

Git RepoComment

Summary

Mint is an offline AI memory for edge devices: it remembers what you tell it, files that knowledge into subjects on its own, answers from it in about 5 ms with no network, and decides what may sync to the cloud.

We built it in about a week for a Qdrant Edge challenge at Code Cubicle 6.0. The stack is Qdrant Edge, fastembed, Ollama, and a Rust engine inside a Tauri desktop app.

The result that matters most is not a feature. It is that we stopped judging the system by eye and started measuring it, and the measurements overturned two of our own fixes.

  • Retrieval: the right source ranked first went from 50% of questions to 100%. Mean reciprocal rank rose from 0.660 to 1.000.

  • Organization: topic grouping F1 rose from 0.733 to 0.944, with zero mixed topics among 150 deliberately similar distractor documents.

  • Privacy: the sync policy classified 44 of 44 labeled cases correctly, with no false positives on the research corpus.

  • Changing facts: version detection never marked a fact outdated wrongly, and caught 7 of 8 real updates.

  • Sync: 7 of 7 two-device scenarios passed against a real Qdrant Server.

  • Tests: 35 automated tests, all passing.

We did not advance past the online round on 3 October 2026. A job-application automation project won it. The last section covers what that taught us about presenting infrastructure.

The problem

The brief asked for an offline-first AI application on Qdrant Edge that keeps semantic memory on the device, works without a network, and syncs with a Qdrant Server when one is reachable.

It named the settings where this matters: robots, industrial systems, kiosks, vehicles, and mobile devices. In all of them connectivity is unreliable, latency matters, and some data must never leave the device.

Three of its eight goals are harder than they look, and they are the ones most projects skip:

  • Decide dynamically what stays local and what syncs. Syncing everything is easy. Deciding per memory, with a reason, is not.

  • Handle evolving memory, updates, and conflicting information. A store that only appends will hold "the exam is Friday" and "the exam is Monday" side by side.

  • Show a meaningful edge-to-cloud workflow, not just a local vector database. Backup is not a workflow.

The other five goals are the entry ticket: local semantic memory, low-latency hybrid search, sync on reconnect, a UI to inspect memory and sync state, and offline operation.

What I built

Mint is a desktop app whose engine is a standalone Rust library, so the app, the benchmark, and the test suites are three separate clients of the same code.

Everything a memory needs happens on the device. Embeddings come from fastembed (all-MiniLM-L6-v2, 384 dimensions) plus BM25 sparse vectors. Storage and search are an embedded Qdrant Edge shard. Reasoning is two local models through Ollama: qwen3:8b for chat and judging, llama3.2 for extraction and summaries.

Screen

What it does

Chat

Answers from memory, streams the model's reasoning, lists the memories each answer used, and shows what it stored

Memory graph

A force-directed map of memories, documents, entities, and topics, with a legend, topic focus, and controls to rename, merge, and move

Vault

Drop in PDFs, text, or code; each file is parsed, chunked, embedded, filed under a topic, and summarized

Timeline

Tasks and deadlines pulled from chat; a rescheduled date replaces the old one

Sync

Online and auto-sync switches, the privacy policy with reasons, and an activity log

The engine does six jobs: ingestion, topic-aware retrieval, organization into topics, version chains for changing facts, the sync policy, and a 3-way merge sync.

How it really went

For the first half of the project we fixed bugs by guesswork, and several of those fixes made the system worse without anyone noticing.

The bug that would not die

The question was "what type of internships are best for me?". The resume was in memory, but the answer came from a study guide, because the guide mentions internships and the resume never uses the word.

We tried four fixes in a row, each judged by re-asking that one question:

  1. A cap on results per document. The resume still was not among the candidates.

  2. Pulling in a document when the question names it. This question names nothing.

  3. Having a language model rewrite the question into several searches. It was slow, opaque, and did not help.

  4. Always adding the resume when a question contains "me" or "my". This fixed the one question we were looking at.

At the same time the capture step was storing junk. Saying "check from my resume" created a note titled "Resume Check", which then came back as a retrieved memory.

The turning point: measuring

We built a benchmark: a labeled corpus, 20 graded questions modeled on real failures, and metrics for retrieval, topic grouping, and capture. Then we ran it on the system as it stood.

The first run showed that our chat pipeline was worse than plain search. Mean reciprocal rank was 0.660 for the chat pipeline against 0.950 for raw hybrid search, on the same data.

What the benchmark exposed

Cause

Fix

Resume ranked first for visa, piano, and running questions

Fix 4 fired on every "my"

Add the profile only for questions about the person: skills, strengths, or "which jobs suit me"

Exoplanet paper injected into unrelated questions

Document names matched as substrings: "for" inside "Informed"

Match whole words, plus compounds and acronyms

One precise note lost to many weak matches

Topic voting was majority rule

Votes decay with rank; the top raw hit always keeps a slot

13 topics for 8 real subjects

New notes were compared with a 3-word topic name

Let already-filed neighbors vote for their topic

"Set a deadline" still created a note

The capture guard missed some command verbs

Add them; the timeline handles the date

The benchmark was wrong too

Our topic metric scored 0.973 on a configuration where a distractor "video" topic had swallowed a real research paper. It only compared labeled items with each other, so it could not see contamination.

We changed it to count every item and to report mixed topics. A sweep of the join threshold then showed a narrow safe value: 0.40 let a paper be absorbed, 0.45 kept every topic clean, and 0.50 split related notes apart.

Engineering deep dives

Five mechanisms carry the system: subject-first retrieval, self-organizing topics, a per-memory privacy policy, version chains, and a sync that never drops an edit.

1. Retrieval: subject first, then details

A question is embedded once, then goes through hybrid search, topic activation, and a weighted fusion of several ranked lists.

An invite-only space for considered discussion. Sign up, and you can comment once you are approved.

Comments

Loading comments…