When Your Notes Meet AI: Assistants That Read Your Archive Without Uploading It

Letting an AI work with your notes usually means handing them to a server. It doesn't have to. How on-device access works, and what to ask before allowing it.

· Privacy & Ownership · 7 min read

The pitch is genuinely appealing. You have four years of notes — meeting records, half-finished ideas, book highlights, the thing you wrote at 1am that you'd recognize but can't find. An assistant that could read all of it would be more useful to you than one that can't, by a wide margin, because most of what makes an assistant useful is context it doesn't have.

The price, as usually configured, is that your entire archive gets copied to a company's servers so a model can index it. And that's a genuinely uncomfortable trade for the exact notes that would benefit most — the private ones, the ones about other people, the ones you write precisely because nobody is reading. It's worth knowing that the trade is a product decision rather than a law of physics. There are at least three architectures here, they differ enormously in what leaves your device, and most people are only shown one.

Three ways an assistant can reach your notes

Upload and index. Your notes are copied to the provider, chunked, embedded into a vector index, and searched at query time. This is what "connect your notes" means in most products. It works well and it means a full, durable, searchable copy of your archive now lives on someone else's infrastructure, subject to their retention policy, their breach surface, and their next change of terms.

Run the model locally. A small model runs on your own hardware and nothing leaves at all. The privacy story is perfect. The capability story is not: local models are meaningfully weaker than frontier ones, and running one comfortably needs hardware most people don't have. Good for classification and summarizing a page; not for the "find the thing I half-remember from 2023" task that motivated this in the first place.

Give a capable model a tool instead of a copy. The middle path, and the one that has quietly become practical. A hosted model keeps its capability, but rather than being handed your archive up front, it's handed a set of actions it can invoke against notes that stay where they are — search, read one note, append to a note. It asks for what it needs, when it needs it.

That third one deserves unpacking, because the distinction it draws is the whole point.

Tool access versus bulk access

Under the tool model, the assistant gets a small vocabulary: search my notes for X, read the note called Y, append this to Z. Each call runs locally — in the app, on your device — and only the result comes back. Ask it to summarize your Q3 meetings and it searches, gets four notes, reads them, and answers. Four notes crossed the boundary. Not four thousand.

Three properties follow, and they're the reasons to care:

  • Exposure is proportional to the question. No question, no data movement. A narrow question moves a narrow slice.
  • There's no durable second copy. Nothing was indexed, so nothing persists on the other side after the conversation ends beyond whatever the conversation itself retains.
  • The surface is inspectable. The tool list is finite and readable. "Can read notes, can append, cannot delete" is a sentence you can verify — very different from auditing what an indexing pipeline retained.

The honest costs, because they're real. Retrieval is only as good as the search underneath it, which is keyword-shaped in most local implementations — so an assistant may miss a note it would have found via semantic search over an index. Content still crosses the wire per query if the model is hosted. And the app has to be open and running, because there's no server holding a copy to answer on your behalf.

In the browser, this pattern now has an emerging standard behind it: WebMCP, which lets a web page declare tools that an AI agent in the browser can call, with the page — not a server — executing them. It's the same idea as MCP on the desktop, moved to the tab.

Indenta implements exactly this shape, which makes it a concrete example to reason about: notes live in the browser's local database with no server anywhere, and the page registers a fixed set of tools — list, search, read, create, append or replace, rename, delete, open — that a browser-based agent can call while the app is open. The archive is never uploaded; individual notes travel only as the answers to specific calls, and the User Guide covers the setup. It's a demonstration that "capable assistant" and "archive stays on device" aren't mutually exclusive, not a claim that the tradeoffs above disappear.

Write access is a different decision from read access

Almost every discussion of AI and personal data stops at reading. But the tools that make an assistant genuinely useful — file this, append that, rewrite this note — are write tools, and they carry a distinct risk that has nothing to do with privacy.

An assistant that can modify notes can modify them wrongly. A "cleanup" that flattens structure you cared about. An append that lands in the wrong note. A rewrite that improves the prose and drops the one detail that mattered. None of this requires bad intent from anyone; it requires an ordinary misunderstanding, which is a thing that happens every day.

There's a sharper version, too. If an assistant reads a note and then acts, anything written in that note is text the model sees — including text someone else put there. A meeting note pasted from an email, a clipped web page, a shared document: content from outside your head, being read by something that takes instructions in the same language it reads content. Treat "the assistant read it" and "the assistant should do what it says" as strictly separate things, and prefer setups where write actions are visible to you rather than silent.

Practical posture:

  • Start read-only and stay there until the assistant has actually earned something.
  • Prefer append over replace. Appending is additive and reversible by eye; replacing a note's body is a silent overwrite.
  • Keep deletion manual. There is no assistant task worth automating a delete for.
  • Have a backup before you start. The backup rules don't change because the actor is a model; a plain-Markdown export is the undo button for everything above.

What to actually ask a tool

Five questions, in the order that matters:

  1. What leaves my device, and when? The answer should be a sentence, not a policy page. "Whole archive, once, at connect time" and "matching notes, per query" are wildly different products.
  2. Is there a durable copy on the other side? Index, embeddings, cache — all of it counts. Ask how long it lives after you disconnect.
  3. Is the content used for training? Default-off, or off with one setting you can find.
  4. What exactly can it write? Read-only, append-only, or full replace and delete. If the tool list isn't published, that's the answer.
  5. Can I get everything out regardless? A plain-text export is what makes any of these decisions reversible. Without it, you're not choosing a setup — you're choosing a landlord.

The trade is narrower than it looks

The default framing — full AI usefulness or full privacy, pick one — describes one implementation, not the problem. Between "upload everything" and "no assistant at all" sits an arrangement where a capable model gets a small, inspectable set of actions against notes that never move, and the exposure is whatever a specific question required.

That's not free of tradeoffs, and this article has tried to name them rather than skip them: weaker retrieval, per-query transmission, the app has to be open, and write access is a genuine risk you should ration. But those are ordinary engineering costs, and you can decide about them. Handing over four years of private writing because the connect button asked for it is not a decision anyone actually made.


Try it in practice: Indenta is a free, offline-first outliner — nested notes, backlinks, tags, and peer-to-peer sync, with no account required. Start writing in your browser, or read the User Guide first.

← All articles