Switching Your Default Classification LLM in Notestream: What Changes, What Doesn't

When you set up Notestream, one of the first choices you make is which LLM handles classification. The default works well for most people, but if you have ever wondered what actually changes when you swap from Claude to ChatGPT or Gemini — or whether the switch is worth making — this is the breakdown.

## What Classification Does

Every time you drop a capture into Notestream — a thought, a forwarded email, a quick note — the classification LLM reads it and decides what it is:

- A **task**: something to do
- A **sub-task**: something that belongs under an existing task
- A **note**: reference information worth keeping
- A **reminder**: time-sensitive

The LLM also assigns the capture to the right thread and extracts any implied due dates or priority signals. This happens automatically, on arrival, without a batch review step later.

The key constraint: classification is fast and targeted. The LLM is not summarizing a 40-page document. It reads a short, single-thought capture and makes a structural decision. That constraint shapes how much the choice of model actually matters in practice.

## The Three Options

Notestream currently supports three classification providers:

**Claude (Anthropic)** — The default. Strong at nuanced disambiguation. When a capture is ambiguous — "look into the pricing thing" could be a task or a reference note — Claude tends to apply the right judgment more consistently. It handles structured captures well: if you paste a formatted list or a short code snippet with a note attached, Claude treats each part correctly rather than collapsing them.

**ChatGPT (OpenAI)** — A solid alternative, especially if you are already deep in the OpenAI ecosystem and want consistency across tools. Classification accuracy on short captures is comparable to Claude. One area where ChatGPT occasionally diverges: captures phrased as passive observations ("might want to check on X sometime") sometimes get filed as notes when they are functionally tasks. Small, but worth knowing if you tend to write captures that way.

**Gemini (Google)** — The newest option. Fast. Works well for simple, clearly-phrased captures. On ambiguous phrasing, it is somewhat more likely to need a light correction. But for anyone who keeps captures short and direct, it performs comparably to the other two.

## What Does Not Change When You Switch

Everything else. This is worth being explicit about:

- Your threads and their structure stay exactly as they are
- Existing tasks and notes are not reclassified
- Dispatch behavior (sending tasks to Claude Code, Lovable, or Claude Cowork) is completely separate from classification and is unaffected
- Your inbox address stays the same
- Billing is not affected by which LLM you use for classification

The classification model is one isolated piece of the pipeline. Swapping it is scoped.

## What Does Change

**How ambiguous captures get filed.** The three models handle uncertain cases differently. If you are precise and concise with your captures, the differences are minimal. If you tend to write captures the way you would say them out loud — "probably need to handle the onboarding flow before launch" — Claude catches the task intent more reliably in those cases.

**Conversation style in the built-in AI chat.** If you use Notestream's in-app chat to follow up on a captured item, the conversation runs on whichever provider you have selected. This matters more than the classification difference for most people, since the chat is where you see the model's reasoning directly. Claude's conversational style is different from ChatGPT's, and you will notice it there before you notice it in silent classification.

**Model familiarity across your stack.** This is softer but real. If Claude is the model you use for everything else — drafting, research, debugging — keeping Claude as your classifier means one fewer mental context switch. The vocabulary and reasoning style stays consistent.

## How to Pick the Right Default

Start with the model you are most familiar with. If you are already using Claude in other parts of your workflow, keep it as the classifier. The mental model stays consistent and you will catch edge cases faster because the output style is familiar.

If your team or company is standardized on ChatGPT and you want uniformity in how things sound across tools, use ChatGPT. The classification output will feel native to what you are used to reading.

If your captures are short, high-volume, and clearly phrased — quick thoughts, URL drops, single-line notes — any of the three providers will handle it well. Spend that decision-making attention elsewhere.

If your captures tend to run longer or carry more context — notes from calls, pasted email threads, multi-sentence observations — Claude's disambiguation holds up better at the edges. That is where the difference between providers is most visible.

## Switching Mid-Stream

You can change your classification model at any time from account settings. The change applies to new captures immediately. Captures that have already landed stay filed as they are.

There is no migration step or re-indexing. Notestream stores the classified output, not the prompt that produced it. Once a capture is a task or a note, it stays that way until you edit it directly.

If you are evaluating two models, give each at least a week with your real capture volume before deciding. The differences show up in edge cases, which accumulate over time with normal usage. Running staged comparisons on the same ten captures misses the tail cases that actually matter in production.

## If You Are Seeing Consistent Misfiling

Before switching providers, check the phrasing of the captures that are getting filed wrong.

"Check on pricing" is ambiguous to any classifier. "Task: revisit the pricing page before Thursday's call" is not. A lot of what looks like a model accuracy problem is a capture clarity problem. Tightening the phrasing usually resolves it faster than changing the provider.

That said, if a specific capture type — forwarded emails, or captures pasted from a particular source — consistently misfires on one provider, that is a real signal. Try the swap for that category and compare.

## Where to Find the Setting

Classification model preference lives in account settings under AI preferences. If you have been on the default since you signed up and have not thought about it, that is likely the right call. If you are noticing patterns in how certain captures land, it is a one-minute change to try a different provider.

[notestream.ai](https://notestream.ai)