MEMO · TOKENS · AUGUST 2026

Six habits that put you ahead of the Claude Code game

9 min read · every claim sourced below

The short version: Claude Code re-sends your entire conversation on every single message, so everything still in it gets paid for again on every turn. Six habits keep it small. /clear between unrelated tasks. Pick your model and effort level before you start, not halfway through. Attach files with @ instead of naming them. Keep noisy output out of the chat with a quiet flag or a subagent. Run /context once in a fresh session to see what loaded before you typed a word. And run /compact before you walk away, while the cache is still warm.

None of this is secret. All of it is in Anthropic’s own documentation, and every claim below links to the page it came from. What the docs do not do is put the six together in one place with the reason each one works, which is what this is.

claude
>your session
/clear
model + effort
@ the file
quiet output
/context
/compact
Six habits, and the session stops carrying everything you already finished.

Why a Claude Code session costs more than it looks

The model does not remember anything between requests. So every time you send a message, Claude Code re-sends the whole lot: the system prompt, your project context, every prior message and tool result, and then your new message on the end.

That is why a session behaves the way it does. A one-line question in a conversation that has been open all day still carries the whole day with it. Prompt caching softens the blow, because the unchanged part is billed at roughly a tenth of the normal input rate rather than full price, but it is still being sent and it still counts.

The cache works by matching the start of each request against what it recently processed. The match is exact, so a change anywhere near the front recomputes everything after it. There is no per-file or per-message caching to fall back on. That single rule is behind habits two, four and six.

The six habits

In the order they matter, with the command and the reason it works.

1

Run /clear when you switch tasks

Everything you already solved is still riding along on every new message. The documentation puts it plainly: clearing between tasks starts fresh when you switch to unrelated work, and stale context wastes tokens on every subsequent message.

between unrelated tasks
/clear

The part people miss is that this one is free. Anthropic’s caching page says so directly: when you want a fresh start instead of continuity, /clear costs nothing. Compare that to /compact, which has to read the conversation in order to summarise it.

If you might want the thread back, name it before you clear and reopen it later:

keep it findable first
/rename auth-refactor

Then /resume brings you back to it. There is no reason to keep a finished task loaded just in case.

This matters most right after anything that puts images in the conversation. Reading a video with the /watch skill can drop 50,000 to 80,000 image tokens in one go, and every one of them is re-sent with every message afterwards until you clear.

The same arithmetic decides whether a heavier Claude surface is affordable. Anthropic says Claude Science burns tokens at a rate comparable to heavy Claude Code use, drawn from the same plan limits, so every habit on this page applies there too.

2

Pick your model and effort level before you start

Each model has its own cache. Switching with /model partway through means the next request reads the entire conversation with no cache hits at all, even though the content is identical.

at the top of the session, not halfway
/model

Not sure which one to pick in the first place? The short answer is Sonnet, and the long one is in Claude Sonnet vs Opus vs Fable: which to pick.

Effort works the same way. The cache is keyed by effort level as well as model, so changing it mid-session recomputes the whole request too. Claude Code will show you a confirmation dialog before applying a change that would do that, which is a useful tell in itself. Setting an effort that resolves to the level already in force skips the dialog and keeps the cache.

same, up front
/effort
The one that catches people out

If your model setting is opusplan, it resolves to Opus during plan mode and Sonnet during execution. That means every toggle in and out of plan mode is a model switch and starts a fresh cache. Worth knowing before you decide it is free to flip in and out.

Turning on fast mode has the same one-off cost, because it adds a request header that is part of the cache key. That is why enabling it at the start of a session is cheaper than enabling it deep into a long one.

3

Attach files with @ instead of naming them

Say “the auth file” and Claude has to go and find it, which usually means a search and then one or more reads, all of which land in the conversation and stay there. Point at it with @ and the file’s full contents go into the message directly, with no waiting for a read.

the file arrives attached
Explain the logic in @src/utils/auth.js

You can reference several in one message, and paths can be relative or absolute. Type @ on its own to open the path suggestion menu rather than typing it out.

Two things worth knowing before you use it everywhere

A directory reference gives you a file listing, not the contents of everything inside, which is usually what you want and occasionally not.

And an @ file reference also pulls in CLAUDE.md from that file’s directory and its parents. That is helpful when you want the local rules, and it is extra context you did not ask for when you do not.

4

Keep loud output out of the chat

A command that prints thousands of lines dumps all of them into the conversation, and they stay there for the rest of the session, re-sent with every subsequent message. Test runners, log files and build output are the usual offenders.

The cheap fix is to ask for less in the first place, with whatever quiet or reporter flag the tool offers.

The thorough fix is to make it automatic. A PreToolUse hook can rewrite noisy commands before they run, so only the failures come back:

~/.claude/hooks/filter-test-output.sh, the filtering line
filtered_cmd="$cmd 2>&1 | grep -A 5 -E '(FAIL|ERROR|error:)' | head -100"

The other route is to hand the job to a subagent. It reads and runs in its own context window, and only a summary comes back to your conversation.

the verbose part happens somewhere else
use a subagent to investigate how our auth system handles token refresh

Your own cache is untouched by this: from the parent’s side, the subagent’s call and its result simply append to the conversation. Two things to know, though. A subagent starts its own conversation with its own system prompt, so its first request does not read your cache and has to build its own. And subagents use the five-minute cache even on a subscription, where the main conversation gets an hour.

Agent teams are a different order of magnitude

Agent teams use roughly seven times the tokens of a standard session when teammates run in plan mode, because each teammate is a separate Claude instance with its own context window. Every teammate keeps consuming until it exits. Keep teams small, keep their tasks self-contained, and shut them down when the work is done.

5

Run /context once in a fresh session

Before you have typed a single word, your context window already contains the system prompt, every tool definition, your MCP tools and your memory files. /context shows you the breakdown, including how much free space is left.

in a brand new session, before you type
/context

You are looking for anything large that you are not actually using today. The two usual finds:

A sprawling CLAUDE.md. It loads at session start and is present on every message, even when you are doing something it has nothing to say about. Anthropic’s own guidance is to keep it under 200 lines and move workflow-specific instructions into skills and config that load on demand instead. If yours has grown into something nobody wants to prune, our free CLAUDE.md generator asks six questions and hands you a short one to start again from.

MCP servers you are not using. Tool definitions are deferred by default on supported models, so only names enter context until a tool is actually used, but not every setup gets that. Run /mcp and turn off what you are not using. Where a CLI exists, it is more context-efficient than an MCP server, because it adds no per-tool listing at all.

6

Run /compact before you walk away

Compaction replaces your message history with a summary. To produce that summary, Claude Code sends a separate request carrying the same system prompt, tools and history as your conversation.

at a natural break, while the cache is still warm
/compact

That is the whole reason the timing matters. While the cache is warm, that request reads your prefix from the cache, so compacting costs a fraction of what the size of the conversation suggests. After a break longer than the cache lifetime there is nothing left to read, and the same request reprocesses the entire history as uncached input. Compacting is at its most expensive exactly when you resume an old session.

You can also tell it what to keep:

steer the summary
/compact Focus on code samples and API usage

Or set that once for the project, in CLAUDE.md:

CLAUDE.md
# Compact instructions — When you are using compact, please focus on test output and code changes

How long the Claude Code prompt cache actually lasts

This is the number the reel could not carry, because it is not one number. The cache expires after a period of inactivity, and every request that hits it resets the timer. How long a gap it survives depends entirely on how you sign in.

How you authenticateCache lifetime
Claude subscription (Pro, Max, Team, Enterprise)1 hour, requested automatically
Subscription that has gone over and is on usage credits5 minutes
API key, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS5 minutes by default

If you are on an API key and want the longer one, or you are on a subscription drawing on usage credits and want to keep it:

opt back into the one-hour cache
ENABLE_PROMPT_CACHING_1H=1

FORCE_PROMPT_CACHING_5M=1 forces the short one, which is mostly useful for debugging or for overriding a setting handed down in managed settings.

Everything else that quietly resets the Claude Code cache

Habits two and six cover the two you control most often. This is the rest of the list, so a slow turn stops being a mystery. Each of these makes the next request read the whole conversation again.

  • Switching model, including every plan-mode toggle under opusplan, and the automatic model fallback when a safety classifier reroutes a request.
  • Changing effort level.
  • Turning on fast mode, once per conversation. Turning it off again, and the automatic fallback to standard speed after a rate limit, both keep the cache.
  • Connecting or disconnecting an MCP server whose tools are loaded into the prefix rather than deferred. This one can happen without you doing anything: a stdio server’s process exits, an HTTP session expires, or a server reconnects after a blip.
  • Enabling or disabling a plugin that ships an MCP server. Plugins that only add skills, commands, agents, hooks or themes are free, because what they add is appended rather than inserted at the front.
  • Denying a whole tool by bare name in permissions, such as Bash or WebFetch, because that removes it from context entirely. Scoped rules like Bash(rm *) are fine, and so is every allow and ask rule.
  • Compacting, by design.
  • Upgrading Claude Code. Auto-update applies on the next launch rather than mid-session, so you see it as one uncached first turn after a restart.
The most expensive request you will ever send

Resuming a long session after an upgrade reprocesses the entire history with no cache hits, because that history now sits behind a different system prompt. The cost scales with how long the conversation is. If you are coming back to a very long session on a fresh version, expect the first turn to hurt.

Four things that surprise people

Editing CLAUDE.md mid-session does nothing. It is read once at session start and held in memory. The edit does not break the cache, but it does not apply either, and Claude keeps working from the version loaded when the session began. It lands on the next /clear, /compact or restart. Output style behaves identically. Nested CLAUDE.md files in subdirectories are the exception, because they load later, when Claude first reads a matching file.

/rewind is cheaper than /compact when you are abandoning a path. It truncates the conversation back to an earlier turn, and that remaining history is exactly what the cache was built from, so the next request hits the earlier entry. Compaction builds a new prefix; rewinding returns to one you already have.

The cache is scoped to one machine and one directory. The system prompt embeds the working directory, so two sessions in different directories never share a cache. That includes two worktrees of the same repository. Sessions run in parallel in the same directory do read each other’s.

Idle sessions are not free. A scheduled task fires on its interval and sends your full context each time. A message from another of your sessions is delivered as a new turn and does the same. Active agent teammates keep consuming until they exit.

How to check whether any of this worked

Two token counts tell you almost everything. Cache reads are billed at roughly a tenth of the standard input rate; cache creation is the write. A high read-to-creation ratio means caching is doing its job. If creation stays high turn after turn, something in your prefix keeps changing, and the list above is where to look.

tokens, cost and the cache split for this session
/usage

On a paid plan, /usage also flags behaviours accounting for ten per cent or more of recent usage, including long context and cache misses, which is the fastest way to find out which of the six you are getting wrong. A statusline script can display context usage continuously if you would rather watch it than ask.

What we would actually do

Set the model and effort once, at the top. Use @ by default rather than describing files. Run /context the first time you work in a new project and fix whatever is unnecessarily large. Then, for the rest of your life, do the two that cost nothing and take a second: /clear when you finish something, /compact before you walk away from something unfinished.

The rest of the list is worth reading once so that a slow turn is explicable rather than alarming. It is not worth memorising.

Questions people actually ask

How do I reduce token usage in Claude Code?

Keep the conversation short and stable. Claude Code re-sends the entire conversation on every message, so anything still in it is paid for again on every turn. Run /clear when you switch to unrelated work, pick your model and effort level before you start rather than halfway through, attach files with @ instead of naming them, keep noisy command output out of the chat with a quiet flag or a subagent, run /context once in a fresh session to see what loaded before you typed, and run /compact at a natural break rather than waiting for it to trigger mid-task.

Does /clear cost anything?

No. Anthropic's own documentation is explicit that when you want a fresh start rather than continuity, /clear costs nothing. /compact is the one with a price, because producing the summary means sending the conversation to be summarised. That is the whole reason to prefer /clear when you genuinely do not need what came before.

Why does my usage climb in a long session even when I barely type?

Because every request carries the full conversation. A one-line question in a session that has been open all day still sends the whole day with it. Prompt caching means that history is billed at the much cheaper cached rate rather than full price, but it is still being sent and still being counted. Scheduled tasks, cross-session messages and active agent teammates also fire while the session sits idle, each one sending the full context again.

Does switching model mid-session really cost extra?

Yes. Each model has its own cache, so switching means the next request reads the entire conversation history with no cache hits, even though the content is identical. Effort level works the same way: the cache is keyed by effort as well as model, and Claude Code shows a confirmation dialog before applying a change that would invalidate it. Setting an effort level that resolves to the one already in effect skips the dialog and keeps the cache.

How long does the Claude Code prompt cache last?

It depends on how you sign in. On a Claude subscription Claude Code requests the one-hour cache automatically, so it survives breaks of up to an hour. On an API key, Amazon Bedrock, Google Cloud, Microsoft Foundry or Claude Platform on AWS it is five minutes by default, and a subscription that has gone over its limit and is drawing on usage credits also drops to five minutes. Set ENABLE_PROMPT_CACHING_1H=1 to opt back into the hour.

Should I use /compact or /clear?

Use /clear when you are done with what came before, because it costs nothing. Use /compact when you need continuity and want to keep working in the same thread. If you are abandoning a path rather than finishing one, /rewind is better than either: it truncates back to a prefix that is already cached, where compaction builds a new one.

Why did editing CLAUDE.md mid-session do nothing?

Because it is read once at session start and held in memory. Editing it does not invalidate the cache, but the edit does not apply either, and Claude keeps working with the version that was loaded when the session began. The new content loads on the next /clear, /compact or restart. Output style behaves the same way. Nested CLAUDE.md files in subdirectories are different: they load when Claude first reads a matching file, so editing one before it loads does take effect.

Want this built for you?

We write these memos because we build this stuff every day. If you want it working in your business instead of sitting on your reading list, that is literally our job.