Claude Code - Using Context More Efficiently
I’m trying to use Claude Code’s context more efficiently on long sessions. Every turn sends the whole conversation to the model, and even at the prompt cache’s discounted rate one long session costs more than the same work spread over a few short ones. Anthropic’s Maximizing the value of your Claude Code sessions covers the basics for keeping sessions short, like /clear between tasks, @-mentioning files instead of naming them, and keeping /model and /effort settled so you don’t bust the prompt cache. Read that first.
On a 1M-window model (Opus 4.7 and later on the Anthropic API), Claude Code doesn’t auto-compact until about 967k tokens, well past where I want to hear about it. /autocompact 300k makes it compact earlier on its own (model config). I use a Stop hook that warns me instead, plus a skill that hands the next task to a fresh session.
The hook goes in ~/.claude/settings.json and needs jq. It runs after every response, so it catches a bloated session before I notice the agent getting slower or worse:
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "jq -r '.transcript_path' | { read -r t; tail -n 200 \"$t\" | jq -s -r '[.[] | select(.type==\"assistant\" and .message.usage != null)] | last as $l | if $l == null then empty else (($l.message.usage.input_tokens // 0) + ($l.message.usage.cache_creation_input_tokens // 0) + ($l.message.usage.cache_read_input_tokens // 0)) as $tot | if $tot > 150000 then {systemMessage: (\"Context ~\" + (($tot/1000)|floor|tostring) + \"k tokens this turn (over 150k). Consider /compact, or run /kickoff and start a fresh session.\")} else empty end end'; } 2>/dev/null || true"
}
]
}
]
}
}
Every API response lands in the session’s transcript JSONL with a usage block, and the hook sums the last one locally, so the check costs no tokens. The two cache fields count because on a long session most of the context is read from the prompt cache, and input_tokens alone badly undercounts it. Past 150k it prints a systemMessage, which Claude Code shows me without sending it to the model (hooks reference). To watch the number all the time instead, the status line can show it as context_window.used_percentage (statusline docs).
A warning needs a good next step, and starting a new chat throws away real, unplanned-for work. By the time the warning shows up, what’s left is usually the next task rather than a half-finished one, so kickoff writes a prompt for that task instead of a summary of the session. I kept it to one paragraph in ~/.claude/skills/kickoff/SKILL.md:
---
name: kickoff
description: Write a prompt that starts the next task in a fresh session.
---
Write it as a prompt for the next session, not a document for a human.
That agent has the same repo access you do, so don't summarize or quote
code; point at the files and areas that matter and let it read them
itself. Include the actual task, why it matters, and whatever took real
investigation to establish, especially constraints or dead ends already
found. End on the first concrete action, specific enough to act on
without re-deriving it.
Then paste the prompt it writes into a new session.