> For the complete documentation index, see [llms.txt](https://docs.warp.dev/llms.txt).
> Markdown versions of each page are available by appending .md to any URL.

# Use tokens more efficiently with AI coding agents

Get more out of every token your coding agents use in Warp with model choice and routing, focused context, conversation management, and Rules.

Tokens are the text a model reads and generates. Inference cost depends on the model’s prices and the mix of input, output, cache-read, and cache-write tokens. The same token count can cost different amounts.

Compare token details and cost when choosing a model or reducing context. Dollar-billed plans report usage charges in dollars; credit-based agreements retain their units. See [usage and billing](https://docs.warp.dev/support-and-community/plans-and-billing/credits/).

## Track your usage first

-   **Per-turn breakdown** - In the Warp app, expand the usage amount below a response. Where detailed usage is available, compare each model’s token categories and cost.
-   **Account usage** - In the Warp app, open **Settings** > **Billing and usage** for your allowance, balance, and reset date.
-   **CLI summaries** - See [CLI usage and cost](https://docs.warp.dev/agents/cli/models-and-usage/#usage-and-cost) for `/usage`, `/cost`, and the conversation statusline. These surfaces do not all use the same display unit.

## Match the model to the task

Models have different token prices. Choose the capability the task needs, then compare its reported cost.

-   **Use a cost-efficient model for routine work** - Choose **Auto (Cost-efficient)** (`auto-efficient`) or a lower-priced model for edits, lookups, and short questions.
-   **Prefer open-weight models** - If you’d rather run open-source models, choose **Auto (Open-weights)** (`auto-open`), which routes among the best open-weight models for low cost and fast speed.
-   **Reserve high-reasoning models for hard problems** - Save heavier models like Claude Opus for deep debugging, architecture decisions, and planning, where the extra reasoning is worth the cost.
-   **Pick a model and stay with it** - Switching models mid-conversation can reset prompt caching and reprocess your context. Choose a model at the start of a task and keep it for the duration when you can.

Change models with the model picker in the input, or run `/model`. See [Agent model choice](https://docs.warp.dev/agents/inference/model-choice/) for the full model list.

## Automate model selection with custom routers

Use a custom router to automatically choose a model. You define the routing logic once, and Warp resolves a concrete model for each prompt instead of defaulting every task to your most expensive model.

-   **Route by complexity** - Warp classifies each task as easy, medium, or hard and routes to the model you mapped to that level. Simple tasks run on a lightweight model; only difficult tasks reach a high-reasoning model.
-   **Route by rules** - Write natural-language rules that pair a description (such as “debugging or fixing failing tests”) with a model. Warp matches rules top to bottom and uses the first one that fits.
-   **Set a cost-efficient default** - Every router falls back to a default model for anything your tiers or rules don’t cover, so choose a lighter model for the default.

A router appears in the model picker like any other model and resolves per conversation, so token usage matches whichever model it picks. Create one in the Warp app under **Settings** > **Agents** > **Warp Agent** in the **Custom Routers** section. See [Custom routers](https://docs.warp.dev/agents/inference/custom-routers/) for setup steps and YAML examples.

## Keep each conversation focused

Because every turn re-sends the current conversation to the model, long or unfocused threads keep paying for the same context. Tight, well-scoped conversations keep token usage low.

-   **Scope tasks and work incrementally** - Break large changes into smaller, contained steps instead of one sprawling request. Well-scoped tasks need less back-and-forth and fewer correction cycles.
-   **Start a new conversation for a new task** - Run `/new` when you switch topics so unrelated history doesn’t ride along in every turn.
-   **Compact long conversations** - When a useful thread grows long, run `/compact` to summarize the history and free up the context window. Use `/fork-and-compact` to branch into a fresh, summarized copy that keeps the relevant context and trims the rest.

See [Conversation forking](https://docs.warp.dev/agents/local-agents/interacting-with-agents/conversation-forking/) and the full [Slash Commands](https://docs.warp.dev/agents/capabilities/slash-commands/) reference for more.

## Be selective about the context you add

Context you attach becomes tokens the model has to process. Adding only what’s relevant keeps each turn lean.

-   **Attach focused snippets, not full dumps** - When sharing logs, code, or command output, include only the relevant portion instead of an entire file or output.
-   **Add context deliberately** - Attach the specific [blocks](https://docs.warp.dev/agents/local-agents/agent-context/blocks-as-context/), files, or images the agent needs for the task, rather than broad, just-in-case context.

## Let Codebase Context retrieve code for you

When an agent explores your repository by reading files one by one, each read is a tool call that consumes tokens. Indexing your codebase lets Warp find the right code with semantic search instead.

-   **Index your repository** - Run `/index` so Warp can locate relevant code by meaning, reducing the number of exploratory tool calls and the amount of code you paste in manually.
-   **Let the agent search instead of pasting** - With an indexed codebase, ask about a feature or file directly rather than copying large sections into the prompt.

Learn more in [Codebase Context](https://docs.warp.dev/agents/capabilities/codebase-context/).

## Set up Rules and AGENTS.md

Without persistent guidance, agents re-derive your preferences every session and sometimes drift off course, which wastes tokens on corrections and rework. Rules encode that guidance once.

-   **Capture preferences as Rules** - Store your tools, conventions, and standards as [Rules](https://docs.warp.dev/agents/capabilities/rules/) so you don’t re-explain them in every conversation. Add one with `/add-rule`.
-   **Add a project AGENTS.md** - Run `/init` to generate a project `AGENTS.md` that gives agents the context they need up front, reducing exploration and missteps.

For examples, see [Set coding best practices with Rules](https://docs.warp.dev/guides/configuration/how-to-set-coding-best-practices/).

## Plan large or complex tasks first

For big or ambiguous tasks, jumping straight to implementation often leads to wrong turns and expensive rework. A short planning pass keeps execution on track.

-   **Create a plan before executing** - Run `/plan` to have the agent research and outline the work in phases before it changes code. A clear plan reduces wasted exploratory work and backtracking on large tasks.

See [Planning](https://docs.warp.dev/agents/capabilities/planning/) for details.

## Next steps

-   [Use Agent Profiles efficiently](https://docs.warp.dev/guides/configuration/how-to-use-agent-profiles-efficiently/)
-   [Agent model choice](https://docs.warp.dev/agents/inference/model-choice/)
-   [Custom routers](https://docs.warp.dev/agents/inference/custom-routers/)
-   [Slash Commands](https://docs.warp.dev/agents/capabilities/slash-commands/)
