Skip to main content

Overview

AgentUse automatically manages conversation context to keep your agents running efficiently within token limits. This ensures long-running agents don’t hit token limits while preserving important information.

Automatic Compaction

When conversations approach token limits, AgentUse automatically:
  1. Detects when context usage exceeds the threshold
  2. Compacts older messages into a concise summary
  3. Preserves recent messages for continuity
  4. Continues the conversation seamlessly
Compaction happens transparently - your agent continues working without interruption.

How It Works

Token Tracking

AgentUse tracks token usage throughout the conversation and compacts when approaching limits. The system uses character-based estimation (approximately 4 characters per token) and updates with actual usage data from the AI models.

Compaction Strategy

AgentUse uses a summarization-based approach:
  • Older messages are summarized using the same model as the agent
  • Recent messages are preserved intact for continuity
  • The summary captures key decisions, tool results, and progress
  • System prompts and tool definitions are always preserved

Configuration

Environment Variables

Control context management globally using environment variables:

What Gets Preserved

Always Preserved

  • System prompts
  • Recent messages (last 3 by default, configurable)
  • Tool definitions remain available

What Gets Compacted

  • Older conversation history
  • Previous tool calls and results
  • Assistant responses from earlier in the conversation
The compaction creates a summary that preserves:
  • Key decisions and outcomes
  • Important tool results and errors
  • Current state and progress
  • Critical information needed for continuation

Compaction Details

Compaction Process

When compaction is triggered:
  1. Older messages are separated from recent ones
  2. A summarization request is sent to the same model
  3. The summary is created with a specialized system prompt
  4. Recent messages are kept intact for continuity
  5. The conversation continues with the compacted context

Compaction Prompt

The system uses this prompt for summarization:

Monitoring

Token Usage

AgentUse displays token usage at the end of each run:
For agents with sub-agents:

Compaction Events

When compaction occurs, you’ll see:

Context Limits by Model

AgentUse automatically detects context limits for different models:
  • Anthropic models: Retrieved from models.dev API
  • OpenAI models: Retrieved from models.dev API
  • Unknown models: Default to 32,000 tokens (conservative)
The system updates context limits dynamically based on the latest model information.

Performance Considerations

Compaction requires an additional API call to summarize context, adding 1-3 seconds to processing time
Compaction itself uses tokens (up to 2000 for the summary), but saves significantly more in long conversations
Summaries preserve key information but some nuance may be lost. Recent messages remain fully intact

Best Practices

1. Set Appropriate Thresholds

2. Adjust Recent Messages

3. Tune Approval-Boundary Compaction

Approval-boundary compaction is useful for long runs with human review gates. It can compact before the model window is close to full because the goal is to avoid resending a large active context after the session resumes.

4. Tune Step-Boundary Compaction

Step-boundary compaction runs after the agent has already made at least one tool call. It preserves the initial prompt, then summarizes older context before later model calls in long single-turn runs.

5. Disable if Not Needed

6. Cap and Reduce Tool Output

Every tool result is re-sent to the model on each subsequent step, so a single large output (a big diff, a verbose log, a huge file read) is one of the biggest drivers of input-token usage, often before compaction ever triggers. The most effective fix is to not generate the bloat in the first place, for example git diff --stat instead of a full git diff of high-churn files. As a safety net, direct JSON and text results over 10KB become a reusable handle with one partial value in preview and a compact omitted map showing what was left out at each jq-style path. The read-only results tool can return up to 20KB of selected text matches or JSON fields without repeating the preview or creating another handle. Provider-native and binary results using native model-output conversion remain on their direct delivery path. Bash separately retains a head + tail preview up to 30KB and saves a complete session artifact when its stream is larger. read_file defaults to 2000 lines. Tune these with the Tool Output environment variables:

Technical Details

Token Estimation

  • Uses character-based estimation: ~4 characters per token
  • Updates with actual token usage from AI models
  • Tracks both input and output tokens

Model Integration

  • Context limits are fetched from models.dev API
  • Falls back to conservative 32,000 token limit for unknown models
  • Caches model information for 24 hours

Error Handling

  • If compaction fails, creates a fallback summary
  • Continues execution even if compaction encounters errors
  • Logs compaction failures for debugging

Troubleshooting

  • Increase COMPACTION_KEEP_RECENT for more preserved messages
  • Lower COMPACTION_THRESHOLD to compact earlier with more context
  • The same model is used for summarization, so quality should be consistent
  • Lower COMPACTION_THRESHOLD to compact earlier
  • Reduce COMPACTION_KEEP_RECENT for more aggressive compaction
  • Consider the model’s actual context limit vs your usage
  • Check that CONTEXT_COMPACTION is not set to ‘false’
  • Verify the model supports the context limits being used
  • Check logs for compaction errors

Example: Long-Running Agent

The agent will automatically:
  • Track token usage throughout the analysis
  • Compact context when approaching 80% of the model’s limit
  • Preserve the last 5 messages for continuity
  • Continue analysis without interruption

Next Steps

Creating Agents

Build efficient agents

Environment Variables

Configuration options