Skip to content

Context Engineering

The discipline of designing what information enters a model's context window, how it is structured, and what is excluded — to maximize the quality and reliability of agent output.

Fundamentals

Core concepts that define context engineering as a practice and establish the structural patterns every other technique builds on.

  • Context Engineering: The Discipline of Designing Agent Context — Context engineering is the practice of designing what information enters a model's context window, how it is structured, and what is excluded
  • Context Quality as a Leading Indicator of Agent Reliability — Audit an agent's context across seven dimensions to predict where it will drift, hallucinate, misuse tools, or fall to injection before blaming the model
  • Context Priming — Load relevant context before asking an agent to act; the order information enters the context window shapes the quality of everything that follows
  • Layered Context Architecture — Ground agents in multiple distinct context sources — schema, code, institutional knowledge, and persistent memory — rather than relying on any single signal
  • Context Budget Allocation — Context is a finite budget; every token preloaded into the context window displaces a token available for reasoning, tool results, and implementation
  • Context Lifecycle Management — Treat context as a managed lifecycle — decide, extract, store per type, consolidate, compact — rather than a passive store, to curb missing recalls and token cost that grows every turn
  • Discoverable vs Non-Discoverable Context — Only put non-discoverable information in agent instruction files; if the agent can find it in the codebase, let it find it
  • Instruction-Guided Code Completion — Functional correctness and instruction adherence are independent capabilities; explicit implementation constraints and model selection close the gap
  • Re-Auditing Context Engineering Across Model Generations — Context-engineering best practices are model-generation-dependent; a capability jump lets you delete guardrails the new model no longer needs while keeping the security-critical ones
  • Coding-Agent Working-Set Coverage (Coherence Debt) — Repository-scale coding agents fail when the coupled facts an upcoming edit needs are neither in active context nor in parametric memory; availability decides the outcome, not distance

Attention & Positioning

Models do not attend uniformly across the context window. These pages cover where attention concentrates, where it drops off, and how to structure content accordingly.

  • Attention Sinks — Transformer models disproportionately attend to initial tokens regardless of their semantic content; position determines attention weight, not importance
  • Lost in the Middle — Model attention is strongest at the start and end of a context window; content in the middle receives significantly less focus regardless of its importance
  • Context Window Dumb Zone — Output quality degrades as context fills, but the onset depends on task type; retrieval, reasoning, and code generation hit different thresholds
  • Manual Compaction as Dumb Zone Mitigation — Auto-compaction fires at ~95% context fill, long after reasoning quality has degraded; manual compaction reframes context management as reasoning quality preservation
  • Observation Masking — Strip intermediate tool results from conversation history once they have served their purpose to keep active context lean without losing the work product
  • Context Window Anxiety — Advanced models exhibit behavioral shortcuts as context limits approach; strategic buffers, counter-prompting, and token budget transparency counteract premature task closure
  • Turn-Level Context Decisions — Every completed turn is a branching point with five options: continue, rewind, clear, compact, or delegate to a subagent; choosing well is the core skill of context management
  • Conversation Registers for AI Coding Sessions — Name which of four interaction modes you are in with an LLM — exploring, brainstorming, deciding, implementing — and start a fresh context when the register changes
  • Comment Content as Code-Generation Context — A model conditions the code it writes on the comments it just wrote; only content correctness moves pass@1, and a comment written for a different problem costs 20.8%

Compression & Caching

Strategies for fitting more useful content into less space, and for making repeated prefixes cheaper through provider caching mechanisms.

  • Context-Window Diagnostic Tooling — Surface which tool calls are inflating the context window so you can optimize specific culprits rather than prune blindly
  • Reducing System-Prompt Token Bloat in Coding Agents — Measure the shipped system prompt with /context, then switch off unused tools, skills, and features so the fixed prefix stops crowding out task context
  • Proprioceptive Context Dashboard — Give a long-horizon agent a live view of its own context blocks — size, age, and usage — so it makes competent keep-or-archive decisions itself instead of a hidden layer compressing blindly
  • PEEK: Orientation Cache for Recurring-Context Agents — A constant-sized prompt artifact that caches reusable orientation knowledge — what is in a recurring context, how it is organized, which entities matter — distinct from trajectory replay and playbook strategy memory
  • Context Compression Strategies — Long-running agents accumulate context that eventually fills the window; tiered compression — offloading large payloads and summarizing history — lets agents continue working without losing task continuity
  • Per-Type Retention Policy for Agent Compaction — Compaction summarizes a safety rule at the same rate as a chat log; pin the constraints first, then classify each line by type and set retention fidelity per type
  • Action-Gated Context Trimming for Long-Horizon Agents — Vary trimming aggressiveness with workflow state: hold every item an open dependency needs, and ease off compression before an irreversible write
  • Specification Memory: What a Shared Agent Workspace Keeps — Persist four categories of user-authored context across sessions and discard the transcript; the store that kept less completed 12/12 trials at the token cost of keeping nothing
  • The Handoff Tax: What a Receiving Model Should Inherit — Switching models mid-task costs quality and money, and the trajectory the receiver should inherit reverses with direction: drop the cheap model's on escalation, keep the strong model's on downshift
  • Selective Rewind Summarization — A user-chosen cut point compresses earlier turns to a summary while the recent turns stay verbatim — a targeted alternative to whole-session compaction
  • Addressable Recall Compaction — Store every tool observation verbatim under a stable ID and show a compact citation once context fills, so the agent recovers the exact record instead of a summary or a similarity match
  • Usage-Reinforced Memory Decay for Long-Running Agents — Score retention with a forgetting curve whose stability compounds on every recall, so frequently-consulted facts outlive the idle stretches that silently evict them from a fixed recency window
  • Agent-Initiated Rubric-Gated Self-Compaction — Give the agent both a compaction tool and a firing rubric so it compacts on trajectory structure — sub-task resolved or converging — instead of a fixed token threshold
  • Reasoning Retention and Compaction as Harness Settings — Whether the harness passes private reasoning back across tool calls and compacts instead of truncating changes measured agent performance, so a benchmark score is a model-plus-harness measurement
  • Shortening Old Tool Results Under Context Pressure (Half-Life Truncation) — Mechanically shortening older tool results as the window fills raises coding-agent fix rates under a tight window, and converges with the untouched transcript once the window is wide
  • Stage Elision Before Summarization — Fire deterministic elision at a soft context threshold and the summarizer only at a hard one; across 176 matched coding-agent settings the staged policy held success rates at the lowest cost
  • Version-Controlled Agent Context — Commit, branch, merge, and read agent memory as a durable file tree so the abstraction level is chosen at read time; the retrieval path carries the benefit, not the checkpoints
  • Elastic Context Orchestration — A per-turn vocabulary of context operations — Skip, Compress, Snippet, Rollback, Delete — that lets long-horizon search agents tier retention by current task relevance instead of accumulating raw trajectory
  • Prompt Compression — Write instructions that convey the same guidance in fewer words; shorter, denser instructions improve agent compliance and reduce token cost
  • Choosing a Compression Budget for Agent Control Context — Set how far you compress an agent's always-loaded instructions from environment-verified task success rather than a token-reduction ratio, and keep semantic sections whole
  • Measuring Reacquisition Cost Under Context Compaction — Split the agent's tool calls into retrieval and execution to expose the interaction cost a flat task-completion rate hides, and confirm the cause by restoring one class of dropped state
  • Injected-Context Cost Attribution in Agent Workflows — Tokenize each node's prompt twice, with and without retrieved memory, to split billed input tokens into what the node carries and what it was handed
  • Four Reporting Levels for Agent Working Memory Evaluation — Two memory policies at the same token cap deliver different context and spend different management work, so report stored state, delivered context, and management work alongside the outcome
  • Prompt Caching: Architectural Discipline for Agents — Treat prompt caching as a structural constraint on prompt composition, with cross-provider economics and extended-TTL guidance folded in
  • Static Content First for Cache Hits — Place static content at the beginning of the prompt and variable content at the end to maximize prompt cache hits and keep inference costs linear
  • Prompt Cache Keepalive for Agent Pauses — Replay a cached prefix on a timer so it survives tool runs and approval waits, and the billing, pause-length, and interval conditions that decide whether it saves money or costs 4x more
  • Mask Tools Instead of Removing Them — Restrict which tools an agent may call through a per-request field instead of editing the tools array, so a changing action space costs no cache write — and when removing them is still cheaper
  • Stateful Iteration State-Carry — Carry typed persistent state across long agent loops through a state-read tool instead of replaying the full transcript each turn; converts O(n²) total token cost to O(n) when loops are long and observations are large
  • Exclude Dynamic System Prompt Sections for Cross-Machine Cache Sharing — Move per-machine context (cwd, OS, shell, memory paths) out of the Claude Code system prompt so identical fleet configurations share one prompt-cache entry across users and machines
  • KV Cache Invalidation in Local Inference — When Claude Code prepends an attribution header to prompts sent to local models, it invalidates the KV cache on every request and causes ~90% slower inference
  • Semantic Density Optimization — Maximize task-relevant tokens in a codebase by eliminating zero-information ceremony while preserving naming, documentation, and commit context that agents cannot reconstruct without inference cost
  • Validating Token-Optimized Formats Inside Agentic Loops — Switching tool schemas from JSON to TOON or TRON saves up to 27% tokens but regresses accuracy by 9-14 percentage points in end-to-end agentic loops; input-side and output-side compression carry different risk
  • Source Code Minification for State-in-Context Agents — Stripping comments, whitespace, and shortening identifiers cuts input tokens 42% but drops SWE-bench Verified resolution rate from 50% to 38% — apply only when measured savings beat the accuracy cost
  • Cross-Lingual Prompt Preprocessing (Local-LLM Token Arbitrage) — A local small model translates non-English prompts to English and rewrites them into compact task-oriented form before send; cuts input tokens 34–47% only when latency, accuracy, and fidelity costs do not erase the savings

Assembly & Composition

How to build, layer, and route context to the right agent at the right time rather than dumping everything into a single prompt.

  • Dynamic System Prompt Composition — Build system prompts from modular, priority-ordered sections rather than monolithic static text, enabling mode-specific variants and efficient API caching
  • Narrative Problem Reformulation for Code Generation — Rewriting a fragmented coding problem as a coherent three-part narrative measurably shifts which algorithms a code LLM selects, with reported 18.7% zero-shot pass@10 gains concentrated on harder competitive-programming tasks
  • Phase-Specific Context Assembly — Optimize the orchestration layer that prepares each agent per phase; planners get summaries, workers get targeted file excerpts and validation commands
  • Prompt Chaining — Decompose a complex task into a sequence of LLM calls where each step processes the output of the previous one, enabling verification and gate-checking at each stage
  • Prompt Layering — Agent instructions arrive from multiple sources simultaneously; understanding the precedence order and conflict resolution prevents unpredictable behavior
  • Filter and Aggregate in the Execution Environment — Run data processing logic inside the code execution sandbox before surfacing results to the model, so only the relevant subset of data enters context
  • Evolving Playbooks — Replace monolithic prompt rewrites with structured delta entries that accumulate, refine, and organize agent strategies without losing domain knowledge
  • Typed Context Buys Addressability, Not Token Savings — Schema-bearing context records cost more tokens than the prose they replace; adopt them when a program must prune, validate, or gate the context before the model reads it

Loading & Retrieval

Techniques for getting the right context into an agent on demand, whether from code repositories, APIs, or structured knowledge bases.

  • Context Hub — Fetch current, versioned API documentation into agent context at generation time so agents write against the live spec rather than stale training-data snapshots
  • Codebase-Derived Pattern Libraries as Agent Context — Mine your own repositories for proven implementations and serve them to an agent as intent-searchable context instead of generic public examples
  • Retrieval-Augmented Agent Workflows — Pull context into the agent at the moment it is needed rather than preloading it at session start
  • Query-Conditioned Reuse of Retrieved Agent Trajectories — Hand the agent a four-field target-bound note instead of the retrieved trajectory, so the procedure transfers and the source run's stale values do not
  • Live Browser as Agent Context Channel — Subscribe an agent to the developer's running browser tabs as live context — lower friction than copy-paste, but the developer's logged-in session enters the indirect-injection blast radius
  • App-Window Snapshot as Agent Context — Bind one hotkey to send the active app window — rendered screenshot plus accessibility-tree text — as a single context unit; the richer payload changes which cross-app handoffs are plausible to delegate
  • Reproducibility Artifacts as Agent Context — Tests, commit history, layout, instructions, and decision records are agent context, but each loads through a different channel and only the non-discoverable slice belongs in always-on instructions
  • Repository Map Pattern — Parse source files with tree-sitter to extract structural symbols, rank them by graph importance, then binary-search fit the most relevant entries into the agent's available token budget
  • Deterministic Anchoring — Inject call-graph, inheritance, and config-dependency facts as plain-text comments so code-agent navigation converges run-to-run; the win is reproducibility, not capability
  • Context Compiler — Compile a task-scoped payload — full target file, skeletonized dependencies, everything else excluded — instead of buying a bigger window; the dependable payoff is reproducibility, not accuracy
  • Semantic Context Loading — Query codebases through Language Server Protocol semantics — symbol lookup, reference finding, type navigation — rather than reading raw files
  • Symbol Ranking for Agent File Pickers — Rank @ mention candidates by the symbols each file defines rather than by filename similarity alone; the win depends on repository size, index coverage, and whether you select files by hand at all
  • Seeding Agent Context — Strategically place files, comments, and markers that agents discover during exploration and use to shape their behavior
  • Grounding Agents in Code the Model Has Never Seen — When the model has no training signal for a proprietary SDK or custom framework, it generates against the closest public API in training; provisioning must displace that prior, not just supplement it
  • Environment Specification as Context — Feed dependency versions, lock files, and runtime constraints into agent context to prevent the 50–70% accuracy drop caused by environment-blind code generation
  • Runtime Resource Limits as Prompt Context — State the memory ceiling and time budget the generated code must run under; peak memory fell in 13 of 14 comparisons, but disclosure shifts the distribution rather than guaranteeing the budget
  • Repository-Level Retrieval for Code Generation — AI coding agents that retrieve cross-file context from dependency graphs, ASTs, and semantic embeddings generate more accurate code than those limited to local file context
  • Give the Model the Target's Contract, Not Similar Solutions — On repository-level tasks, a target function's own test and required helpers beat a retrieved solution to a similar problem
  • Agent-Tuned Code Search — A search tool that runs its own loop and returns file paths plus line ranges cuts search latency sharply, but the token win against local grep is modest and bounded by index freshness
  • Silent Handoff Failure in Delegated Code Search — For read-only questions over an indexable repository, a pre-built vector index beat a search subagent 65.2% to 46.2%, and 41.8% of the subagent's failures broke silently at the planner handoff
  • AOCI: Symbolic-Semantic Repository Indexing — A persistent, query-independent blueprint pairing architectural coordinates with semantic content — read whole before any task, distinct from on-demand retrieval and token-fitted repo maps
  • Structured Domain Retrieval — Combine hierarchical knowledge graphs with coverage-driven case selection to retrieve domain-specific context that flat vector search misses
  • Schema-Guided Graph Retrieval — Use one shared domain schema across graph construction, query decomposition, and typed retrieval to improve multi-hop reasoning precision over private knowledge bases
  • Fact Supersession Memory for Code Assistants — Key each remembered fact by subject and relation and retire the old row when the value changes; covers the 18.4% of real fixes that reduce to one value swapping under a stable key
  • Claim-Scoped Invalidation for Agent Memory — Ask whether one stored claim still holds after a commit rather than whether the commit preserved behavior; on identical diffs the question moves precision further than extra evidence or a bigger model does
  • Budgeted Verification of Inherited Agent Constraints — A constraint that reads as settled is the memory an agent never re-checks; the limit-preferring allocation rule recovers it, but only where the verification budget is genuinely scarce
  • Cross-Reference Dereference Hop in Retrieval Loops — A retrieved passage that says "see Section 7.2" ranks first and answers nothing; a one-hop resolution pass fixes it where references resolve to a key and the corpus churns faster than it is queried
  • Chunking Strategy for RAG-Based Code Completion — Function-based chunking is dominated by every other strategy on line-level code completion; Sliding Window and cAST sit on the Pareto frontier, and doubling cross-file context length matters more than chunking choice
  • Component-Wise RAG Prioritization — A 21+ model component-wise empirical study finds retriever choice dominates generator choice for SE-task RAG, and BM25 is robust across code generation, summarization, and repair — under specific conditions
  • Corpus Shape as a Retrieval Design Constraint — Three yes/no questions about a document collection separate the retrieval failures a re-ranker can fix from the recall ceilings that sit above tuning entirely
  • LLM-Driven Logical Retrieval — When the agent LLM is frontier-capable, letting it emit AND/OR/NOT Boolean queries against an inverted index matches an agentic hybrid baseline at 41× lower indexing cost — under specific lexical-overlap conditions
  • Exhaustive Retrieval for Listing Questions — When the correct answer is every matching passage, ranked top-k truncates silently; the completeness signal you hold decides whether an enumeration loop earns its cost
  • Gate Generation on Retrieval Sufficiency, Not Model Confidence — On a private codebase the model has no prior to be uncertain against, so a pre-generation gate has to score retrieval coverage from the repository; the coverage price and the compiler alternative decide whether it is worth building
  • State-Conditioned Evidence Selection for Mid-Task Retrieval — Subtract the source units the agent has already read, then score what remains as a combination; the payoff concentrates on decisions needing three or more facts together, and the controller costs tokens
  • Long Context vs Retrieval: The Break-Even Decision — Feeding a whole corpus into a million-token window buys answer completeness at roughly sixteen times the per-query cost, so corpus size and query volume decide, not token price
  • Hypothetical Classification for Large Label Vocabularies — Let the model invent a label that does not exist, then resolve it to the nearest real one by embedding; the unguided form lost to plain retrieval on both published benchmarks
  • Compositional Skill Routing — Decompose a query into atomic sub-tasks, retrieve one skill per sub-task, then compose the plan — earns its cost only above hundreds of skills, where decomposition quality caps the system
  • Skill Loadout Curation for Coding Agents — Curate the skills and MCP servers an agent loads before a session; above roughly 30 skills the gain comes from removing colliding descriptions, not from saving tokens
  • Choosing a Skill Loading Method for Agents — Preload, on-demand blocks, a reference catalog, or a hybrid of stubs and fetches; skill size and per-turn need decide which wins, measured on cache-corrected input
  • When a Skill Graph Cannot Beat the Ranker — A typed skill graph whose candidate edges come from the retriever's own embedding neighbors can relabel what retrieval found but cannot extend its reach; overlay the edges and count what they newly connect
  • Coverage-Aware Skill Selection Under a Token Budget — Score candidate skills against the capabilities the selected set still leaves uncovered rather than against the query; the fit needs pass/fail logs and does not transfer across libraries or executors
  • Per-Object Context Allocation — Group retrieved fragments by canonical code object and cap each object's token share, so repeated views of one function stop displacing independent dependencies before the model sees them
  • Organizational Context Layer for Agents — A shared context layer maps many source systems into one governed store that many agents query; deterministic identity and rebuildable index projections transfer well below that scale
  • Generated Questionnaires: Eliciting Someone Else's Context — Have the agent draft a questionnaire for the one person who holds the missing knowledge, instead of guessing the gap; the technique holds only where that knowledge is statable rather than tacit

Error Handling & Drift Prevention

Keeping agents on track across long sessions by preserving failure signals and reinforcing goals.

  • Context-Injected Error Recovery — When a tool call fails, inject structured error context into the next inference call to prevent retry loops before they form
  • Error Preservation in Context — Keep failed actions and error traces visible in the agent's context window; error history acts as negative examples that shift model behavior
  • Goal Recitation — Periodically rewrite objectives, to-do lists, and status summaries at the tail of context to exploit recency bias and prevent goal drift in long-running sessions
  • Trajectory Attribution for Context Repair (TRACE) — Mine stored trajectories for implicit dissatisfaction cues to attribute a failure to the skill, knowledge-base entry, or tool description that caused it, then propose the edit