
Context Engineering Explained: How to Give AI Models the Right Information
In 2025, Shopify’s CEO Tobi Lütke posted a short tweet that ended up reshaping how serious AI practitioners talk about their own work. He said he preferred the term “context engineering” over “prompt engineering” — not because it sounded better, but because it described the actual skill more honestly: providing everything a model needs to make a task solvable, not just phrasing a clever question.
Andrej Karpathy amplified it within days. Independent AI commentator Simon Willison predicted the term would “have sticking power.” They were right. By 2026, Anthropic had published its own engineering guide treating context engineering as the natural next stage after prompt engineering, and production AI teams had quietly stopped obsessing over magic phrasings altogether.
This guide explains what actually changed, why the research says “just add more context” is often the wrong move, and how to think about the problem in a way that holds up whether you’re chatting with an AI casually or building a production agent.
None of this requires a computer science background to follow. The underlying research is technical; the practical takeaway isn’t. By the end of this guide, you’ll have a concrete, three-question test you can run against any AI system you use or build, plus a map of the specific techniques and dedicated guides that go deeper on each piece.
What This Article Covers
This piece covers what context engineering actually is, why it replaced prompt engineering as the skill that matters most for reliable AI output, the real research on why more context can make a model worse, and a practical framework for deciding what belongs in a model’s context and what doesn’t.
It does not re-explain what a context window is mechanically — this site’s existing guide on how AI memory works already covers that distinction between context windows and long-term memory. It also doesn’t re-explain the Model Context Protocol, covered in its own dedicated article, or basic prompt-wording tips, covered in this site’s earlier prompting guides.
This piece assumes you’re comfortable with the basic idea of talking to an AI model and want to understand the layer above that: not what to type, but what surrounds what you type, and why that surrounding layer increasingly determines whether the AI actually succeeds.
Five dedicated guides go deeper on individual pieces introduced here at overview level: few-shot versus zero-shot prompting, writing system instructions, structured/JSON prompting, role prompting, and compressing long documents without losing what matters. Three more compare the tools and models practitioners actually use to do this work: prompt libraries and marketplaces, prompt management tools and ChatGPT, Claude and Gemini for prompt-heavy workflows. Treat this guide as the map; those guides are the terrain.
A helpful mental model, popularized by Karpathy, treats the whole setup like a small computer: the model is the processor, the context window is RAM or short-term memory, and whoever builds the system is the operating system deciding what gets loaded at each step. That framing is worth keeping in mind through the rest of this guide, since every technique below is really just a more disciplined answer to “what does the operating system load, and when.”
Why “Context Engineering” Replaced “Prompt Engineering”
Prompt engineering asks one question: how do you word an instruction so a model understands it correctly? For years, that was genuinely the highest-leverage skill in working with AI, because early use cases were mostly single questions with single answers.
Context engineering asks a bigger question: what does the model need to know, in what form, before it can even attempt the task? Anthropic’s own engineering team frames this as the natural progression of prompt engineering — as AI moved from one-shot chat answers to multi-step agents handling long, complex work, the challenge shifted from wording a good instruction to assembling the right configuration of information at every step.
Gartner has since published a formal definition treating context engineering as designing and structuring the relevant data, workflows, and environment so an AI system can understand intent and produce enterprise-aligned outcomes — without leaning on someone hand-crafting the perfect prompt each time. That’s a meaningful shift: from a skill one clever person could deploy, to an engineering discipline a whole system can be built around.
The distinction is easiest to see side by side. A prompt engineer writes careful instructions: be concise, use a professional tone, format as a bulleted list. A context engineer makes sure the model also has the specific data, prior conversation, and tool access it needs — the instructions were never the part that was missing.
How Models Actually Use Context
A large language model doesn’t read your context the way a person reads a document — top to bottom, building understanding as it goes. It processes every token in its context window through an attention mechanism that relates each token to every other token at once.
That matters because attention is a finite, competed-for resource, not an unlimited spotlight. Anthropic describes this as an “attention budget”: transformer architecture creates a pairwise relationship between every token and every other token, so as context grows, the model’s effective ability to weigh any one piece of it shrinks.
This is the real, technical reason “just give it more information” isn’t automatically a good strategy. Every token you add doesn’t just cost you money — it competes with everything else already in the window for the model’s limited attention, whether that new token turns out to matter or not.
It’s worth being precise about what this does and doesn’t mean. A model isn’t literally running out of room the way a hard drive fills up; it’s that the quality of attention spread across every token declines as the token count grows, which is a subtler and, in practice, more common failure than simply overflowing a window.
The Real Problem: Context Rot
For a while, the AI industry’s working assumption was that context windows would just keep growing — a million tokens, then more — until stuffing everything relevant into one giant prompt made careful curation unnecessary. Research published by Chroma directly contradicts that assumption.
Chroma’s researchers tested 18 frontier models, including versions of GPT, Claude, and Gemini, and found every single one degraded in accuracy and reliability as input length grew — even when every added token was genuinely relevant and nowhere near the model’s actual context limit. They named this pattern context rot, and it’s distinct from simply running out of room: rot sets in well before a window actually overflows.
This isn’t a new observation dressed up in a new name. Earlier academic research, often cited as the “lost in the middle” finding, documented a U-shaped attention bias: models reliably use information placed near the start or end of a long context, and lose track of the same information when it’s buried in the middle. One frequently-cited result tested 20 retrieved documents totaling roughly 4,000 tokens and found accuracy fell from 70-75 percent, when the needed fact sat at position 1 or 20, down to 55-60 percent when it sat in the middle.
Chroma’s research adds a sharper edge to that finding: models don’t just struggle with volume, they struggle specifically with distractors — information that’s semantically similar to the right answer but subtly wrong. A pile of obviously irrelevant text is easier for a model to ignore than one plausible-sounding wrong answer sitting right next to the correct one.
Anthropic’s own vocabulary for this is worth knowing, since it shows up across the industry now: context poisoning (a hallucination or error early in a session gets treated as fact and compounds later), distraction (too much accumulated history crowds out the actual current task), confusion (irrelevant tools or data change how a model interprets a request), and clash (different pieces of context actively contradict each other). Each is a slightly different failure, but all four trace back to the same root cause this guide keeps returning to: an unmanaged attention budget.
This is also the honest rebuttal to a real industry debate worth naming directly: some practitioners argued that million-token context windows made retrieval-augmented generation obsolete, since you could supposedly just paste everything in. Context rot research is the reason that argument didn’t hold up — selective, well-curated retrieval still outperforms dumping everything into a huge window, which is exactly why our RAG guides remain as relevant now as when context windows were far smaller.

Where This Shows Up Most: Coding Agents
No domain makes context engineering more concrete than autonomous coding agents, because the failure mode is immediate and visible: the agent starts making worse decisions the longer it runs, not because the underlying model got dumber, but because its own accumulated context did.
The current generation of coding agents makes this trade-off explicit through how long they’ll run unattended. Cognition Labs’ Devin is described in its own documentation as an autonomous AI software engineer that can write, run and test code, with example sessions that run from prompt to pull request. Cursor’s Agent Mode, by contrast, runs inside an editor where the developer stays close to the loop, which limits how much context can accumulate before a human checkpoint resets it.
Anthropic’s own Claude Code sits at a third point on that same spectrum: conversational and tool-using, with the developer actively driving each step rather than letting context accumulate across a long unattended run. Teams adopting Devin are explicitly advised to budget real time for reviewing its output rather than trusting it as production-ready, precisely because errors compound without a human checkpoint — which is context rot’s practical consequence, stated as a product warning rather than a research finding.
The pattern that emerges across all three tools is the same one this guide’s Budget Test is built around: more autonomy means more accumulated context between human checkpoints, which means more exposure to rot unless something — compaction, structured memory, a shorter autonomy window — is actively managing it.
None of this is unique to coding specifically. The same pattern shows up in any long-running agent: a research assistant working through many search steps, a customer-support agent handling an extended back-and-forth, or a planning agent coordinating several tools across a multi-stage task. Coding agents simply make the failure visible fastest, because a broken pull request is an obvious, checkable outcome in a way a subtly worse customer support answer often isn’t.
A Short History: From Chatbots to Context Problems
Early consumer chatbots didn’t have a context engineering problem worth naming, because they mostly didn’t need one. A single question, answered from the model’s training knowledge, with no tools, no retrieved documents, and no multi-step task to sustain, simply doesn’t generate enough accumulated information for position or decay to matter.
The problem became unavoidable once AI systems started doing more than answering isolated questions: calling tools, retrieving documents, remembering earlier turns, and working across many steps toward a single goal. Each of those additions is, structurally, another source of tokens competing for the same finite attention budget this guide described earlier.
That’s the real timeline behind the 2025 terminology shift. Prompt engineering was sufficient for a world of one-shot answers. Context engineering became necessary the moment AI products started being agents rather than chatbots — which is also, not coincidentally, when Tobi Lütke’s tweet found an audience that had already started feeling the problem without a name for it yet.
It’s worth naming what changed on the research side too. “Lost in the middle” predates the term “context engineering” by roughly two years — the underlying problem was documented before the industry had a name for the discipline meant to address it, which is a useful reminder that the term caught up to a real, already-measured phenomenon rather than inventing a new one.
The AI Hustle World Context Budget Test
Given a finite, competed-for attention budget and evidence that position and relevance both matter, it helps to have a repeatable way to decide what earns a place in a model’s context. We call this the Context Budget Test, and it’s three checks, not a checklist to blindly maximize.
The Necessity Check: does the model need this right now, or could it be fetched just-in-time only if the task actually requires it? Anthropic’s own guidance favors just-in-time retrieval over loading everything upfront specifically because unused context still costs attention budget even when the model never needed it.
In practice, this check is easiest to apply as a simple question asked of every piece of context before it goes in: if this turned out to be unnecessary, would removing it have changed anything about how the model needed to behave? If the honest answer is no, it was a candidate for just-in-time retrieval instead of upfront inclusion.
The Position Check: if something is genuinely critical, is it placed where a model will actually attend to it — near the start or end of the context — rather than buried in the middle of a long block of retrieved text or conversation history?
This check matters most for anything a system genuinely cannot afford to have missed — a hard safety constraint, a non-negotiable business rule, the single most important fact in a retrieved document set. Repeating that specific piece near the end of the context, immediately before the model generates its response, is a simple, low-effort mitigation directly grounded in the lost-in-the-middle research above.
The Decay Check: as a task continues across many turns, does this specific piece of context still deserve its spot, or has it become stale, already-used, or safely summarizable? Long-running agents that never revisit this question are the ones most exposed to context rot in practice.
A practical version of this check: at regular intervals in a long session, ask whether each major piece of accumulated context would still be included if you were starting the task fresh right now, knowing everything you know at this point. Anything that fails that test is a compaction or removal candidate, not a permanent fixture.
| Check | Question | If it fails |
| Necessity | Does the model need this now, or only if the task requires it? | Wasted attention budget on information that never gets used |
| Position | Is critical information placed where the model actually attends? | The exact ‘lost in the middle’ failure this guide describes |
| Decay | Does this still deserve its spot as the task continues? | Context rot compounds the longer a session runs |

The Five Components of Context
Whatever specific technique you use, almost every context-engineering decision is really about one of five components. We have dedicated guides going deeper on several of them — treat this section as the map, not the destination.
System instructions set the behavioral rules the model operates under for the entire task — not just “be helpful,” but specific tone, constraints, and boundaries. Our guide on writing system instructions for consistent outputs goes deep on getting this layer right.
Weak system instructions tend to fail in a specific, recognizable way: they work in the first exchange and quietly stop being followed several turns later, once other context has crowded in. That’s not the model forgetting the rule so much as the rule losing its share of a shrinking attention budget — another reason position and repetition both matter more than people expect.
The immediate request is the specific thing being asked right now. This is where classic prompt engineering still lives, and where techniques like few-shot versus zero-shot prompting and role prompting — both covered in our dedicated guides — make their biggest difference.
This is also the component most people already have some intuition for, since it’s the part of context engineering that looks most like traditional prompting. The mistake is assuming it’s the only component that matters, when the research this guide covers shows the other four are frequently what determines whether a well-worded request actually succeeds.
Retrieved or grounding data is information pulled in from outside the conversation — documents, database records, search results — most commonly through retrieval-augmented generation, already covered in depth in our RAG guides. Tool and function definitions tell a model what actions it can take and how to take them, frequently delivered through the Model Context Protocol, which this site covers in its own dedicated explainer rather than repeating here.
A detail worth flagging because it surprises people: a tool the model never ends up calling still costs context budget just by being defined and available. Loading every tool a system might conceivably need, rather than the ones a specific task actually requires, is a quiet, easy-to-miss version of the Necessity Check failing before the conversation has even started.
Conversation history and memory carry forward what’s already happened in a session or across sessions. This site’s existing guide on how AI memory works covers the distinction between a context window and true long-term memory in detail.
Notice that all five compete for the same finite attention budget described earlier — adding more of one (a longer conversation history, say) doesn’t just cost tokens, it reduces the effective share of attention available to the other four. That’s the real reason this is a curation problem and not just an inventory-management one.
The Tools and Techniques Practitioners Actually Use
Beyond the five components above, a specific set of named techniques has emerged for actually managing context well in production. Each earns its own explanation rather than a passing mention, since each solves a genuinely different piece of the problem.
A note on how this section was researched: the specific figures and techniques below come from Anthropic’s own published engineering guidance and independent reporting on production coding agents, not from us running these systems ourselves. Where a number is a company’s own reported result rather than independently audited, that’s flagged explicitly.

Anthropic’s Context Editing and Memory Tool
Anthropic ships two specific, named features aimed directly at context rot: Context Editing, which prunes or compresses accumulated context as a task runs long, and a Memory Tool, which lets an agent read, write, and maintain its own files across a session rather than keeping everything crammed into the live conversation.
In Anthropic’s own internal testing, combining the Memory Tool with Context Editing improved agent-based search performance by 39 percent; context editing alone produced a 29 percent improvement on its own. In a 100-round web search task, the same combination reportedly cut token consumption by 84 percent. These are the company’s own reported figures, not independently audited, but they’re a concrete illustration of how much headroom exists when context is actively managed rather than left to accumulate.
The design logic behind both features maps directly onto this guide’s Budget Test: Context Editing is essentially an automated Decay Check running continuously through a session, while the Memory Tool gives an agent a place to put information that’s still useful but doesn’t need to occupy the live conversation’s attention budget every single turn.
Model Context Protocol (MCP)
MCP standardizes how a model discovers and calls external tools — search, code execution, databases — which matters for context engineering specifically because every tool definition a model has access to consumes context budget, whether or not that tool gets used in a given turn. This site’s dedicated MCP explainer covers the protocol itself in depth; the context-engineering angle is simply that tool definitions are context too, and unused ones aren’t free.
Retrieval-Augmented Generation (RAG)
RAG is the most established technique for solving the “don’t load everything upfront” problem: instead of pasting an entire knowledge base into context, a system retrieves only the specific passages relevant to the current query. Our RAG guides cover the full mechanism, chunking strategy, and vector-database landscape; the context-engineering framing is that RAG is really just the Necessity Check from this guide’s Budget Test, implemented as infrastructure.
This reframing matters for anyone who’s already read this site’s RAG content: none of that knowledge becomes obsolete here. RAG is one specific, mature implementation of a much broader curation principle, and understanding the principle makes the specific technique easier to apply well rather than by rote.
Compaction and Structured Note-Taking
Compaction means periodically summarizing or compressing what’s accumulated in a long conversation so far, rather than letting raw history grow unchecked. Structured note-taking is a closely related technique where an agent maintains its own running notes or state file, similar in spirit to how a person keeps a project log instead of trying to remember every detail of a long project in their head.
Both techniques exist for the same reason: raw conversation history is an inefficient way to preserve what actually matters from a long session. A compacted summary or a structured note takes up a fraction of the tokens a full transcript would, while preserving the specific facts a task actually still needs going forward.
Multi-Agent Architectures
Rather than asking one agent to hold an entire complex task’s context at once, some systems split the work across multiple agents, each with its own narrower, more focused context. This trades coordination complexity for individually cleaner context per agent — a direct, structural response to the attention-budget problem this guide opened with.
This is also why some coding-agent workflows now explicitly combine tools rather than picking one: a fast, editor-native agent for active development paired with a slower, fully autonomous agent for well-scoped backlog work is really a multi-agent architecture at the product level, each half deliberately carrying a narrower context than a single tool trying to do both jobs at once.
What This Actually Costs
Context isn’t just an accuracy problem — it’s a direct line item. Every token in a model’s context window is a token you pay for, whether or not it ends up mattering to the model’s answer.
Karpathy has pointed out that roughly 90 percent of a typical AI coding bill can be spent on unnecessary context, with meaningful savings available just from optimizing what gets included and how tool calls are routed. That’s a striking number precisely because it implies most teams aren’t paying for intelligence — they’re paying for waste.
Anthropic’s own reported 84 percent token-consumption drop, from combining Context Editing with its Memory Tool over a 100-round web search task, is the concrete counter-example: that headroom is real and addressable, not just a theoretical inefficiency.
The honest framing: cost and accuracy aren’t two separate problems here, they’re the same problem viewed from two angles. Context that’s poorly curated is both more expensive and less reliable at the same time, which is a rare case where doing the right thing and doing the cheap thing point in exactly the same direction.
This is also why the Necessity Check from the Budget Test above pays for itself twice: every piece of context it correctly excludes is both a token cost avoided and an attention-budget cost avoided, and those two savings compound rather than trade off against each other.
Common Mistakes to Avoid
Treating a bigger context window as a substitute for curation is the most common and most expensive mistake — the context rot research is explicit that this doesn’t hold, regardless of how large a model’s advertised window is.
Ignoring where information sits in the context, not just whether it’s there, is a close second. A fact your system genuinely needs is only as useful as the model’s actual ability to attend to it, and burying it in the middle of a long block of retrieved text or history is a self-inflicted version of the lost-in-the-middle problem.
Never revisiting accumulated context in a long-running agent is a third — conversation history and tool outputs that made sense five steps ago can become noise by step fifty, and treating context as something you only add to, never prune, guarantees rot over a long enough session. Treating every tool or data source as free to include “just in case” is a fourth: unused tool definitions and unreferenced documents still occupy attention budget and cost tokens, whether the model ever touches them or not.
Assuming this only matters for exotic, agentic use cases is a fifth mistake worth naming — even a single well-designed customer support response benefits from the same Necessity and Position thinking this guide describes, just at a smaller scale than a long-running coding agent. Confusing a large advertised context window with a solved problem is a sixth: a model that can technically accept a million tokens hasn’t been shown to use all of them equally well, and the context rot research means the marketing number and the practical, reliable capacity are two different things.

Who Actually Needs to Practice This
Casual, single-question chat use barely needs any of this — asking an AI model one clear question with the information already in front of it doesn’t require curation, and prompt-wording tips still matter more at that scale.
A useful rule of thumb: if you could screenshot the entire relevant context for a task in one image, you probably don’t have a context engineering problem yet. The discipline this guide describes earns its keep specifically once a task’s relevant information stops fitting comfortably in view all at once.
Anyone building an AI feature that runs multiple steps, calls tools, retrieves documents, or maintains a conversation across many turns needs this discipline specifically, because that’s exactly where context rot, cost waste, and lost-in-the-middle failures actually show up in practice. Teams shipping AI agents into production need it most of all — the gap between a demo that works on a clean, short example and a production system running hundreds of real, messy, long-running sessions is almost entirely a context engineering gap, not a model capability gap.
There’s a useful test for where you personally fall on this spectrum: if you’ve ever had an AI system that worked great in testing and got noticeably worse the longer a real user’s session ran, you’ve already experienced the exact problem this discipline exists to prevent, whether or not you had a name for it at the time.
Why Simple Prompting Still Has a Place
None of this makes prompt wording obsolete, and it’s worth being specific about why. A single, well-scoped question with all its relevant information already visible in a short conversation doesn’t have a context-curation problem to solve — there’s nothing to prune, reorder, or compact.
The realistic split is by task shape, not by skill level: a one-off question benefits most from a clearly worded prompt; a multi-step agent working over many turns benefits most from the context discipline this guide describes. Most real AI use sits somewhere between those two poles, which is exactly why this guide covers both prompt-level techniques (few-shot, role prompting, structured output) and the broader context discipline in the same place.
It’s worth naming the coding-agent example from earlier one more time here, since it makes the split concrete: Cursor’s roughly 20-minute checkpoint isn’t a limitation, it’s a working example of exactly this trade-off being made deliberately — short enough that prompt-level clarity still carries most of the weight, long enough to get real work done between resets. There’s also a cost argument for keeping things simple when the task allows it: building retrieval, memory, and compaction infrastructure for a task that was always going to be a single clean question is its own kind of waste — the Necessity Check cuts both ways, and sometimes the answer is that a whole context-engineering system isn’t necessary either.
What Happens If You Ignore This
Skipping context engineering doesn’t usually break a demo — demos are short, clean, and exactly the conditions under which context rot hasn’t had a chance to set in yet. It breaks in production, quietly, exactly when a system has been running long enough for accumulated, poorly-curated context to start degrading answers.
The compounding cost is the one teams notice too late: a support agent that answers well for the first ten exchanges and gets noticeably worse by exchange fifty isn’t experiencing a model limitation — it’s experiencing the accumulated result of never revisiting what’s actually still useful in its own context. There’s a second, quieter cost worth naming: teams who never adopt this discipline tend to respond to unreliable output by switching models or upgrading to a bigger context window, rather than fixing the actual curation problem — which means they pay more for infrastructure that was never going to fix the thing that was actually wrong.
How to Know If Your Context Engineering Is Working
Track task success rate specifically across long sessions, not just on short, clean test cases — the whole point of context rot is that failures concentrate later in a session, so testing only short interactions will systematically miss the problem. Track token cost per completed task, not just per API call — Karpathy’s 90 percent figure is a per-task waste number, and it only shows up if you’re measuring efficiency at the level of the actual work getting done, not just raw usage.
Run your own Budget Test audit periodically on a long-running agent’s actual context: what’s in there right now, does it still pass Necessity, is anything critical buried in the middle, and what’s stale enough to prune. This is the same three-check framework from earlier in this guide, just applied to your own system instead of a new one you’re evaluating.
Watch for the specific symptom pattern researchers describe as attention dilution: a long-running agent’s tool choices start to drift, or it stops following an instruction it was clearly given earlier in the session. That’s usually a decay problem, not a capability problem, and it points directly back to the Decay Check rather than to a need for a more capable model.

What’s Next
The clearest trend to watch is automation of the curation work itself — Anthropic’s Context Editing and Memory Tool are early examples of the model’s own tooling handling compaction and relevance decisions that a human engineer used to have to hand-code. A second-order effect worth watching is on hiring and job titles: “context engineer” and similar framings are already appearing as the practitioner community’s answer to “prompt engineer,” and the skill being valued has shifted from clever wording to systems thinking about information flow.
This shift is already visible in how job postings for AI-adjacent roles are worded: fewer ask for “prompt writing” specifically, and more ask for experience with retrieval systems, agent memory, and evaluation — the practical building blocks of context engineering, even when the posting doesn’t use the term itself yet. A related labor-market effect is on what “AI experience” even means on a resume: a year of writing clever prompts is a very different, and increasingly less valuable, credential than a year of shipping production agents that stay reliable across hundreds of long-running sessions.
A third is on AI product design generally: as more builders internalize that bigger context windows don’t solve reliability on their own, expect more products to be explicitly designed around retrieval, memory, and compaction from day one, rather than treating a large context window as a shortcut around that design work. A fourth thread worth watching: as evaluation tooling matures, expect “context quality” metrics — not just accuracy, but measures of how efficiently a system uses its attention budget — to become as standard in AI engineering as latency and cost metrics already are.
A fifth, tied directly to the coding-agent example above: expect more products to make autonomy duration itself a user-facing setting, the way Cursor’s roughly 20-minute checkpoint and Devin’s 60-plus-minute runs already represent two different, deliberate answers to the same underlying context-accumulation trade-off.
Final Thoughts
Tobi Lütke’s original point wasn’t really about terminology. It was that the actual skill in working with AI was never just finding the magic words — it was making sure the model had what it needed to succeed, in a form it could actually use.
The research since then has only sharpened that point. Every frontier model tested degrades as context grows, position matters as much as presence, and the fix isn’t a bigger window — it’s the discipline of deciding what earns a place in that window, where it sits, and how long it stays. That discipline is what this guide is built to teach, one component at a time.
The five components, the named techniques, and the Budget Test in this guide are the map. Five dedicated guides — on few-shot versus zero-shot prompting, system instructions, structured prompts, role prompting and document compression — give each piece of that map its own detailed treatment.
Read together, this guide and those five guides are meant to leave you able to answer a single practical question for any AI system you touch: given everything this model could see right now, does it actually have what it needs — no more, no less — to do this specific job well?
Ready to Go Deeper on One Component?
This guide mapped the whole discipline. Our guide on writing system instructions goes deep on the one component every AI interaction actually starts with.
Read the System Instructions Guide →Frequently Asked Questions
What is context engineering?
Context engineering is the discipline of curating everything an AI model sees before it answers — system instructions, retrieved data, tool definitions, examples, and conversation history — rather than just wording a single instruction well. It’s the practical answer to why the same underlying model can perform brilliantly in one setup and unreliably in another.
How is context engineering different from prompt engineering?
Prompt engineering focuses on wording a single instruction effectively. Context engineering addresses the entire configuration of information a model has access to across a task, including data it retrieves, tools it can call, and history it carries forward — prompt wording is one small piece of a much larger picture.
Does a bigger context window solve the problem?
No. Research from Chroma tested 18 frontier models and found every one degraded in accuracy as input length grew, even well before hitting the model’s actual context limit — a phenomenon researchers call context rot.
What is “lost in the middle”? It’s a documented bias where models use information well when it appears near the start or end of a long context, and lose track of the same information when it’s buried in the middle — one study found accuracy dropped from 70-75 percent to 55-60 percent depending purely on position. Who coined the term context engineering?
Shopify CEO Tobi Lütke popularized the specific term in a 2025 social media post, which AI researcher Andrej Karpathy quickly amplified; Anthropic has since published its own formal engineering guide built around the concept. Does context engineering make retrieval-augmented generation (RAG) unnecessary? The opposite: context rot research is often cited specifically to show that dumping large amounts of text into a huge context window performs worse than selective, well-curated retrieval, which is exactly what RAG is built to do.
How much does poor context management actually cost? Andrej Karpathy has pointed out that roughly 90 percent of a typical AI coding bill can go toward unnecessary context, while Anthropic has reported cutting token consumption by 84 percent in a long-running task through active context management. What is context rot?
Context rot is the degradation of an AI model’s accuracy and reliability as its input context grows, even when the added information is genuinely relevant — distinct from simply running out of context-window space. Do I need to think about context engineering for everyday AI chat use?
Not much — a single, clearly worded question with its relevant information already visible doesn’t have a curation problem to solve. This matters most for multi-step agents, tool-calling systems, and long-running conversations.
What’s the first practical step toward better context engineering? Run the Necessity, Position, and Decay checks from this guide’s Context Budget Test against whatever AI system you’re already building: is everything in its context actually needed, is critical information placed where the model will attend to it, and is anything stale enough to remove.
Written by
Muntasir Ahmad Chowdhury
Founder, AI Hustle World
Muntasir Ahmad Chowdhury is the Founder of AI Hustle World, an independent publication dedicated to making Artificial Intelligence practical, trustworthy, and easy to understand. He researches AI tools, automation, customer service, productivity, and real-world business applications, helping readers make smarter technology decisions through research-driven, experience-backed content.
Expertise:
AI Tools • AI Automation • AI Customer Service • AI Productivity • Generative AI • AI Workflows
Get Smarter With AI
Enjoyed this guide? Get practical AI tools, tutorials, and honest reviews delivered to your inbox.
4 thoughts on “Context Engineering Explained: How to Give AI Models the Right Information”