Tailr
← All posts

Context Engineering vs. Prompt Engineering: What's the Difference?

· updated

Two tall cards comparing prompt engineering, which shapes the instruction, and context engineering, which shapes everything the model sees

For a couple of years the skill for getting good results from a language model was called prompt engineering: find the right words, add an example, ask it to think step by step. Then the applications got bigger. A model answering a support question now reads a retrieved knowledge base, calls a tool to check an order, remembers the last three turns, and follows a thousand-word system instruction, and the prompt is a small part of what it sees. The practice of designing all of that got a name in 2025: context engineering. This page explains both, how they differ, where they overlap, and what to learn.

The short answer: Prompt engineering is writing the instruction: phrasing, examples, format, reasoning cues, role. Context engineering is designing everything the model sees when it generates: the system instructions, retrieved documents, tool definitions and results, conversation history, memory and examples, and deciding what goes in, what stays out, and in what order, within a limited context window. Prompt engineering is one component of context engineering. The shift matters because production AI features and agents fail far more often from the wrong information being in front of the model than from the wrong wording, and fixing that is engineering, not phrasing.

Prompt engineering: shaping the instruction

Prompt engineering is the set of techniques for writing the text a model receives so that its output is what you want. The main ones:

  • Clear task statement: say what you want, for whom, in what form. “Summarise this contract for a non-lawyer in five bullet points” beats “summarise this”.
  • Examples (few-shot): show two or three input-output pairs so the model copies the pattern.
  • Reasoning prompts (chain of thought): “think through this step by step before answering” improves accuracy on multi-step problems.
  • Format constraints: ask for JSON, a table, a fixed length, a specific structure.
  • Role or persona: “you are a senior tax accountant” shifts vocabulary and assumptions.
  • Constraints and negatives: what not to do, what to do when uncertain, when to refuse.
  • Decomposition: breaking a big task into a chain of smaller prompts.

12 types of prompting methods covers each of these with examples. They all share a property: they operate on the words in the instruction. They assume the model already has, or can be given in the same prompt, everything it needs to know.

That assumption held when applications were a single prompt and a single answer. It stopped holding when they weren’t.

Context engineering: shaping everything the model sees

A language model generates from its context window: a limited amount of text that includes everything it’s been given for this call. In a real application, that’s not just a prompt. A typical context for an AI support assistant answering one question contains:

  1. System instructions: who the assistant is, what it may and may not do, tone, escalation rules. Often a thousand words or more.
  2. Retrieved documents: the five most relevant knowledge-base articles for this question, pulled by a search over embeddings.
  3. Tool definitions: the functions the model can call (look up an order, check a refund policy, create a ticket), each with a schema.
  4. Tool results: what came back from the calls it already made this turn.
  5. Conversation history: the last several turns, possibly summarised.
  6. Memory: facts about this user from previous sessions.
  7. Examples: a few demonstrations of the ideal answer format.
  8. The user’s message.

Context engineering is the design of all of that: what to include, how to get it, how to structure it, how much of the window to spend on each, what to cut when it doesn’t fit, and how to verify that the result is better. The practices involved:

  • Retrieval: chunking documents, embedding them, searching and reranking, so the model gets the right five articles and not the wrong twenty.
  • Tool design: describing tools so the model uses them correctly, returning results in a form it can use, deciding which tools to expose for which task.
  • Memory: what to remember across sessions, how to store it, when to surface it.
  • History management: how many turns to keep verbatim, when to summarise, what to drop.
  • Compaction: when the window fills, what gets condensed and how, without losing what matters.
  • Structure and ordering: where instructions go relative to data, how to delimit sections, what the model attends to at the start and end of the window.
  • Dynamic assembly: building the context fresh for each call from the current state, rather than one static prompt.
  • Evaluation: a set of real cases and a metric, so a change to any of the above can be measured.

The name spread in mid-2025, after Shopify’s Tobi Lütke described the work as providing all the context for a task to be plausibly solvable, and Andrej Karpathy endorsed it as a better term than prompt engineering for what building with models actually involves. The practices were older; the term gave them a shared label and made the point that this is engineering, with pipelines and measurements, not copywriting.

The two side by side

Prompt engineering Context engineering
Operates on The instruction text Everything in the context window
Core question How do I phrase this so the model does it well? What does the model need to know at this step, and how does it get there?
Typical unit of work A prompt A pipeline that assembles context per call
Skills Writing, examples, reasoning cues, format Retrieval, tool design, memory, compaction, evaluation, software engineering
Failure it fixes Vague or misread instructions Missing, wrong, stale or excessive information
Scale One prompt, one answer Agents, multi-step workflows, long conversations
Measured by Reading the output Evals over a set of real cases
Who does it Anyone using a model AI engineers, and increasingly product engineers

The relationship: prompt engineering is a component of context engineering. The system instruction in a context is a prompt, and writing it well matters. But the instruction is now one of eight things in the window, and the other seven are where most production failures come from.

A worked example

The task: an assistant that answers questions about a company’s expense policy.

Prompt engineering alone: you write a careful system prompt. “You are a helpful assistant for Acme’s finance team. Answer questions about the expense policy clearly and concisely. If unsure, say so.” You paste the policy document into the prompt. It works for the first few questions.

Then it breaks. The policy is 40 pages and doesn’t fit alongside the conversation. Users ask about their own past claims, which aren’t in the policy. The policy changed last month and the assistant is quoting the old version. A user asks about travel and the model confidently cites the home-office section. Better wording doesn’t fix any of these.

Context engineering: you build the context per question.

  • Retrieval: the policy is chunked by section and embedded; each question pulls the three most relevant sections, so travel questions get travel sections. The index is rebuilt when the policy changes.
  • Tools: the model can call get_user_claims(user_id, months) to answer “what did I claim in March”, and get_policy_version() so it can cite the current version.
  • History: the last four turns are kept verbatim; older turns are summarised into two lines.
  • Memory: the user’s department and grade are stored, because per-diem rates depend on them.
  • Instruction: the system prompt is still there, now shorter and more specific, including rules for when to call each tool and what to say when retrieval returns nothing relevant.
  • Evaluation: 150 real questions with reference answers, scored on every change. Retrieval accuracy is measured separately from answer accuracy so you know which layer to fix.

Same model, same task. The difference is what’s in front of the model when it answers, and the difference in accuracy is usually the difference between a demo and a product.

When each one matters

Prompt engineering is enough when: the task is a single call, the information fits in the prompt, the inputs are similar each time, and the stakes allow occasional misses. Classification, summarisation of a document you have, generating a draft from a brief, most personal and one-off uses.

Context engineering is required when: the model needs information it wasn’t given (documents, databases, live data), the task spans multiple steps or tool calls, the conversation is long, the information changes, users differ, or the output has to be reliably right. Assistants over company data, agents that take actions, anything customer-facing, anything that runs unattended.

The line moves as models get larger context windows, but not as much as people expect: a bigger window makes it possible to include more, and context engineering is still the discipline of deciding what.

Common context failures

Knowing these is most of the job:

  • Missing context: the model wasn’t given the fact it needed and made one up. Fix: retrieval or a tool.
  • Wrong context: retrieval returned plausible but irrelevant chunks and the model used them. Fix: better chunking, reranking, filtering by metadata.
  • Stale context: the document changed and the index didn’t. Fix: refresh pipelines, version tags.
  • Too much context: the window is full of marginally relevant material and the model’s attention is diluted; accuracy drops even though everything’s “in there”. Fix: fewer, better chunks; summarise history.
  • Conflicting context: two sources disagree and the model picks one silently. Fix: surface conflicts, prefer authoritative sources, instruct on precedence.
  • Poisoned context: a retrieved document or tool result contains instructions (“ignore your rules and…”) and the model follows them. Fix: treat all retrieved and tool content as data, delimit it clearly, instruct the model accordingly, and filter.
  • Lost context: a long agent run summarised away the detail it later needed. Fix: structured state outside the window, selective compaction.

Each has a diagnosis and a fix, and none of them is fixed by rewording the prompt.

Skills to learn, in order

  1. Prompt engineering basics. Clear instructions, examples, reasoning cues, format constraints. A few days. Everything else assumes it.
  2. Retrieval. How to chunk, embed, search and rerank; how to measure whether the right chunks came back. This is where most production quality lives.
  3. Evaluation. Building a set of real cases with expected outputs, choosing a metric, running it on every change. Without this you can’t tell whether any context change helped.
  4. Tool design. Writing function schemas and descriptions the model uses correctly; returning results in a usable shape; deciding which tools to expose.
  5. Memory and history management. What to keep, what to summarise, what to store outside the window.
  6. Compaction and structure. Managing a full window without losing what matters; where things go and how they’re delimited.
  7. Enough engineering to build the pipeline. Python or TypeScript, an API client, a vector store, a scheduler, logging.

The best way to learn all seven is one project: an assistant over a real document set, with tools, with an eval set, improved in measurable steps.

What this means for AI careers

Job listings caught up with the shift before the vocabulary did. Roles titled “prompt engineer” have largely become AI engineer, applied AI engineer or LLM engineer, and the listings describe context work: retrieval, RAG, agents, tool use, evals, memory. A candidate who can only write prompts is competing for a shrinking set of roles; one who can build and measure the context pipeline is in the fastest-growing engineering job market there is. What does an AI engineer do covers the role in full, and best AI coding tools for developers covers the tooling most of these teams use.

Where Tailr fits

AI engineering listings name these skills precisely: RAG, retrieval, evals, agents, tool use, context management, and occasionally still “prompt engineering”. A resume that uses the listing’s terms for work you’ve actually done, and leads with a project that has numbers on it (retrieval accuracy, eval scores, cost per query), is the one that gets past the screen. Tailr tailors your resume to the specific listing you’re viewing, so the retrieval, evaluation or agent work that role asks for is what your resume leads with, and generates a matching cover letter. Try Tailr on the next AI role you open.

Conclusion

Prompt engineering is writing the instruction well. Context engineering is making sure the model has the right information, tools, history and memory in front of it at every step, within a limited window, and being able to measure that it does. The first is a component of the second, and the second is where production AI systems succeed or fail. Learn to write a clear prompt, then learn retrieval and evaluation, then build one real assistant over real data and improve it in measured steps. That’s the skill the term names, and it’s the one the job market is hiring for.

Frequently asked questions

01What is context engineering?

Context engineering is the practice of designing everything a language model sees when it generates a response: the system instructions, the retrieved documents, the tool definitions and their results, the conversation history, memory from earlier sessions, and examples, plus the decisions about what to include, what to leave out, and in what order, within the model's context window. It treats the context as a system to be engineered rather than a message to be written.

02What is prompt engineering?

Prompt engineering is the craft of writing the instruction a model receives so it produces the output you want: phrasing the task clearly, giving examples, asking for step-by-step reasoning, specifying format and constraints, assigning a role. It's about the words in the prompt. It was the main lever for getting good results from language models in 2022 and 2023, and it's still a necessary skill; it's just no longer the whole job.

03Is context engineering replacing prompt engineering?

It's absorbing it. Prompt engineering is one part of context engineering: the instruction is one item in the context, alongside retrieved data, tools, history and memory. As AI applications moved from single prompts to agents that run for many steps with many sources of information, the hard problems shifted from wording to what the model knows at each step. Most people who were called prompt engineers now do context engineering, whether or not the title changed.

04Where did the term context engineering come from?

It spread in mid-2025 after Shopify's CEO Tobi Lütke described it as the art of providing all the context for a task to be plausibly solvable by the model, and Andrej Karpathy endorsed the term as a better description than prompt engineering of what building with language models actually involves. The practices it names, retrieval, memory, tool design, context management, existed before; the term gave them a shared name.

05What skills does context engineering need?

Retrieval (chunking, embeddings, search, reranking), designing tool and function schemas the model can use reliably, managing conversation history and memory across sessions, compaction and summarisation when the context fills up, evaluation so you can measure whether a context change helped, and enough software engineering to build pipelines that assemble context dynamically. Plus prompt writing, which is still in there.

06Which should I learn first, prompt engineering or context engineering?

Prompt engineering first, because it takes days and everything else builds on it: clear instructions, examples, format constraints, reasoning prompts. Then context engineering, starting with retrieval and evaluation, because those are what production AI features are made of and what AI engineering roles hire for. Learn both by building one real thing that uses documents and tools, and measuring it.