"Prompt engineering" has a bad reputation among developers: it sounds like guessing magic phrases. In production it is something more ordinary — writing clear, testable instructions for a component whose behaviour you cannot fully specify in code. A good production prompt is closer to a well-written spec than to a clever trick. This guide covers the techniques that consistently matter, how to structure prompts as code, and how to change them without breaking what already works.
What changed with modern models
Early prompting advice was full of tricks: "take a deep breath", threats, tips, all caps. Modern instruction-tuned models follow plain, specific instructions well, and vendors' own guides converge on the same fundamentals (Anthropic prompt engineering overview; OpenAI prompt engineering guide). A large survey of prompting techniques catalogued dozens of methods, but most production gains come from a handful of them (Schulhoff et al., The Prompt Report, 2024).
The shift for developers: stop thinking of prompts as incantations and start thinking of them as interfaces — inputs, behaviour, outputs, error cases.
Anatomy of a production prompt
A reliable structure for system prompts:
- Role and context. Who the model is acting as, who the user is, what product this is.
- Task. What to do, in one or two sentences.
- Rules and constraints. What to always and never do, with reasons.
- Input description. What data the model will receive and how it is delimited.
- Output specification. Format, length, language, schema.
- Examples. A few input/output pairs covering typical and tricky cases.
- Edge cases. What to do when information is missing, ambiguous or out of scope.
You are a support assistant for Acme Accounting, a SaaS for small businesses
in Ukraine. Users are accountants; answer as a knowledgeable colleague.
Task: answer the user's question using only the documentation in <docs>.
Rules:
- If the docs do not contain the answer, say so and suggest contacting
support@acme.ua. Do not guess — wrong tax advice causes real harm.
- Answer in the user's language (Ukrainian or English).
- Keep answers under 150 words unless the user asks for detail.
- Never mention internal ticket numbers or employee names from the docs.
<docs>
{{retrieved_documents}}
</docs>
Output: a direct answer first, then a "Sources:" line listing doc titles.
Notice the reason after "Do not guess". Explaining why a rule exists helps the model generalise it to cases you did not list — Anthropic's guidance calls this out explicitly.
Techniques that consistently help
Be specific about what you want
Vague: "Summarise this document." Specific: "Summarise this contract for a sales manager in five bullet points: parties, term, price, termination conditions, unusual clauses. Quote clause numbers." The specific version is longer and much more reliable. The test: would a new human colleague, reading only this prompt, produce what you want?
Separate instructions from data with tags
When a prompt contains documents, user input or tool results, delimit them clearly. XML-style tags work well across models: <docs>, <user_question>, <previous_answer>. Tags make it easier for the model to tell your instructions from content, and easier for you to parse the output. They also help — modestly — against prompt injection, though they are not a security boundary; see AI agent security.
Use examples deliberately
Few-shot examples are the most powerful way to show format and tone (Brown et al., 2020). Rules for good examples:
- 3–5 examples, diverse enough that the model does not copy surface features.
- Include at least one edge case (missing information, refusal).
- Wrap them in
<example>tags so they are not confused with real input. - Keep them in sync with your rules — contradictions between examples and instructions confuse models.
Let the model think when the task requires it
For multi-step reasoning — calculations, classification with many criteria, planning — asking the model to reason before answering improves accuracy (Wei et al., 2022). Modern models often have built-in extended thinking modes that do this without prompting; use them for hard tasks, and skip reasoning for simple extraction where it only adds latency and cost. If you need the reasoning separate from the answer, ask for it in a <thinking> block and the answer in an <answer> block, then parse only the latter.
Specify the output precisely
Say what format, length and structure you want, and what to leave out ("no preamble", "no markdown headings"). For machine-readable output, use structured outputs or tool calling with a JSON schema instead of asking nicely for JSON; see structured outputs.
Put long documents first, the question last
For prompts with long context, placing documents at the top and the question and instructions at the end tends to improve answers. Models also attend unevenly across very long contexts, a phenomenon documented in Lost in the Middle. Less, more relevant context usually beats more context — the theme of context engineering.
Prompts are code: structure them like code
In production, prompts should live in the repository, not in a dashboard someone edits on Friday evening.
prompts/
support_answer/
system.md # system prompt template
examples.yaml # few-shot examples
schema.json # output schema
CHANGELOG.md # what changed and why
ticket_classifier/
...
from pathlib import Path
from string import Template
PROMPT_DIR = Path(__file__).parent / "prompts"
def render(name: str, **vars) -> str:
text = (PROMPT_DIR / name / "system.md").read_text()
return Template(text).substitute(**vars) # fails loudly on missing vars
system = render("support_answer", product="Acme Accounting")
Practices that save pain:
- Version prompts with the code that calls them. A prompt change is a code change and goes through review.
- Log the prompt version with every request, so you can correlate quality changes with prompt changes; see LLM observability.
- Fail on missing template variables instead of silently sending
{{customer_name}}to the model. - Keep dynamic content at the end and stable instructions at the beginning; this also maximises prompt cache hits.
Changing prompts safely
Every prompt change can fix one case and break three others. The only reliable way to know is to run evaluations:
- Maintain a golden dataset of 50–200 real inputs with expected properties.
- Run the current and the new prompt on the full set.
- Compare scores per category, not just overall.
- Read a sample of changed outputs before shipping.
The full method is in how to evaluate LLM applications. Without evals, prompt engineering really is guessing.
Model upgrades are prompt changes too
When you switch to a newer model, re-run your evals. Newer models often follow instructions more literally — a prompt that relied on the old model "filling in the gaps" may now produce narrower output. Typical fixes: state the desired level of detail explicitly, remove emphatic language ("CRITICAL", "MUST") that newer models over-apply, and re-check examples. Vendors publish migration notes for this reason; read them.
Anti-patterns
- Prompt soup. A 3,000-word system prompt accumulated over months, with contradictory rules nobody dares to delete. Refactor prompts like code; delete rules that evals show are unnecessary.
- Negative-only instructions. "Don't use markdown" works less reliably than "Write in plain prose paragraphs".
- Hidden requirements. Business rules that exist only in the prompt and nowhere in the code or docs. Put important constraints in validation code too.
- Testing on five examples. Feels productive, proves nothing.
- Asking the model to do arithmetic or exact lookups that code could do. Compute in code, pass results in.
A checklist before shipping a prompt
- Role, task, rules, input, output and edge cases are explicit
- Data is delimited with tags; user input cannot be confused with instructions
- 3–5 examples cover typical and edge cases and agree with the rules
- Output format is enforced with structured outputs where machine-readable
- The prompt lives in the repository, versioned and reviewed
- Evals pass on the golden dataset, with per-category scores
- Prompt version is logged with every request
FAQ
Should prompts be in English if users write in Ukrainian? Instructions in English work well with major models, and you can instruct the model to answer in the user's language. For domain-specific terms, include examples in the target language.
How long should a system prompt be? As long as needed and no longer. Many production prompts are 300–1,500 words. Length is fine; contradictions and filler are not.
Is prompt engineering still needed with agents? More than ever, but its scope widens to tool descriptions, context selection and memory — see context engineering and designing tools for agents.
Can an LLM improve our prompts? Yes — models are good at critiquing and rewriting prompts. Treat their suggestions as candidates and accept them only if evals improve.
Sources
- Anthropic. Prompt engineering overview.
- OpenAI. Prompt engineering guide.
- Schulhoff et al. (2024). The Prompt Report: A Systematic Survey of Prompting Techniques.
- Brown et al. (2020). Language Models are Few-Shot Learners.
- Wei et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.
- Liu et al. (2023). Lost in the Middle: How Language Models Use Long Contexts.