Before you prompt an agent, write one goal with one number.
A cashier screen went from six taps to two in one afternoon because the number came before the prompt. How to write the goal, and how to hold an agent to it.
On this page
- Key takeaways
- What does “start at the end” mean in a design sprint?
- Should a sprint goal have a number?
- Why does an agent need a number more than a team does?
- How do you write one goal with one number?
- Where does the number come from?
- How do you hold the agent to the number?
- What goes wrong when the number becomes the target?
- Is plan mode dead?
- When should you not write a number?
- What changes when the product is in Arabic?
- Which templates help?
- FAQ
- Sources

On 29 September 2026 the cashier screen in Dot, my loyalty app, was slow. A cashier at a busy counter had to start a new claim, tap into the amount, open a country picker, pick the country, tap into the phone and press Generate. Six taps plus typing, and nothing was in focus when the screen opened.
At 17:50 that day a design spec landed in the repo. It had a lot in it: eight screen states, the copy for every error, the layout down to the pixel. The table that mattered was four rows long and had a plain title, “Tap budget”. For a returning customer who gives a phone number, it had two columns. Today: six taps plus typing. New: two taps plus typing.
Everything else in the spec followed from that cell. Two taps meant the cashier could not tap from one field to the next. Phone keypads on iOS have no Return or Next key, so the only way to skip that tap is to move focus automatically, and you can only do that out of a field with a known length. A phone number has one. An amount does not. So the spec put the phone first, and when the last digit lands, focus jumps to the amount by itself.
At 19:13 the build commit landed with Claude as co-author: “phone first, two taps for a returning customer”. A minute later came the test commit, also co-written with Claude. It opens the counter on a 375 pixel phone, in Arabic and in English, types a returning customer’s number and an amount, taps Add points, taps Next customer, and checks that the counter is empty again with the cursor back in the phone field. The test is called “takes a returning customer from idle to success in two taps”. The pull request merged at 19:44.
Under two hours from spec to merge. Claude co-wrote the code and the test. The number told both of them what done meant.
Key takeaways
- Before you prompt an AI agent, write one goal with one number in it. “Returning customers earn points in two taps” is a goal. “Make the cashier flow faster” is a wish.
- GV’s design sprint opens the same way: on Monday morning the team starts at the end and agrees on a long-term goal. On Wednesday it uses that goal to choose between solutions.
- An agent needs the number more than a team does. Anthropic’s own Claude Code guide says Claude stops when the work looks done, and without a check it can run, looking done is the only signal it has.
- Put the number in three places: the first line of the prompt, a test that fails until the number holds, and the review.
- A number on its own invites gaming. Write down what must not break on the way there, the way the Dot spec gave up a tap rather than risk a silent late earn.
- Plan mode may be dead. The goal is not. A plan is pages. A goal is one line, and it survives every round of iteration.
What does “start at the end” mean in a design sprint?
GV’s design sprint is a five-day process for answering critical business questions through design, prototyping and testing ideas with customers. Jake Knapp began running sprints at Google in 2010 and brought them to GV in 2012. Monday maps the problem, Tuesday sketches, Wednesday decides, Thursday builds a prototype, Friday tests it with real people.
Monday morning comes first, and GV’s own description of it is short: you start at the end and agree to a long-term goal. Then you make a map of the challenge, ask the experts in the company what they know, and pick a target, a piece of the problem small enough to solve in one week.
The goal comes back on Wednesday. By then the team has a stack of sketched solutions and can only prototype one, so it critiques each and decides which ones have the best chance of reaching the long-term goal. The goal written on Monday is the ruler used on Wednesday. Without it, Wednesday is a vote on taste.
GV credits John Zeratsky with this part of the method. In their words, he helped them start at the end and focus on measuring results with the key metrics from each business. The end comes first, and the end is something you can measure.
My design process guide puts the same step at the top of the empathise phase, before any interview: one sentence on what the project must change and how you will know it did, with the one number that would prove it. On a team, that sentence keeps a room of people pointed the same way. With an agent, it does a second job.

Should a sprint goal have a number?
Most guides say no, or not necessarily. The Design Thinking Toolkit’s long-term goal page, one of the top results for this question, builds the goal from three parts: a timeframe, a specific aspirational achievement, and a metric marked optional. Its examples look years ahead and read like “we will be the number one choice for”. Ask ChatGPT or Perplexity the same question today and you get the same advice back, with the numbers moved out of the goal and into separate success criteria.
For a sprint room, that advice holds. A team spends five days together, carries the context in its head, and tests with real customers on Friday. The people in the room do the measuring.
A solo builder with an agent has no room. The agent cannot be held to a moonshot, because a moonshot never returns pass or fail. A line like “the counter cashiers love” is a fine thing to believe. The agent can only work toward “two taps”.
So keep both, and write them on two lines. The aspiration on top, if it helps you and your team. The number directly under it, because that is the line the agent, the test and the reviewer will actually use. GV’s own credit to Zeratsky already joins the two: start at the end, then measure the results.
Why does an agent need a number more than a team does?
A team argues back. Give five people a vague goal and someone asks what “faster” means. An agent does a plausible version of the work and stops.
Anthropic says this plainly in its Claude Code best practices: Claude stops when the work looks done. Without a check it can run, “looks done” is the only signal available, and you become the verification loop. Every mistake waits for you to notice it. The guide’s fix is to give Claude something that produces a pass or a fail, so the loop closes on its own: Claude does the work, runs the check, reads the result and keeps going until the check passes.
That is a goal with a number, seen from the agent’s side. “Two taps for a returning customer” can pass or fail. “A smoother checkout” cannot.
We saw what happens without one in May. When we ran ux-skill, the design engine I build, on its own homepage brief, the brief was all adjectives: editorial, confident, calm, cinematic. The engine obeyed every word and returned almost exactly the claude.ai landing, the most likely answer in the category. I told that story in can AI replace user research. Adjectives point at an average. A number points at a place.
Anthropic’s newer /goal command makes the same point in product form. You type a completion condition, and after every turn a separate model checks whether it holds. If not, Claude keeps working. The docs say a condition that holds up across many turns usually has three parts: one measurable end state, a stated check that proves it, and the constraints that matter on the way there. Their example is “all tests in test/auth pass and the lint step is clean”. It reads exactly like a sprint goal with the number filled in.
How do you write one goal with one number?
Write it before you open the agent. One sentence, five parts.
How to run it
- The person. Who is this for, in a real moment. “A tired cashier at a busy counter, one hand on the phone and one on the cup” is a person. “Users” is not.
- The outcome. What they get done. For the counter: a returning customer earns points.
- The number. How you will know it worked. Taps, seconds, steps, errors, a percentage of people who finish. Pick one.
- Today’s value. Count it on the current product before you write the target. Six taps was the counter as it stood. A target with no baseline is a guess with a decimal point.
- How it gets measured. Who or what checks it, and where. On the counter it was a browser test on a 375 pixel phone, in both languages.
Put together: “A returning customer earns points in two taps plus typing, down from six, measured on a 375 pixel phone in Arabic and English.”
The same pattern works for a product that does not exist yet. In the Bunn example I use across my free templates, the problem statement reads: Rana needs to order coffee for three people in under a minute, because the 10:00 meeting starts on time. Under a minute is the number. Bunn is fictional, but the shape is the one I use on real work.

What you walk away with: a filter. Every feature the agent proposes either moves the number or waits. Every review comment either protects the number or is a preference.
The common mistake: a goal that names a solution. “Build a loyalty app” and “add a QR scanner” are tasks. The goal is what the task is for. The counter spec did add a scan button, but the scan path has its own row in the tap budget, three taps plus the amount, so it had to earn its place against a number too.
The second mistake: a goal with five numbers. Time, taps, errors, satisfaction and conversion all on one line means the agent can always point to the one that moved. Pick the number that tracks the person’s outcome, the way nobody wants your product argues, and put the others in the constraints.
Where does the number come from?
From the person and from today’s product, in that order.
Look back at what I wrote for Wfrah in 2023. After ten breeder interviews, the goal was a user experience that educates and builds trust in modern tools while meeting the needs of traditional breeders. It is a good direction, and every word of it came from the research. Reading it now, there is no number in it. It could tell us which way to walk. It could not tell us when we had arrived.

Three sources give you a number you can trust.
Count the current product. The six taps were the old counter as it stood, written into the spec’s “today” column next to the target. Do this before the target, because the target only means something next to the baseline. On a product with analytics, the baseline is already in a dashboard like the Bysooq one above. On a new product, it is how long people take with whatever they use today.
Take it from the person’s moment. Two taps came from the cashier and the keypad. One hand is on the cup, the queue is waiting, the light is poor at night, and a phone keypad has no Next key. A number that comes from the moment is easy to defend when someone asks why it is two and not four.
Take it from the research you already did. If interviews say people abandon the order when the fee appears late, the number is how many finish it. If people fail a task in a usability test, the number is how many pass it.
When it goes wrong: you pick the easiest number to measure. Clicks and page views are easy. Whether the cashier got the customer through is the one that matters.
How do you hold the agent to the number?
Write the number into three places, and check that all three agree.
One: the first line of the prompt. Before context, before files, before style rules. A brief that opens with the goal and closes with the check keeps both in view. Here is the shape I would use for the counter, written as an example:
Goal: a returning customer earns points in 2 taps plus typing, down from 6.
Check: a browser test on a 375px phone, in Arabic and English, that fails
if the flow takes more than 2 taps. Run it and show me the output.
Constraints: never clear what the cashier typed. A retry never awards twice.
No auto-submit when the connection comes back.
Spec: the counter spec, section 1.Two: a test that fails until the number holds. This is the part that turns a goal into a leash. On the counter, the test commit landed one minute after the build commit, and its name is the goal. Anthropic’s guide lists the options from loosest to strictest: ask Claude to run the check in the same prompt, set it as a /goal condition across a session, or make it a Stop hook, a script that blocks the turn from ending until the check passes. Each step trades setup for attention.
Three: the review. The guide recommends a reviewer that runs in a fresh context and sees only the diff and the criteria you give it, not the reasoning that produced the change. Give it the goal as the criterion. It will ask the one question the author forgot: does this actually take two taps?

Ask for evidence instead of a claim: the test output, the command, a screenshot. The counter tests save fourteen screenshots of the counter’s states, in Arabic and English, two of them in dark mode.
What goes wrong when the number becomes the target?
The economist Charles Goodhart noticed in 1975 that a statistic tends to collapse once people put pressure on it to control something. The anthropologist Marilyn Strathern later put it in the form most people quote: when a measure becomes a target, it ceases to be a good measure. Goodhart’s law applies to agents more than to people, because an agent pushes on the number without any sense of why the number exists.
Tell an agent “two taps” and nothing else, and there are cheap ways to get there. Submit automatically when the connection comes back. Drop the confirmation. Skip the undo. Each one removes a tap and breaks the counter.
The Dot spec is worth reading for what it refused. It says there is no auto-submit when the connection returns: the button comes back and the cashier taps again, because the cashier may have moved on and a silent late earn is worse than one more tap. It says never to clear what the cashier typed on an error. It gives every attempt an idempotency key, so a nervous double tap can never award points twice, and there is a test for exactly that. It keeps a ten minute undo with a required reason. Every one of those rules costs taps or code, and every one is written next to the number.
That is the answer to Goodhart. The goal has one number. The constraints carry everything the number cannot see.
What to write in the constraints
- What must never happen. Lost input, a double charge, a silent failure.
- What the number is not allowed to trade away. Accessibility, the undo, the error message that names the field and the fix.
- Where the change must stop. Files not to touch, behaviour that must stay the same.
If you cannot list at least two constraints, you do not yet understand the number well enough to hand it to an agent.
Is plan mode dead?
On 24 September 2026 Ayman Nadeem published Plan mode is dead. It reached 591 points on Hacker News. His argument is that plan modes made sense when models needed hand-holding, and that the separate steps of chat, spec, review, approve and implement now feel artificial. People do not read long generated specs. Work moves in a cycle instead: understand, act, inspect, clarify, adjust, act again. Planning still happens. It just does not need to be a document called the plan.
He is mostly right about plans. Anthropic’s own guide says that if you could describe the diff in one sentence, skip the plan.
Notice what the argument keeps. Nadeem writes that thinking through what to build, why it matters, and evaluating decisions remains important. The Hacker News thread split along the same line. One commenter, thiht, wrote that plan mode helps ensure the model will actually do what you have in mind. Another, loeg, described the alternative: the model makes some plausible choices, and you can ask it to make different ones later.
Simon Späti’s essay from the same week, the problem is not the AI code, names what goes missing when the plan disappears entirely: teams stop knowing the intent behind why choices were made. His line is blunt. You end up with no plan whatsoever.
My position sits between them. Drop the plan if it slows you down. Keep the goal. A plan describes the route and goes stale on the first turn of the cycle. A goal describes the destination, so it survives every turn. Understand, act, inspect, adjust: each of those steps needs something to inspect against, and that something is one line with a number in it. The Dot spec runs to 385 lines. The part that set the build was the four-row table.
When should you not write a number?
When the question is still open. Some work is exploration, and forcing a number on it too early measures the wrong thing very precisely.

Taste. Whether a map feels like a place, or a brand feels like its people, does not reduce to a count. You can still set a frame: three directions to compare, a deadline, a person who decides. The judgement stays with a human, and the agent’s job is to produce options fast.
Discovery. Before the interviews, you often do not know which number matters. On Wfrah we could not have written “breeders calculate feed in under a minute” before we learned that the calculation was the problem. In that phase the goal is a question, and the number you write is how many people you will talk to.
When the number would be fake. If you cannot measure it on this product, in this release, it is decoration. “Increase retention by 20%” on a feature that ships next week tells the agent nothing it can check. Pick something it can run.
What changes when the product is in Arabic?
A search for writing on goal setting for Arabic or right-to-left products turned up nothing. The gap matters, because a number checked in one language is not checked in the other.
The counter test runs the same two-tap flow twice, once in Arabic and once in English. The spec’s visual checks include right-to-left details a single-language test would miss: the phone number row stays left to right even on an Arabic screen, because that is how people read a number out loud, and the amount aligns to the right. Arabic and English lines rarely break in the same place, so a layout that passes in English can push a button below the fold in Arabic.
Write the language into the goal. “Two taps, in Arabic and English” is a different goal from “two taps”, and the agent will test what you write. When we checked 179 funded MENA startups, 44% had no Arabic website at all. Arabic left out of the goal tends to be left out of the product.
Which templates help?
Every free template ships in Arabic and English and uses one worked example, Bunn.
| Step | Free template | What it gives the goal |
|---|---|---|
| Write the goal | Problem statement | The person, the moment, the cost |
| Find the path | User flow | Where the taps and steps actually are |
| Cut the ideas | Prioritisation matrix | What moves the number first |
| Check it | Usability test | Pass, partial and fail for each task, with evidence |

All four are also in UX AI Kit Pro Max, with Claude prompts for each.
FAQ
What is a long-term goal in a design sprint?
It is the outcome the team agrees on first thing Monday morning, before mapping the problem. GV describes it as starting at the end. The goal comes back on Wednesday, when the team decides which sketched solutions have the best chance of reaching it. Write it with a number so Wednesday’s decision has something to measure against.
How do you write a measurable project goal?
One sentence with five parts: the person, the outcome they need, one number, today’s value of that number, and how it will be checked. “A returning customer earns points in two taps plus typing, down from six, checked on a 375 pixel phone” is measurable. “Improve the cashier experience” is not.
How do I give an AI coding agent a goal?
Put the goal with its number in the first line of the prompt, give the agent a check that returns pass or fail, such as a test, and list the constraints it must not break. Anthropic’s Claude Code guide recommends exactly this, and its /goal command keeps Claude working until a separate model confirms the condition holds.
What is Goodhart’s law and why does it matter for AI agents?
In Marilyn Strathern’s wording, when a measure becomes a target, it ceases to be a good measure. An agent pushes on whatever number you give it, so it will find the cheap ways to move it. Write constraints next to the number: what must never break, and what the number cannot trade away.
Should a design sprint long-term goal be a metric?
Most sprint guides treat the long-term goal as an aspiration and keep metrics optional or separate. That works for a team in a room. If an AI agent will do the building, write the aspiration and then a second line with one number, because the agent can only be held to something that passes or fails. “A returning customer earns points in two taps” is a good example.
Sources
- design sprint gv.com
- Design Thinking Toolkit’s long-term goal page designthinkingtoolkit.co
- Claude Code best practices code.claude.com
- docs code.claude.com
- Goodhart’s law en.wikipedia.org
- Plan mode is dead aymannadeem.com
- Hacker News news.ycombinator.com
- the problem is not the AI code ssp.sh