A task that worked once but drifts every time the context changes
20 minutes to write + 3 pilot runs
Write the five decisions before you rewrite the prompt
The problem: a prompt is not a process
“Check our competitors every morning” sounds complete until the first run exposes everything it left unsaid. Which competitors? Which pages? What counts as a change worth your attention? Where does the result go, and who hears about it? A prompt captures the request but silently assumes the answers — and the second run is when those assumptions drift and the output quietly stops being trustworthy.
Do one thing before you touch the wording: pick a task whose result you can verify in under five minutes — a daily change summary, a draft weekly report, a list of unanswered support questions. If you cannot check it quickly, you cannot tell a good run from a bad one, and the agent inherits that blindness.
A common failure: a founder schedules a “morning news” prompt, and for a week it summarizes the same cached page because nobody defined what “new” meant. The trade-off to accept up front: writing the process down costs twenty unglamorous minutes, and it is the only thing that stops you paying that time back on every single run.
The 5-part contract, in full
A workflow you can trust comes from five decisions, not one clever paragraph. Write them on a single page before you rewrite the prompt:
- Trigger — when the work starts, with a timezone (“8:00 AM on weekdays, in your timezone”).
- Inputs — the exact, approved sources it may use, named one by one.
- Output — the precise artifact it produces and where it lands.
- Boundary — what it may never change, send, publish, or buy.
- Owner — the person who reviews exceptions and holds the “off” switch.
Why five and not one: each part closes a gap the prompt leaves open. A named input list stops the agent wandering the whole web; a fixed output makes the result checkable; a written boundary keeps a read-only job from ever quietly becoming an outbound one.
How to judge it: hand the page to someone else and ask them to run the task by hand. If they have to guess at any of the five, the contract is not finished. The trade-off: a tight contract does less than an open-ended one — and that narrowness is exactly what makes it safe to repeat.
Use case: a weekly competitor brief that finally holds still
Priya runs a solo skincare brand and wants a Friday brief on five competitors. Her first instinct is the prompt she already typed once: “tell me what changed this week.” It worked in the demo, so why write more?
Because “what changed” has no boundary. She faced a real choice: keep the open-ended prompt and re-check everything herself, or spend twenty minutes writing a contract that names the five pages, defines a “material change” (price, product, or announcement — not a blog post), fixes the output to a dashboard draft, and forbids any outbound message. She chose the contract.
The pilot earned its keep. On run one the review took eleven minutes because the agent counted a reposted blog and a cached price as changes; she added one rule excluding both. Run two surfaced a genuine pricing-page change, so she kept that source. By run three the review was four minutes and the scope stopped moving — only then did she schedule it.
The lesson: the prompt was never the hard part. The five decisions around it — trigger, inputs, output, boundary, owner — were what turned a lucky first result into work she could safely ignore until Friday.
Priya’s competitor brief, across three pilot runs
Illustrative figures from the worked example below—an example of how a scope tightens across a pilot, not a customer result.
Two rules missing: it read a reposted blog and a cached price as “changes.”
Same contract, tighter inputs — now fast enough to trust on a schedule.
The draft-only boundary held every run; nothing left the dashboard.
Define success before the first run
Pick a measure that is visible in the output itself, not a feeling about it. For a research brief that might be: every claim carries a link, duplicates are removed, and a person can review it in under three minutes. “Looks good” cannot be checked twice the same way, so it cannot teach the agent — or you — anything.
If a person cannot quickly decide whether the work is complete, an agent cannot reliably decide either.
A common failure: success is defined as “a useful summary,” so every run gets graded on mood — great on a calm Tuesday, a failure on a busy one — and nothing about the workflow actually changes. The trade-off: a strict, checkable rule will sometimes reject a run you would have accepted by eye, and that friction is the price of a standard that holds when you are not looking.
Run a 3-run pilot before you schedule
Run the workflow by hand three times before it goes on a schedule. After each run, write down exactly one rule that was missing and one that was unnecessary — no more, or you will rebuild the whole thing every round. Resist adding tools mid-pilot; you are testing the contract, not the toolbox.
Three runs is the smallest number that separates a fluke from a pattern. One good run proves nothing, two can agree by luck, but by the third a stable scope stays stable and a fragile one has already broken twice.
A common failure: the task looks perfect on run one, gets scheduled, then fails silently on run four when a source page changes — because it was never stressed. The trade-off: three manual runs delay launch by a few days, which is cheap next to a scheduled workflow quietly producing wrong output for a month. Keep it only if it saves review time without creating new cleanup work.
A launch checklist you can reuse
When the scope has held for three runs, walk the contract through one last list before you let it run unattended:
- The trigger and its timezone are unambiguous.
- Every input source is named — no “whatever is relevant.”
- The output has one fixed location and format.
- Sending, publishing, buying, and deleting all require approval.
- A named owner knows the exact conditions for stepping in.
How to judge it: if any line makes you pause, that pause is the workflow’s weakest point — fix it before the schedule hides it from view. The trade-off worth naming: a checklist cannot catch a risk you never wrote down, so treat it as the floor, not the ceiling. In eeky AI, keep the first version draft-only and add a single approval gate before anything leaves the dashboard.
Try this next
- Write the five decisions — trigger, inputs, output, boundary, owner — on one page before touching the prompt.
- Define one success rule you could grade the same way twice, visible in the output itself.
- Run the contract by hand three times, changing exactly one rule after each run.
- Schedule it only after the third clean review, and keep the first version draft-only.
Sources and further reading
These primary references support the article’s approach to defining success, keeping a person in control, and governing a workflow before it runs unattended.
Supports defining a checkable success measure before building, rather than grading each run on feel.
Google PAIRFeedback + ControlBacks the boundary and owner decisions—keeping a person in control, especially where stakes are higher.
NISTAI Risk Management FrameworkA govern-measure-manage framework behind assigning an owner, measuring results, and the launch checklist.
Ready to put one useful workflow to work?
Start with one clear job, a result you can review, and boundaries you understand.
See launch pricing