A growing memory store you have not reviewed in a while
30-minute memory audit
Sample ~15 items each from most-used, oldest, sensitive, and random
Sample four groups, don’t scroll forever
Pick roughly fifteen memories from each of four groups: the ones the agent uses most, the oldest entries, anything you would flag as sensitive or personal, and a random handful. For a store of a few hundred items that is about sixty—enough to see the shape of the whole without reading every line. If your dashboard shows how often each memory is used, sort by that; if not, start with the facts your main workflow leans on.
Each group exposes a different failure. Most-used items carry the widest blast radius, because an error there reaches every task. The oldest are the likeliest to have quietly gone stale. Sensitive items carry the highest privacy cost. The random group catches what the other three miss and gives you an honest read on the store overall. How to judge it: if the random sample is clean but the oldest group is full of errors, your problem is age, not everything.
The tempting mistake is to read the list from the top and work down—you burn the half hour on the first forty low-risk entries and never reach the risky ones. Sampling can miss a rare bad item, and you accept that in exchange for finishing today instead of never.
Score each item on four fast tests
Run each sampled memory through four quick checks. Current: is it still true today? Sourced: can you see where it came from? Useful: does it change a real task, or is it trivia? Appropriate: would you be comfortable if it surfaced in a draft to a client? A memory earns its place only by passing all four, and a single failing check points to what you do next.
Keep the pace to a few seconds per item—if you are deliberating, that is already a signal the item needs correcting or removing, not keeping. How to judge it: you should clear sixty items in about the time it takes to make a coffee. The failure here is turning each entry into a research project and scoring twelve items in the half hour instead of sixty.
Yes-or-no scoring is deliberately coarse and will occasionally mislabel a borderline case. A detailed rubric would be more accurate, but the accurate version is the one that never gets finished—speed is what makes this a monthly habit rather than a weekend.
Use case: Priya’s drifting memory store
Priya runs a solo marketing consultancy, and her agent has quietly built up about 120 memories over three months—client preferences, brand facts, pricing notes, a few stray personal details. She catches the drift in a draft proposal: the agent quotes a day rate she retired in spring. If one fact is wrong, others probably are too.
She had three ways to respond. Reading all 120 in order would take over an hour, and she knew she would be skimming by item forty. Wiping the store and starting fresh would clear the errors but throw away the genuinely useful context that makes the agent worth having. Instead she sampled: fifteen most-used, fifteen oldest, all nine she considered sensitive, and fifteen random—fifty-four items in about half an hour.
The sample turned up twelve problems, roughly one in five. She corrected the retired rate first because it was high-use, expired four dead facts, deleted two entries holding a client’s personal detail the agent never needed, and set aside six unsourced claims until she could confirm them. Then the pattern surfaced: every unsourced item had been saved in the first month, before she had thought about what the agent should keep.
The lesson: the corrections took about ten minutes, but the fix that mattered was changing what the agent is allowed to remember going forward. Individual cleanups keep a store tidy for a week; a tighter intake is what keeps the next audit short.
Priya’s 30-minute audit, by the numbers
Illustrative figures from one founder’s audit of a 120-memory store—an example of how a timed sample becomes a short list of actions, not a customer result.
About fifteen each from most-used, oldest, sensitive, and random.
Roughly one in five—stale, unsourced, or over-personal.
Change the intake once instead of re-auditing the same errors.
Take exactly one of four actions
Give every scored item one action. Keep what passed all four tests. Correct a fact that is right in spirit but wrong in detail—the retired rate, the old address. Expire what was genuinely useful once but no longer applies. Delete what should never have been stored at all, and note why, because that reason is what you fix later.
Do not open a fifth “review later” pile unless it has a named owner and a date—otherwise it just relocates the mess and returns, larger, at the next audit. Unsourced facts that would carry real weight in a decision are the one exception worth pausing on: set them aside until you can confirm them, rather than trusting a claim with no origin.
How to judge it: after the pass, every sampled item sits in one of the four buckets; a growing “not sure” pile means your four tests are too vague, not that the item is special. The trade-off is direction—a wrong “delete” loses a little useful context, but a wrong “keep” quietly steers future work and is far harder to notice, so when a low-use item is borderline, lean toward removing it.
Fix the pattern, not just the items
Individual fixes decay. If you correct twelve stale items but leave the intake unchanged, you will audit the same twelve kinds of error three months from now. So step back from the items and read the sample as a diagnosis of the store itself.
Match each pattern to a cause. Many memories missing a source means you are letting the agent keep claims it cannot back up—tighten what it is allowed to remember. Expired information still sitting active means the store has no habit of retiring facts—decide now when the next review happens and put it on your calendar. Items nobody can explain mean the memory policy is too broad and should be narrowed to what your work actually uses.
How to judge it: a good pattern fix lowers next month’s issue count without you touching a single individual item. The trade-off is real—tighten too hard and you starve the agent of useful context—so loosen deliberately for specific, sourced facts rather than leaving the intake wide open.
Record the result and set the next date
Write down four things before you close the tab: how many items you sampled, what you found by type, what you did about it, and the date of the next review. A few lines is enough. A single audit is only a snapshot; the value lives in the trend.
At the next audit you compare—are stale and unsourced items falling as a share of the sample? If they are, your pattern fixes are working. If they are not, the change you made last time did not hold, and that is worth knowing before you spend another half hour. How to judge it: across two or three audits the issue rate should trend down; a flat line means you are treating symptoms.
The failure is auditing but keeping no record, so every pass feels like the first and you can never tell whether the store is getting healthier. Recording costs a few minutes now; skipping it costs you the only evidence that the work is paying off. Monthly suits a fast-growing store; quarterly is plenty once the issue rate settles.
Try this next
- Pull a sample: about fifteen memories each from most-used, oldest, sensitive, and random.
- Score each for current, sourced, useful, and appropriate—a few seconds apiece.
- Give every item one action—keep, correct, expire, or delete—and set aside unsourced high-impact facts until verified.
- Note the issue count and next review date, then fix the one pattern behind most of the errors.
Sources and further reading
These primary references support the article’s approach to keeping stored information accurate and minimal, protecting sensitive data, and managing risk as the store grows.
The accuracy, data minimisation, and storage limitation principles behind the four scoring tests and the expire/delete actions.
OWASP GenAI Security ProjectLLM02:2025 Sensitive Information DisclosureWhy sensitive stored data deserves its own sample group, with sanitization and minimization as mitigations.
NISTAI Risk Management FrameworkA framework for measuring and managing AI risk over time—support for recording each audit and tracking the trend.
Ready to put one useful workflow to work?
Start with one clear job, a result you can review, and boundaries you understand.
See launch pricing