AI design
Memory, evals, and AI you can keep
UX100 · 6 October 2026
A model that remembers, and a team that can tell whether the model still behaves, are design problems. Memory the person cannot see is surveillance with a friendly name. A prompt change no one can regression-test is a redesign shipped in the dark.
Memory is a list the person can read
Products now store facts about a person so later answers can use them: a role, a preference, a project, a correction. Human-centric memory shows that list. Each item has a source, a date, and a way to delete or fix it. The next answer that uses a memory says so, the way a citation says which page it used. Hidden personalization that changes a price, a ranking, or a recommendation without a trace fails the same test GiveWell’s layout passes for a charity: the conclusion and the reason travel together.
There is a difference between the conversation in front of you and a profile that persists. The conversation can be long. The profile should be short, specific, and boring. A memory that stores a passing joke and then repeats it for a year is a design bug. Write the rules for what is allowed to stick, show them, and let the person empty the store without closing the account.
Context the model is allowed to see
Before memory, there is context: the document, the matter, the calendar, the mailbox, the screen. A harness decides what is attached to this run. The interface should say what is attached, in a line the person will actually read, and offer a way to remove a source before the model sees it. Consent here is concrete. It is the list on this screen, for this task.
Workplace products have a second reader, the organization. A design that is human-centric for the employee also states what the employer can inspect. Transcripts, memories, and files sent to a model are records. JobzMall’s talent work and Harvey’s legal work both sit next to sensitive material. The pattern is the same: scope the read, show the scope, and keep the write behind a person.
Evals are how design reviews a model
An eval is a set of tasks with an expected shape of answer, run whenever the model, the prompt, or the tools change. Designers belong in that set. The tasks are the jobs the interface claims to support: summarize this filing, cite a source, refuse this request, draft this message without inventing a fact. A change that wins on a generic benchmark and loses on those tasks is a regression in the product, and it should block release the way a broken layout does.
Review the output as an interface. Is the citation present? Is the uncertainty visible? Did the harness stop before a write? Did the empty state still name the job? These are design checks. Teams that leave them to a later “AI quality” group discover the failure in front of a customer. The eval set is a design artifact, kept next to the components.
What to ship, and what to leave in the lab
The trendy surface in 2026 is an agent that goes off and returns with the work done. The durable version is smaller. A person states a job. A harness takes a few visible steps. A draft comes back with sources. The person edits and commits. Memory helps the second visit if the person can see it. Evals keep the next release from quietly getting worse.
That is enough AI for a serious product. A mascot, an unbounded loop, and a memory the person cannot find are optional, and in the UX100 classes they were not what won. Perplexity, Claude, Harvey, and the ChatGPT award that opened the category all point the same direction: a model inside a job, with the work on the table.
Questions
- How should AI memory be designed?
- As a short list the person can read, correct, and delete. Each item needs a source and a date. When a later answer uses a memory, the interface should say so. A profile that changes the product in secret is hidden personalization.
- What is an eval in AI product design?
- An eval is a fixed set of real tasks, with an expected shape of result, run when the model, prompt, or tools change. Designers use it to check citations, uncertainty, refusal, and whether the harness still waits before a write.
