Kevin Weil, OpenAI's CPO, told product managers that writing evals is the most important thing a PM can do in the AI era. Elle Shwer, a PM on GitHub Copilot, said the responsibility to define what "good" looks like for users naturally fell to Product. Braintrust published a framework mapping the old PRD onto evals: functional requirements become test scenarios, acceptance criteria become eval rubrics. The spec document has a successor. Design was not in the room when it was written.
The eval runs
The old PRD described what the product should do. Teams debated it, built against it, then quietly ignored half of it.
The eval is different. It executes. It runs against every model update and returns a score. Braintrust describes the PM's new directive to engineering: here is the eval, make this number go up.
Whoever writes the eval criteria defines what the product optimizes for. The spec was negotiable. The eval just runs. And right now, product managers and engineers set the measurement in a closed loop.
Amplitude's guide says PMs and engineers share ownership. Shwer learned evals through workshops with engineers. Aman Khan of Arize AI framed the job as making sure LLMs don't embarrass the company or the brand. Every guide I read follows the same pattern. Design does not appear in any of them.
Which public does the eval serve
I use a lens from the foundation of this site: every product has four audiences. The user, the buyer, the regulator, and the surrounding ecosystem. Most teams design for the user, remember the buyer when revenue comes up, and discover the other two when something breaks.
Evals inherit the blind spots of whoever writes them. When PMs set the criteria, they optimize for what PMs measure: task completion and brand safety. Khan's framing is the buyer public talking. Legitimate, but partial.
The user public asks different questions. Did the response match the context the person was actually in? Did it earn trust or dissolve it?
Amplitude spotted the gap: a high eval pass rate tells you the model performs on a test set, not whether successful interactions drive retention. The signals underneath retention live in design's territory. Design has no seat at the eval table.
The structural repeat
This happened before. Design was excluded from writing requirements because the PRD belonged to product management, engineering consumed it, and design entered after the document was locked.
The eval is the spec's successor. The exclusion is playing out again. The old PRD was a document that teams could negotiate around. The eval is automated. It runs on every deploy.
The criteria it encodes keep shaping model behavior without anyone revisiting them. If the user public is absent from those criteria, no amount of design review after the fact compensates. The optimization function has already been set.
Designative put it plainly: "If designers and PMs do not define what good experience means in scoreable terms, most evaluation systems will default to optimizing technical correctness over experiential quality."
What goes unrepresented
Two publics have zero representation in current eval frameworks. The regulator public asks about transparency and harm prevention. The ecosystem public tracks downstream effects on adjacent systems. No eval rubric I found addresses either.
Design carried the user public's perspective into the old spec process. The eval replaced that process. Without a new entry point, the user public loses its voice in the definition of quality. The regulator and ecosystem publics never had one.
Your eval defines what your product optimizes for. Which of your four publics wrote it?








.webp)
.webp)
.webp)

