Quick answerOnce you have decided, as a governance matter, whether AI-agent-assisted output can inform an employment decision at all, a separate and harder question follows: how do you evaluate that employee fairly once it is allowed. Update the review rubric to measure judgment, escalation quality, and how well the employee directs and corrects the agent, not raw output volume or polish, since the agent inflates both regardless of the underlying employee's skill. Recalibrate the bar periodically as the agent's own capability improves, so employees are not quietly held to a rising standard that reflects the tool getting better, not them.
A fairness question, not a consent question
A related post in this series covers the governance boundary of whether AI agent interaction data can be used in an employment decision at all, a consent and data-use question that has to be settled with HR and legal before any of this applies. This post assumes that boundary has already been drawn and starts one step later: given that a manager is now looking at AI-agent-assisted output as part of a review, how do they evaluate the human fairly, without either crediting the employee for what the tool did or penalizing them for the tool's limitations.
Why raw output metrics stop being a fair signal
Before an AI agent assisted the role, output volume and polish were reasonable proxies for skill, since producing more, cleaner work took more underlying ability. Once an agent drafts, formats, or accelerates that same output, volume and polish measure tool adoption as much as skill, and two employees with very different judgment can produce visually similar output if both are competent at directing the same agent. Continuing to score primarily on those old proxies rewards employees who happen to lean on the tool well for reasons unrelated to the underlying skill the review is supposed to measure.
What to measure instead
Shift review weight toward the parts of the work the agent cannot do for the employee: recognizing when the agent's output is wrong or off-target before it goes out, knowing when a task should not be delegated to the agent at all, and the judgment involved in editing, correcting, or escalating rather than accepting the agent's first pass. These are the skills that actually differentiate employees once the agent has leveled the baseline output quality across the team, and they are also the skills most exposed if the agent is later restricted, upgraded, or removed.
The moving-target problem
Agent capability improves over time, often without the employee doing anything differently, which means an employee's assisted output can look meaningfully better this quarter than last quarter for reasons that have nothing to do with their own growth. A review process that does not account for this risks quietly raising the bar on employees every time the underlying tool improves, then treating flat individual growth as a performance problem when it is actually the agent, not the employee, that stayed the same relative to a rising baseline.
Make the criteria change visible to employees
Employees should know explicitly that the review rubric has shifted toward judgment and correction rather than raw output, and why, before that shift shows up in a review score. Introducing this quietly, only surfacing it when an employee is surprised by a lower score than their visible output volume would have predicted under the old rubric, reads as moving the goalposts rather than as a considered response to a tool that changed what a review needs to measure.
FAQ
Should employees be told the review criteria changed because of the AI agent?
Yes. Communicate the shift in advance of any review cycle it affects, along with the reasoning, rather than letting employees discover it through a review score that surprises them.
Does this apply differently to new hires than tenured employees?
It is worth being especially deliberate with new hires, since they may never have experienced the pre-agent baseline and can unintentionally be evaluated against a moving target with no prior reference point to notice the shift.
Who should own updating the rubric as the agent improves?
The manager closest to the role should propose the specific criteria, but HR should own reviewing the rubric on a recurring cadence to make sure it keeps pace with how much the agent's capability has actually changed.

