The Evidence Standard for Performance Decisions

Most performance reviews claim to be “data-driven.” Here’s a practical hierarchy for what should actually count as evidence, and what’s just opinion.

Catch Up AI Team

The Evidence Standard for Performance Decisions

Evidence in performance decisions: what should actually count

Evidence in a performance decision should be ranked, not treated as equal: verified outcomes and completed work carry the most weight, followed by structured feedback collected across the review period, relevant work context, the employee’s own account, and finally the manager’s interpretation, which matters, but should not be the only input.

Most reviews invert this order by default, leaning almost entirely on what a manager remembers.

The Evidence Standard for Performance Decisions: What Counts, What Doesn’t

Almost every performance management vendor describes its product as “data-driven.” Almost none define what data actually means in a review, or how it should be weighted against a manager’s memory and opinion.

That gap is where recency bias, halo effects, and inconsistent ratings live.

Research on manager rating behavior is blunt about the scale of the problem. A landmark 2000 study published in Journal of Applied Psychology by Scullen, Mount, and Goff found that idiosyncratic rater effects accounted for 62% of the variance in performance ratings — more than double the 21% attributable to actual job performance.

Separately, surveys of managers find that a large majority admit their ratings are shaped more by what an employee did in the last few weeks than across the full review period.

Neither finding means managers are acting in bad faith. It means unaided human memory is a poor instrument for evaluating a year of work, and most review processes do not compensate for that.

A five-level evidence hierarchy

Rather than treating “data” as one undifferentiated pile, it helps to rank inputs by how much weight they should carry in a decision.

  1. Verified outcomes and completed work. Goals hit or missed, deliverables shipped, and metrics moved. This is the closest thing to ground truth, though it still needs context.
  2. Structured feedback collected across the period. Peer input, upward feedback, and check-in notes gathered continuously, not reconstructed from memory the week before a review.
  3. Work context. Team changes, shifting priorities, external dependencies, and constraints that explain why an outcome looks the way it does. Outcomes without context can be misleading in either direction.
  4. The employee’s own account. Self-assessment matters, particularly for surfacing work that was not visible to the manager, but it is naturally self-interested and should not stand alone.
  5. Manager interpretation. Judgment is still necessary. Someone has to weigh the other four levels, but it should be the synthesis step, not the starting and ending point.

Most annual review processes run this hierarchy backwards: managers write from memory with maybe a self-assessment attached, and call it evidence-based because a form asked for “specific examples.”

The same situation, viewed at level 1 versus level 5

The gap between the top and bottom of the hierarchy is easier to see next to a single situation than in the abstract.

Take a missed deadline, a common enough event that most reviews end up characterizing it one way or another.

At level 1, verified outcomes, the record shows a deliverable was due on a Friday and shipped the following Wednesday, four business days late.

That fact alone is neutral. It does not say whether the delay mattered, who caused it, or whether it was foreseeable.

Add level 3, work context, and the picture changes: a dependency team delivered its part of the work eight days late, and the four-day slip is what happened after the employee absorbed most of that lost time.

Without the context, the outcome alone reads as a missed deadline. With it, the same outcome reads as a recovery.

At level 5, manager memory unaided, the same situation might get written up months later as “struggled to hit deadlines this quarter,” because the missed date is what stuck, and the dependency delay that explains it was never logged anywhere and has since been forgotten.

Nothing about that description is dishonest. It is just working from a thinner, less accurate slice of what actually happened, and the rating built on it will reflect that thinner slice rather than the full picture.

The point is not that level 5 is wrong and level 1 is right. A bare outcome without context can mislead just as easily as unaided memory can.

The point is that a rating built from levels 1 through 3 together is working from more of the real situation than a rating built from level 5 alone, and most annual reviews default to exactly the input that carries the least information.

This is also where an AI drafting tool can make things worse rather than better if it is not deliberate about which level it is pulling from.

An AI system asked to summarize “how the deadline went” will happily generate fluent prose from whatever it is fed: a level-5 recollection typed in by a manager, or the level-1 through level-3 record if that is what is available.

The tool cannot tell the difference between a well-supported account and a thin one. It can only make either one read more confidently.

That is a reason to be more deliberate about the input, not less, when AI drafting is part of the process.

For a closer comparison of drafting tools versus evidence-grounded review systems, read AI Performance Review Software: Drafts vs. Evidence.

What does not count as evidence

A few categories deserve explicit exclusion from any evidence standard, because they show up in reviews disguised as data:

  • Activity volume. Hours logged, messages sent, and “always online” signals measure presence, not contribution.
  • A single standout moment. One unusually good or bad moment should not be generalized into an overall rating.
  • Unstructured recall gathered for the first time during the review itself. Evidence should be logged as it happens, not reconstructed for the first time when the review is due.

These categories can still provide context, but they should not carry the same weight as verified outcomes, structured feedback, and documented work context.

Why this matters for AI-assisted reviews specifically

As more performance tools add AI drafting features, the evidence question becomes more urgent, not less.

An AI system that drafts fluent, confident-sounding prose from whatever inputs it is given will produce a fluent, confident-sounding draft regardless of whether those inputs were level-1 evidence or level-5 opinion.

Fluency is not a proxy for accuracy.

Organizations adopting AI-assisted reviews should decide their evidence hierarchy before choosing a tool, not after.

The Catch Up AI perspective

Catch Up AI’s Platform is built around this same premise: performance decisions are better when they are grounded in a broader base of evidence connected work signals, structured feedback, and manager context rather than a single manager’s memory of the last few weeks.

Review Now AI drafts are explicitly meant to sit inside that hierarchy: a structured starting point built from real inputs, still requiring a human to weigh context and take ownership of the final call.

Conclusion

“Data-driven” is a claim every vendor makes and almost none defines.

Before adopting any performance tool, AI-assisted or not, decide what counts as evidence in your organization, rank it, and hold every input, including AI-generated ones, to that standard.

FAQs

What counts as evidence in a performance review?

Verified outcomes, structured feedback collected over time, and relevant work context carry the most weight. Self-assessment and manager interpretation matter, but should not stand alone.

How do you reduce recency bias without removing manager judgment?

By requiring evidence to be logged continuously throughout the review period rather than reconstructed from memory at review time, so a manager’s synthesis has more than the last few weeks to work from.

What is the difference between an opinion and evidence in a review?

Evidence is verifiable: an outcome, a dated piece of feedback, or a documented change in context. An opinion is a manager’s unaided interpretation, which is necessary but should not be mistaken for the underlying data.

How many sources of evidence should a performance decision use?

There is no fixed number, but a decision built on only one source, usually manager memory, is the pattern most associated with bias and inconsistency. Combining outcomes, structured feedback, and context is more defensible.

Does AI automatically make a performance review more evidence-based?

No. AI can draft fluent text from any input, including low-quality ones. The evidence standard has to be applied to what feeds the AI, not assumed because AI was involved.