
Using Lattice? Add Live Signals, Early Alerts, and Timely Manager Action
Catch Up AI integrates with Lattice to add live signals, early alerts, private check-ins, and timely manager action to existing people workflows.
Most performance reviews claim to be “data-driven.” Here’s a practical hierarchy for what should actually count as evidence, and what’s just opinion.

Evidence in a performance decision should be ranked, not treated as equal: verified outcomes and completed work carry the most weight, followed by structured feedback collected across the review period, relevant work context, the employee’s own account, and finally the manager’s interpretation, which matters, but should not be the only input.
Most reviews invert this order by default, leaning almost entirely on what a manager remembers.
Almost every performance management vendor describes its product as “data-driven.” Almost none define what data actually means in a review, or how it should be weighted against a manager’s memory and opinion.
That gap is where recency bias, halo effects, and inconsistent ratings live.
Research on manager rating behavior is blunt about the scale of the problem. A landmark 2000 study published in Journal of Applied Psychology by Scullen, Mount, and Goff found that idiosyncratic rater effects accounted for 62% of the variance in performance ratings — more than double the 21% attributable to actual job performance.
Separately, surveys of managers find that a large majority admit their ratings are shaped more by what an employee did in the last few weeks than across the full review period.
Neither finding means managers are acting in bad faith. It means unaided human memory is a poor instrument for evaluating a year of work, and most review processes do not compensate for that.
Rather than treating “data” as one undifferentiated pile, it helps to rank inputs by how much weight they should carry in a decision.
Most annual review processes run this hierarchy backwards: managers write from memory with maybe a self-assessment attached, and call it evidence-based because a form asked for “specific examples.”
The gap between the top and bottom of the hierarchy is easier to see next to a single situation than in the abstract.
Take a missed deadline, a common enough event that most reviews end up characterizing it one way or another.
At level 1, verified outcomes, the record shows a deliverable was due on a Friday and shipped the following Wednesday, four business days late.
That fact alone is neutral. It does not say whether the delay mattered, who caused it, or whether it was foreseeable.
Add level 3, work context, and the picture changes: a dependency team delivered its part of the work eight days late, and the four-day slip is what happened after the employee absorbed most of that lost time.
Without the context, the outcome alone reads as a missed deadline. With it, the same outcome reads as a recovery.
At level 5, manager memory unaided, the same situation might get written up months later as “struggled to hit deadlines this quarter,” because the missed date is what stuck, and the dependency delay that explains it was never logged anywhere and has since been forgotten.
Nothing about that description is dishonest. It is just working from a thinner, less accurate slice of what actually happened, and the rating built on it will reflect that thinner slice rather than the full picture.
The point is not that level 5 is wrong and level 1 is right. A bare outcome without context can mislead just as easily as unaided memory can.
The point is that a rating built from levels 1 through 3 together is working from more of the real situation than a rating built from level 5 alone, and most annual reviews default to exactly the input that carries the least information.
This is also where an AI drafting tool can make things worse rather than better if it is not deliberate about which level it is pulling from.
An AI system asked to summarize “how the deadline went” will happily generate fluent prose from whatever it is fed: a level-5 recollection typed in by a manager, or the level-1 through level-3 record if that is what is available.
The tool cannot tell the difference between a well-supported account and a thin one. It can only make either one read more confidently.
That is a reason to be more deliberate about the input, not less, when AI drafting is part of the process.
For a closer comparison of drafting tools versus evidence-grounded review systems, read AI Performance Review Software: Drafts vs. Evidence.
A few categories deserve explicit exclusion from any evidence standard, because they show up in reviews disguised as data:
These categories can still provide context, but they should not carry the same weight as verified outcomes, structured feedback, and documented work context.
As more performance tools add AI drafting features, the evidence question becomes more urgent, not less.
An AI system that drafts fluent, confident-sounding prose from whatever inputs it is given will produce a fluent, confident-sounding draft regardless of whether those inputs were level-1 evidence or level-5 opinion.
Fluency is not a proxy for accuracy.
Organizations adopting AI-assisted reviews should decide their evidence hierarchy before choosing a tool, not after.
Catch Up AI’s Platform is built around this same premise: performance decisions are better when they are grounded in a broader base of evidence connected work signals, structured feedback, and manager context rather than a single manager’s memory of the last few weeks.
Review Now AI drafts are explicitly meant to sit inside that hierarchy: a structured starting point built from real inputs, still requiring a human to weigh context and take ownership of the final call.
“Data-driven” is a claim every vendor makes and almost none defines.
Before adopting any performance tool, AI-assisted or not, decide what counts as evidence in your organization, rank it, and hold every input, including AI-generated ones, to that standard.
Verified outcomes, structured feedback collected over time, and relevant work context carry the most weight. Self-assessment and manager interpretation matter, but should not stand alone.
By requiring evidence to be logged continuously throughout the review period rather than reconstructed from memory at review time, so a manager’s synthesis has more than the last few weeks to work from.
Evidence is verifiable: an outcome, a dated piece of feedback, or a documented change in context. An opinion is a manager’s unaided interpretation, which is necessary but should not be mistaken for the underlying data.
There is no fixed number, but a decision built on only one source, usually manager memory, is the pattern most associated with bias and inconsistency. Combining outcomes, structured feedback, and context is more defensible.
No. AI can draft fluent text from any input, including low-quality ones. The evidence standard has to be applied to what feeds the AI, not assumed because AI was involved.
Continue reading more insights from CatchUp AI