AI Performance Review Software: Drafts vs. Evidence

Not every “AI-powered” performance review tool does the same job. Here’s how to tell an evidence-based system from one that just writes fast.

Catch Up AI Team

AI Performance Review Software: Drafts vs. Evidence

AI-powered performance reviews: the difference buyers need to understand

“AI-powered performance reviews” can mean two very different things: software that drafts polished text from whatever a manager types in, or software that grounds a draft in verified work signals, structured feedback, and goals, then requires a human to review it before anything is final.

The first saves typing time. The second changes the quality of the decision. Buyers should ask which one they are evaluating before comparing price or interface.

AI Performance Reviews: Drafts vs. Evidence-Based Systems

Every major HR platform now claims “AI-powered performance reviews.” At Lattice’s Lattiverse conference in June 2026, CEO Sarah Franklin went further, announcing Lattice MCP, a way to expose performance data to AI tools like Claude and ChatGPT, alongside “AI Leverage Insights,” a feature that connects AI-tool usage to actual performance outcomes rather than just token counts.

The message was pointed: usage data alone is not a quality signal.

That distinction matters more than it might seem, because “AI-powered” performance review software has quietly split into two categories that get marketed almost identically.

Category one: AI that drafts

Most AI review features work the same way. A manager pastes in some notes, selects a few goals, and the tool returns a polished paragraph.

It saves time on a task nobody enjoys, turning fragments into readable prose.

Lattice AI, 15Five’s AMAYA, and similar assistants are explicit that the AI does not evaluate performance or make the promotion, rating, or termination call; it summarizes what a human already believes and helps phrase it.

That is a legitimate, useful feature. It is also, fundamentally, a writing tool.

Category two: AI that is grounded in evidence

The second category starts somewhere different: not with what the manager types, but with what actually happened.

That means connected signals, goals, completed work, structured feedback collected throughout the period, and manager notes, feeding a draft the manager still has to review, edit, and take ownership of before it is final.

The difference is not “AI vs. no AI.” It is what the AI is working from, and how much a human is expected to check before the words become a decision that affects someone’s pay or career.

What “grounded in evidence” looks like in practice

The distinction is easiest to see next to a real draft, so it is worth walking through what “connected signals” actually means rather than leaving it abstract.

Take a completed project milestone.

An ungrounded draft might say what a manager typed in: “Shipped the Q2 integration on time, good work.”

An evidence-grounded draft instead starts from the actual milestone record, the date the integration was marked complete, whether it slipped from its original target, and which linked tickets or goals it closed out.

The AI is not inventing detail. It is pulling from a record that already exists in the systems a team uses, so the resulting language can say something more specific, like naming the dependency that was at risk and how it was resolved, instead of a generic compliment that could describe almost any project.

Structured peer feedback works the same way.

A manager typing from memory might recall that “a couple of people mentioned she is helpful in code review.”

An evidence-grounded system instead draws on the actual peer feedback collected during the period: who gave it, when, and what specifically they described, such as catching a race condition before it shipped or walking a newer engineer through a debugging approach.

That peer input existed already. The difference is whether the draft is built from it directly or from a manager’s secondhand, months-later paraphrase of it.

A documented goal check-in is a third example.

An ungrounded draft tends to restate the goal itself, such as “made progress on the onboarding redesign goal,” because that is what the manager has in front of them.

An evidence-grounded draft can instead reference the actual check-in notes logged against that goal throughout the quarter: what was flagged as blocked in April, what changed by June, and whether the final state matches what was actually recorded rather than what is remembered at review time.

Memory compresses and flattens a quarter’s worth of change into a single end-of-period impression. A documented trail preserves the shape of how the work actually went.

None of this requires the AI to make a judgment call about whether the milestone, the peer feedback, or the goal progress was actually good.

It still requires a manager to weigh that.

What it changes is the raw material the draft is built from: a documented record instead of whatever happens to be top of mind when the manager sits down to write.

That is the practical difference between the two categories, and it is the difference buyers are trying to get at when they ask what a vendor’s “AI-powered” feature is actually doing.

This is also why the first question in the buyer’s checklist below, what the draft is grounded in, is the one worth spending the most time on.

A vendor demo can make either category look similar on screen. The milestone record, the peer-feedback log, and the goal check-in trail are the parts that do not show up unless someone asks to see them.

Why this distinction is showing up now

Two forces are pushing this into the open.

First, competitive pressure: as every vendor adds an “AI-powered” badge, buyers increasingly ask what is under the hood, and vendors that can show their AI is grounded in real signals, not just prompted by whatever a manager types, have a genuine story to tell.

Second, research on how performance ratings actually form keeps surfacing an uncomfortable number: a landmark 2000 study published in Journal of Applied Psychology by Scullen, Mount, and Goff found that idiosyncratic rater effects accounted for 62% of the variance in performance ratings, more than double the 21% attributable to actual job performance.

If an AI tool simply reflects a manager’s impressions back in nicer prose, it can reproduce that same distortion faster and more confidently than before.

A buyer’s checklist: what to ask before “AI-powered” means anything

Use this five-question framework with any vendor, including Catch Up AI, before evaluating price or interface.

  1. What is the draft grounded in? Only manager-typed notes, or connected work signals plus structured feedback collected over time?
  2. Is there a confidence or explainability layer? Can the tool show why it drafted what it drafted, or is it a black box?
  3. Is human review required, or optional? A tool that lets a draft go out unedited is a different risk profile than one built around mandatory review.
  4. Does it support calibration across managers? Individual review quality matters less if ratings still cannot be compared fairly across a team.
  5. What does the vendor say it will not do? A vendor that is honest about limits, no certainty claims and no automated final decisions, is more trustworthy than one that implies the AI “knows” performance.

The Catch Up AI perspective

Review Now AI is built for the second category: it produces a structured draft grounded in manager notes and, where connected, real work data, not a decision.

The output is explicitly a starting point, not a finished review. A manager still needs to read it, adjust it for context the system cannot see, and take responsibility for what goes to the employee.

That is a deliberate design choice, not a limitation to apologize for. A review is a human judgment about a person’s work, and the goal of AI here is to make that judgment better-informed, not to remove it.

For teams looking at the broader system behind these signals, the Catch Up AI Platform explains how connected work signals, manager context, and review-ready narratives fit together.

Conclusion

“AI-powered” has stopped telling buyers anything useful on its own.

The real question is what the AI is grounded in and how much human review stands between a draft and a decision.

Ask the five questions above of any vendor you are evaluating, including us.

FAQs

Can AI write a performance review by itself?

AI can draft the language of a review, but a responsible system still requires a manager to review, edit, and take ownership of it before it is delivered. No credible vendor claims an AI-written review should go out unedited.

What is the difference between an AI-drafted review and an evidence-based one?

An AI-drafted review turns whatever a manager types into polished prose. An evidence-based review is grounded in connected work signals and structured feedback collected over time, with the manager’s input as one input among several rather than the only one.

Should managers edit AI-drafted reviews before sending them?

Yes. Every credible vendor, including Catch Up AI, treats an AI draft as a starting point that requires manager review, context, and edits, not a finished document.

What is a good calibration workflow when AI drafts are involved?

Calibration should happen after individual drafts are reviewed, comparing ratings and language across managers and teams for consistency, the same as it would with human-written drafts. AI assistance does not remove the need for calibration.

It can, particularly for organizations with EU-based employees given the EU AI Act’s treatment of performance-evaluation systems as high-risk. Organizations should confirm current requirements with legal counsel rather than relying on vendor assurances alone.