Performance Reviews Are Broken: Here is a Simple Alternative
The performance review is the only meeting on the corporate calendar that runs on the open agreement of nearly everyone involved that it doesn't work. Managers draft them at 11pm the night before calibration. Employees skim for the rating and never open the document again. Then the next cycle arrives and the ritual runs again, on schedule, because nobody can say exactly what would break if it stopped.
Three things would break, and naming them is the way out. The review carries three different jobs stapled into one meeting: pay decisions, feedback, and a legal paper trail. Each job is legitimate. Merged, they quietly ruin each other.
Three jobs in one meeting
Job one is pay: someone has to decide who gets the raise and the promotion, and explain why. Job two is feedback: helping a person get better at the work. Job three is documentation: a dated record of what was said, for the day a dispute or a bad exit needs history.
Pay corrupts feedback first and worst. The moment a rating moves money, the review stops being a conversation and becomes a negotiation. Nobody writes "I struggled with estimation all year" in a self-assessment that prices their raise. Managers inflate the language to protect their people's compensation, then wonder why every document reads like a press release.
The legal job corrupts feedback from the other side. A document that might someday be read aloud in an employment dispute gets written for that audience: hedged, drained of specifics, safe rather than useful. "Opportunities for growth in stakeholder communication" is tribunal-proof and coaching-useless.
And the feedback ritual corrupts pay in return. Compensation deserves cold, consistent calibration against a level framework. Route it through individual narrative reviews and it becomes hostage to prose: the raise goes to whoever writes the best story about themselves, which measures self-promotion and not much else.
The annual cadence then finishes the damage. Feedback delivered eight months after the event is archaeology, not coaching.
Run the jobs separately
Feedback is continuous and cheap, and its natural unit is the 1:1. Say the thing within a week of the event, in specific terms, while everyone remembers what happened and the stakes are low. A person who hears about a problem in March can fix it by June. A person who hears about it in December has been billed for a year of something nobody mentioned.
Pay is a calibration exercise on its own calendar. Managers propose, leadership calibrates across teams against a level framework and a budget, and decisions come back with reasons. It needs consistency, not essays. Run it once or twice a year, tell people plainly how it works, and stop pretending the review essay drives it. In most companies everyone already suspects it doesn't.
Documentation is boring and factual, and boring is the standard to hold it to: dated notes on what was discussed and agreed. If there's a genuine performance problem, it needs its own explicit process with concrete expectations and dates, not a coded adjective buried in paragraph four of a review.
The small-company version fits on half a page
Under fifty people you don't need a 9-box or a 360 platform. You need a rhythm that guarantees the conversation happens, and a record that it did.
The version I'd defend: a written check-in each quarter, or twice a year if quarterly feels heavy. The employee and the manager answer the same two questions about the employee's work, in writing, separately, before they meet. Half a page maximum. Ten minutes of writing.
1. Since the last check-in, what piece of your work mattered most, and why?
2. What one thing should change before the next check-in, and what support does that change need?
Then 45 minutes together, comparing answers. The overlaps confirm what's working. The gaps are the agenda: the manager thinks the migration mattered most while the employee names a mentoring effort nobody noticed. The disagreements between the two half-pages will teach you more than any rating scale.
Notice what this quietly produces: the documentation job, done as a byproduct. Dated, factual, in both voices, written while the events were fresh. And notice what stays off the table: money. Pay gets decided at calibration time, informed by the whole trail of check-ins, never negotiated inside one.
The failure modes are old and well mapped
Recency bias. A year rated in November is really September to November rated in November. Shorter cycles shrink the window, and the written trail does the remembering that memory won't.
Central tendency. When differentiating feels socially expensive, everyone drifts to "meets expectations" and the record becomes noise. The cheapest fix at small scale is to stop producing scores entirely. Write claims instead: "shipped the migration early, underestimated the API work twice." A claim can't hide in the middle of a scale.
The surprise rule. This one is folk wisdom in every decent management book because it keeps being true: nothing in a review may be news. If a person first hears a serious criticism inside a formal review, the process failed months earlier. Reviews are where things get written down, not where they get revealed. Hearing "we've had concerns since spring" in December is how a fixable problem becomes a resignation.
You need tooling later than vendors say
At ten people, the right stack is a shared doc per person and a calendar reminder. At thirty it's still that. What fails at this size is never the software. It's a founder skipping the conversation for two quarters, and no platform fixes that.
Tooling earns its keep when coordination costs bite: a dozen managers running cycles that need synchronizing, calibration that wants structured history across a hundred people, retention rules that scattered docs can't guarantee. Around that point, buy the boring kind: it stores check-ins, nudges the late, and shows history side by side. Be suspicious of anything promising to mine insight out of the ratings themselves. The ratings were always the least informative part of the process.
AI can draft the words, not the noticing
Everyone has tried some version of the tempting prompt by now: "write a performance review for a senior engineer, strengths include..." and back comes something fluent and kind. That fluency is exactly the danger.
Used as an editor, AI is genuinely useful here. Paste in your own dated notes (the check-ins above hand you these for free), let it organize, cut the rambling, flag the sentence that reads harsher than you meant. Then verify every line against your memory of real events. That's an hour of typing saved, with the honest work intact.
Used as a substitute for observation, it's corrosive. When the model writes from a job title and a vague impression, you get a document with the shape of feedback and none of the content, and the employee can tell. Increasingly they respond in kind, running their self-assessment through the same tool, and the loop closes: two language models exchanging compliments about a year of work neither of them saw.
The test takes a minute. For every sentence in the draft, can you name the specific event behind it? If yes, the model saved you time. If no, it replaced the part that was the actual job.
The team is a different instrument
Everything above is an individual instrument, and some problems will never show up in it. An overloaded on-call rotation or a morale slide surfaces across a team long before it surfaces in any one person's check-in. That layer needs its own, blunter tools: we've written about anonymous team health checks and eNPS as a team-level number. Neither replaces the half-page conversation. They tell you which team needs it soonest.