Two engineers. Same team, same tenure, same performance cycle. One gets promoted, one does not.
Ask why and you will get an answer with the following shape: she operates at a higher level, she has more impact, she shows more ownership, she is more senior in how she approaches problems. Ask what any of that means and you will get the same four words rearranged. Ask what the other engineer would have to do differently, and you will watch a competent adult discover, live, that they cannot say.
I want to be careful here, because I am not accusing anyone of dishonesty. The manager in that scenario usually has a real and accurate perception. The two engineers genuinely are different. The problem is that the perception exists in a form that cannot be transmitted, checked, disputed, or acted upon. It is a feeling in a person's head wearing a business suit.
And a feeling in a person's head is not an altitude marker. It is a guess about altitude made by somebody standing next to you, at the same altitude, in fog.
You cannot climb toward an adjective. "More senior" is not a destination. It is a compliment somebody pays you after the fact, and it will not tell you a single thing to do on Tuesday.

The problem, stated precisely
Part 1 ended on three questions that organizations cannot answer: how do we define paths, how do we set expectations, how do we measure. Part 3 built the map, which handles the terrain. This part handles the other two, and they turn out to be the same problem in a trench coat.
The problem is language. We do not have words for "better" that mean the same thing to two people. So we borrow words from adjacent concepts, and every one of them fails in the same way.
We borrow from time. Three years is mid, six years is senior. Simple, universally understood, and completely fraudulent, because it measures the ride and not the climb. This is the Escalator wearing a badge.
We borrow from scope. Owns a service, owns a system, owns a domain. Better, but it measures what somebody was handed as much as what they can do, and it quietly punishes anyone on a small but critical piece of ground.
We borrow from adjectives. Strong, deep, strategic, mature. These feel the most descriptive and are the most useless, because every one of them resolves to the speaker's private standard.
What all three share is that none of them describe a behavior. And behavior is the only thing anyone has ever been able to observe, teach, practice, or fairly assess.
Two structural moves before the language
Before the vocabulary, two structural decisions, both of which will feel like a demotion to somebody and are worth it anyway.
First: a small set of core jobs, not a ladder of titles. Most organizations have accumulated a title taxonomy, twelve rungs deep, invented incrementally to solve retention conversations. It describes nothing. Replace it with a handful of core jobs that genuinely exist in your delivery model. Product manager. Analyst. Engineer. DevOps. Four, six, eight, but a number you can say out loud without checking a document.
Second: within each job, defined levels of behavior. Not levels of seniority. Levels of what the person can be observed doing. The career path is then the honest thing it always should have been: you get better at a job, visibly, and the levels describe how far you have come. Moving between jobs is a lateral traverse, not a demotion, and both are legitimate routes.
That is the whole reframe, and it does more work than it looks like. It kills the idea that growth means acquiring a title, and replaces it with the idea that growth means acquiring capability. Those are not the same thing, and every organization that conflates them ends up with senior people who are senior at nothing in particular.
Steal the language from education
Now the vocabulary problem, which has an unglamorous and thoroughly solved answer that our industry ignored for fifty years because it came from the wrong building.
Educators have needed to describe levels of mastery, precisely and comparably, for a very long time. And they built a taxonomy for it. Bloom's, if you want the name. Its insight is almost embarrassingly simple: describe mastery with verbs, and organize the verbs into levels.
At the bottom, a person can recall a thing. Above that, explain it. Above that, apply it in a familiar situation. Above that, analyze a situation to work out which parts of it apply. Above that, evaluate, meaning judge between options and defend the judgment. And at the top, create, meaning produce something new and correct that did not exist before.
Look at what happened there. Every one of those is a verb. Every one is observable. And critically, they are ordered, so a person at "explain" can see that the next thing is "apply," which is a specific and achievable act, in a way that "be more senior" is never going to be.
Apply that to a topic on the map from Part 3, and the difference is stark. Instead of "strong at requirements," you get a ladder that reads: can describe what a well formed requirement contains, can explain why a given requirement is ambiguous, can write requirements for a familiar feature, can decompose an unfamiliar business problem into requirements, can evaluate a colleague's requirements and defend the critique, can design the requirements approach for a new domain.
Any one of those is a thing you can watch a person do. Any one of them is a thing a person can practice on purpose. And two people looking at the same engineer will agree on which rung they are on far more often than they will agree on whether she is "strong."
Acceptance criteria, for humans
Here is where the whole thing gets its teeth, and where I get to make my favorite point in this series.
You already know how to do this. You do it every day, to software.
Nobody on your team would accept a story that said "make the search better." They would demand acceptance criteria: specific conditions, written before the work, that determine whether it is done. Not because engineers are pedantic (they are, but that is not why), but because you learned, expensively, that "done" is unarguable only when it is defined in advance.
Then those same people walk into a performance conversation and cheerfully assess an entire human being against "make the engineer better."
So: for each topic on the map, at each level of behavior, write acceptance criteria. What specifically must a person be able to do to be at this level of this topic? Written down, in advance, visible to the person being assessed.
The effects are immediate and slightly uncomfortable.
Expectations become documents instead of opinions. The Part 1 question finally has an answer, and it is a list, and the list is finite. That is the single most motivating thing you can hand a person who wants to grow: proof that the distance is bounded.
Assessment becomes evidence. The conversation stops being "do I think you are senior" and becomes "here are the criteria, here is what we have observed, here is where you are." Disagreements do not vanish, but they relocate to the evidence, where they can be settled, instead of orbiting the manager's regard, where they cannot.
Fairness stops depending on the assessor's character. Written criteria are the same for the loud person and the quiet one, the one who reminds you of yourself and the one who does not. I made this argument about scorecards in Part 2 and it is the same argument here, because it is the same mechanism. Specificity is very hard to argue with after the fact.
I made a version of this case in Building on Solid Ground, about audit: write down what happened, as it happens, so you can answer with a query instead of a panic. This is the same instinct pointed at people. Define what done means before you argue about whether it is done.
On making it a game, carefully
Behavior levels have an obvious gamified quality. Verbs, levels, progress, unlocking the next rung. Used well, this is genuinely motivating, because it turns a vague ambition into a visible sequence of achievable moves, and people will climb a visible sequence with real enthusiasm.
Used badly, it becomes a points system that people optimize instead of a capability model that people grow into. The failure mode is entirely predictable: the moment a level is tied mechanically to compensation, the map becomes a currency and everybody starts farming it.
The guard is simple. Levels describe capability. Humans, with judgment, make decisions about roles and money, informed by levels but not automated by them. Keep the measurement honest and it stays useful. Wire it directly to the payroll system and you have built a very expensive machine for producing self reported mastery.
Try the two engineer test
Take the two engineers from the opening. Yours, not mine. You have a pair.
Now write down, in verbs, what the promoted one can do that the other cannot. Not adjectives. Verbs, attached to specific topics. She can decompose an unfamiliar domain into a delivery plan. He can execute a plan somebody else decomposed.
If you can write that in ten minutes, your organization is in far better shape than most, and the only thing missing is writing it down where the second engineer can read it.
If you cannot, that is the finding, and it is the most common finding there is. It also means that every promotion decision your organization has made was made in fog, including the good ones. The good ones were luck. Luck is not a system.
GSD's Maturity Assessment exists to replace that fog with a reading, looking at role behaviors as they actually are rather than as the title taxonomy claims. Find the gaps first. Then close them. In that order, always.
We now have a party, a map, and markers. Which is everything a person needs to climb, except the one thing that turns out to matter most, and it is not a system at all.
Next week: the guide.
Nobody summits alone.
Next in the series, Part 5: "The Guide: Nobody Summits Alone."
*Can your people describe their next level in verbs? GSD's Maturity Assessment is the entry point, a structured look at where your delivery organization really stands, people included, before you invest in closing the gaps.*
Back to Blog