Levels become falsifiable.
The signature test replaces adjective rubrics with a commitment the organization acts on, so miscalibrated promotions surface as evidence rather than lingering as politics.
IH Tech · Field Guide 002 · https://ihtech.ba/field-guides/ai-native-leveling-ladder/
FG-002
Field Guide 002 · IH Tech
A custody-based leveling and hiring framework for the AI era: levels defined by what you can be accountable for, a standardized assessment cockpit, and a protected formation budget that keeps the junior-to-senior pipeline alive.
You own an engineering ladder and an interview loop, and both were written for a world where engineers produced the code. That world is ending in your own repositories right now. Agents write a growing share of the diffs; your seniors spend their days directing and validating rather than typing; your juniors can ship work that looks two levels above their experience. Meanwhile your leveling document still prices people by the code they personally produce, and your interview loop still assumes the candidate wrote every line in front of you.
The question is not whether to change the ladder. It is what to change it to, without breaking the three things a ladder exists to protect: fair assessment of candidates, honest measurement of seniority, and a pipeline that reliably turns juniors into the seniors you will need in five years. The current wave of AI adoption quietly attacks all three at once, and most redesigns we see fix one facet while making another worse.
This guide gives you one framework that holds all three together, the evidence for why each is genuinely at risk, the trade-offs you accept, the failure signals to watch, and the specific moves to make this quarter.
Rebuild the ladder around custody, not production: a level is the largest unit of work a person can be accountable for shipping without someone above them re-checking it.
Then make three moves, together, because each one covers a hole the other two open.
01
The signature test sets the level, the slop test proves the judgment behind it, and the reconstruction test proves the judgment has a floor when the agent is wrong.
02
Any interview where a personal AI subscription changes the outcome is measuring disposable income, not engineering ability. Standardize the model, the budget and the harness for every candidate, and assess the delta the human adds.
03
Protect a budgeted share of junior time for manual reps that agents would otherwise absorb, run mentorship as a designed system with named owners, and make the formation record a promotion requirement. This is the capital investment that keeps the senior pipeline producing.
Three separate pressures landed on the ladder at once. Each has been reported on its own this year; the danger is that they interact.
Pressure 01 · The wallet
Interviews have started to price the candidate's wallet. Engineering leaders increasingly probe agentic-tool proficiency in interviews, and that proficiency is bought: premium agentic tools run to real money per month, and building fluency also takes unpaid personal time that candidates with caregiving duties or second jobs do not have. LeadDev reported the pattern in August 2026, including an engineering manager spending about $100 a month across agentic tools to stay current, a student plan capped at 200 monthly credits that users called too small to practice on, and an intern candidate who could not afford any paid tool at all (Lazzaro, LeadDev, 2026-08-12, checked 2026-08-16)[1]. One practitioner in that piece named the compounding effect: the people who can afford the tokens to experiment early build the delegation skills that make them more hirable, which buys them more access. An interview loop that assumes personal tool access converts a temporary resource gap into a permanent career filter, and it does so invisibly, because the rejected candidates simply look worse at the exercise.
Pressure 02 · Production volume
Production volume no longer separates levels. When an agent produces the diff, the diff's existence tells you little about the person who shipped it. What separates engineers now is the quality of the decisions around the diff: what to build, what to accept, what to reject, what the agent got plausibly wrong. The InfoQ Culture and Methods trends report describes the role shift directly: teams compressing toward one or two people plus a swarm of agents, and engineers acting as custodians who build guardrails and validate AI output rather than contributing most of the code themselves (InfoQ Trends Report, 2026-08-07, checked 2026-08-16)[5]. The same report records the paradox that makes validation a real skill rather than a formality: adoption keeps accelerating while 96% of surveyed developers say they distrust AI outputs. A ladder that still awards levels for output volume is awarding them for the agent's work.
Pressure 03 · The pipeline
The pipeline that produces validators is eroding underneath you. Here is the quiet structural problem. Supervising agent output well requires exactly the intuition that engineers historically built by doing the low-level work agents now absorb. InfoQ's report on Alasdair Allan's analysis states the loop plainly: AI removes the entry-level learning experiences, juniors lean on AI summaries instead of years of reading legacy code and debugging production at 3 a.m., and the field therefore produces fewer of the seniors who can supervise AI (Linders reporting Allan, InfoQ, 2026-08-13, checked 2026-08-16)[3]. An organization can run on its existing seniors for years while this happens, which is why the erosion is silent. The day it becomes visible is the day it is expensive.
Two further findings shape the fix. Hands-on practice is what keeps validation judgment sharp even in managers: LeadDev's August piece on building managers argues that the ability to recognize AI slop develops only through direct building, because AI output sits at the median of human taste and only practiced eyes see the gap (Rungta, LeadDev, 2026-08-13, checked 2026-08-16)[2]. And development of juniors was never reliably delivered by office proximity in the first place: cited research shows employees whose managers actively participate in onboarding are 3.5 times more likely to be satisfied, while 18% of remote workers report few or no check-ins at all, which argues for designed development systems over osmosis (Korducki, LeadDev, 2026-08-03, checked 2026-08-16)[4].
Each of these pieces isolates one facet. Fix them separately and they conflict: a judgment-only ladder with no formation budget accelerates the pipeline erosion; a formation program with production-denominated levels punishes the juniors doing the reps; an equity-minded interview loop that bans AI tools entirely stops measuring the job as it now exists. The framework below is built so the three fixes reinforce instead.
The framework has one denominating idea, three axes it measures, four tiers it defines, and three tests that make the axes observable. It is designed to be applied to the ladder you already have, not to replace your titles.
Custody is the largest unit of work a person can be accountable for shipping without someone above them re-validating it.
Not the largest thing they can produce; agents inflated production for everyone simultaneously, which is precisely why it stopped carrying signal. Custody is a claim about trust under absence: what could this person sign off on if nobody checked behind them, and the organization would be right to let them.
Custody is measurable in a way "impact" and "scope" never quite were, because it has a falsifiable form: for any unit of work, either you would let this person be the last pair of eyes on it or you would not.
Every level is assessed on the same three axes. A level requires all three at that tier; the axes are deliberately non-substitutable, because each one covers a specific failure the other two permit.
Axis 01
The size of the unit the person can take custody of: a bounded change, a component or workstream, a system with its dependencies, or the engineering process itself. This is the level's primary denomination.
Axis 02
Decision quality when directing and adjudicating agent output at that span: knowing what to ask for, what to accept, what to reject, and which flagged difference is a bug versus load-bearing behavior. This is the skill the custodian role actually exercises day to day.
Axis 03
Evidence that the judgment has a manual floor: the person has done enough of the underlying work without agent execution that they can still reason when the agent is wrong, unavailable, or confidently misleading. This axis exists because judgment without formation is pattern-matching on outputs, and it collapses exactly at the novel cases where custody matters most.
Map these onto your existing titles; the names describe the custody, not the business card.
| Tier | Axis 01 Custody span Sets the span · Signature test | Axis 02 Judgment under agency Proves judgment · Slop test | Axis 03 Formation depth Proves the floor · Reconstruction test |
|---|---|---|---|
| Tier 04 Process custody Staff, principal |
How the organization builds: plan templates, verification harnesses, the ladder itself |
The agent process, not individual agents; designs what others direct |
Broad floor across domains; audits whether the org's judgment still has one |
| Tier 03 System custody Senior |
A component or system, including all agent output that ships inside it Above the level |
Multi-step agent work; adjudicates the flags that matter; teaches Above the level |
Deep floor in their domain; runs formation for apprentices as a named duty |
|
Tier 02
Independent custody
Mid-level
Custody line |
Changes without review above them; a component with review |
Agents through their own plans; adjudicates routine flags |
Actively banking reps in unfamiliar territory; can reconstruct their own accepted work |
| Tier 01 Apprentice custody Junior |
A bounded change, with review above them |
An agent on scoped tasks, with their direction itself reviewed |
Formation-heavy: protected manual reps are a core duty, not spare-time study |
Two properties of this table do real work. First, teaching appears as a custody duty from system custody upward, which makes the formation budget staffable instead of aspirational. Second, process custody explicitly owns the ladder and the verification harnesses, so the framework has a named maintainer inside your organization once we hand it over.
The axes become observable through three tests. Use all three at every promotion, and the first two in hiring.
Test 01 · Signature
For a concrete unit of work at the candidate tier, ask the people above this person one question: would you let them be the last signature on this, with nobody checking behind them? The level is the largest unit for which the honest answer is yes. The test's value is its falsifiability: a "yes" here is a commitment the organization then acts on, so inflated answers surface as incidents, and the ladder self-corrects in a way adjective-based rubrics never do.
Test 02 · Slop
Give the person real agent output at their claimed span, seeded with a small number of planted defects of the plausible kind: a subtly wrong boundary condition, a regenerated utility that already exists in your codebase, a correct-looking behavioral change to something load-bearing. Score two things: what they caught, and what they wasted attention on. The second matters as much as the first; custody at scale is triage, and a reviewer who flags everything has not exercised judgment, they have exported it back to you.
Test 03 · Reconstruction
Take a piece of work the person recently shipped through an agent and accepted. With the agent off, have them explain why the solution is correct, where it would break, and rebuild the load-bearing part by hand. This is the direct counter to the hollow-competence failure Allan describes: someone performing above their formation level on the agent's strength, whose judgment evaporates precisely when the agent's does. A person who passes slop tests but fails reconstruction is not ready for the custody tier; they are ready for more formation.
Set each axis to the tier you would honestly defend.
Custody line
Apprentice custody · Junior Independent custody · Mid-level System custody · Senior Process custody · Staff, principal
The level is the lowest of the three axes, not their average.
Custody line
The equity problem and the judgment problem have a single fix, and it is cheap relative to a bad hire: the organization provides the assessment environment, and every candidate gets the same one.
Every AI-involving exercise in your loop runs in an environment you supply: same model, same token budget, same harness, same retrieval over the same prepared codebase, for every candidate at that stage. Candidates never use personal accounts and are never asked which tools they subscribe to as a proxy for ability.
Constant for every candidate
The one variable
the delta the human adds to a fixed agent
The rule's justification is the finding above: when proficiency is bought with money and unpaid time, an interview that rewards it filters on wallet and circumstance (Lazzaro, LeadDev, 2026-08-12)[1]. A cockpit converts the tool from a differentiator between candidates into a constant, so the only variable left is the one you wanted to measure: the delta the human adds to a fixed agent.
The candidate receives a work-in-progress agent session on the prepared codebase: a task, the agent's plan, and its current output seeded with planted defects, exactly as in the slop test. They review, catch what matters, and redirect the agent to finish. This measures custody behavior directly and takes no longer than the whiteboard exercise it replaces.
The candidate writes the plan for a bounded task; the cockpit agent executes it while they watch; the candidate then verifies the result against their own plan and decides what to accept. This measures the two ends of the custodian's job (intent and adjudication) while the middle, the typing, is performed by the same agent for everyone.
Retain one exercise with no agent at all, sized to the tier's formation floor: reading unfamiliar code and reasoning about it, or debugging a small failure by hand. This is the hiring-time reconstruction test. Without it, you will hire people who interview well inside a cockpit and hollow out your formation profile from the outside.
Fluency with your specific toolchain is learnable in weeks on your payroll and your budget. Judgment and formation are not. Move the first out of the loop entirely, and say so in the job description, because the candidates you are currently losing to the wallet filter cannot otherwise know your loop is safe to enter.
The pipeline problem cannot be fixed inside the interview loop, because it is not a selection problem. It is an investment problem: the market now lets you consume seniors without producing them, and every AI-reliant organization that declines to fund production is free-riding on a shrinking pool. The fix is to make formation an explicit budget line with an owner, rather than a hope.
Set an explicit percentage of apprentice-tier working time (we start engagements at 25% and tune from evidence) that is spent on work agents would otherwise absorb: debugging real failures by hand, reading and mapping unfamiliar parts of the codebase, participating in incident response, implementing selected changes without agent execution. The selection is deliberate: reps are chosen for what they teach, not for what most needs shipping. On this budget, agents may act as tutors (explaining, quizzing, reviewing the junior's manual work) but never as executors, because the entire point is the rep.
25%
we start engagements at 25% and tune from evidence
Assign every apprentice a named mentor at system custody or above, with a fixed cadence and a curriculum tied to the reconstruction test. The evidence here is blunt: development left to chance was failing even when everyone shared an office, and manager participation in structured onboarding tracks with 3.5-times-higher satisfaction (Korducki, LeadDev, 2026-08-03)[4]. Designed beats colocated; it also survives remote.
Each apprentice and independent-custody engineer keeps a running evidence file: reps completed, incidents worked, reconstructions passed, slop tests taken. Promotion to the next tier requires the file, alongside the three tests. This closes the loop that makes the budget defensible in planning: the time is not a perk, it is the documented production process for the seniors your five-year plan already assumes exist.
The slop test is only as good as the people administering it. Managers and system-custody engineers keep a protected slice of real building time so their own taste stays current, per the evidence that recognition of AI slop develops only through practice (Rungta, LeadDev, 2026-08-13)[2]. A ladder administered by people who no longer build converges on rubber-stamping within a few cycles.
The signature test replaces adjective rubrics with a commitment the organization acts on, so miscalibrated promotions surface as evidence rather than lingering as politics.
A cockpit is auditable: you can show any candidate, or any regulator, that every person at a stage faced the same tools and the same budget. "We ask about their Claude setup" is not defensible in the same way, and after the reporting this year, it will be asked about.
In a production-priced ladder, an apprentice plus an agent looks like a worse deal than an agent alone, which is exactly the reasoning currently strangling entry-level hiring. In a custody-priced ladder with a formation budget, junior hiring is the visible input to future system custody, with an evidence file to show for it.
People shipping above their formation level on agent strength stop levelling up on output alone, because reconstruction is checked. Some current mid-levels will test lower than their title suggests; the honest response is formation investment, not quiet demotion.
A quarter of apprentice time doing by hand what an agent does in minutes is deliberately purchased inefficiency. The alternative is consuming judgment capital without replacing it, which is cheaper every individual quarter and ruinous over five years. Fund it like capital expenditure, because that is what it is.
Standardized environments must be built, budgeted per candidate, and refreshed as your production toolchain moves. A stale cockpit slowly reintroduces the very gap it exists to close, between what you assess and how you actually work. Assign its maintenance to process custody and refresh it on a calendar.
Signature, slop and reconstruction tests take senior time to run well; dashboards of merged PRs are free. You are trading a cheap wrong measurement for an expensive right one, and the expense lands on your most loaded people. That cost is real, and it is the smaller one.
Engineers whose standing rests on volume will read custody as a demotion in waiting. Be straightforward: the ladder is repricing what the organization actually needs from seniors now, the tests are open, and formation investment is available to anyone whose reconstruction floor lags their title.
01
Deadline pressure converts reps time back into delivery, one sprint at a time, and the evidence files go stale.
Repair
The budget is tracked like any other committed capacity, and a tier promotion without a current formation file is invalid, no exceptions signed below process custody.
02
Candidates are assessed on last year's harness while the team works on this year's.
Repair
The cockpit is versioned against the production toolchain, with a named owner and a refresh cadence, and interviewers flag drift as a defect.
03
Planted defects become predictable, candidates and promotees pattern-match the test instead of exercising judgment.
Repair
Defects are drawn from your own recent incident and review history, rotated each cycle, and retired once discussed openly.
04
"Would you let them sign it" becomes polite fiction because nobody wants to block a promotion.
Repair
Pair every signature answer with actual delegated custody in the following quarter, and audit incidents back to the signatures that preceded them.
05
Apprentices nominally direct agents but accept everything, and the 96% distrust the industry reports does not show up in their behavior.
Repair
Sample apprentice acceptances through spot reconstruction tests, and treat a pattern of hollow acceptances as a formation gap to fund, not a performance failure to punish.
06
Managers stop building, their taste dates, and the tests they administer soften without anyone noticing.
Repair
The builder floor is on the manager's own review, and system-custody engineers co-administer slop tests so no single atrophying palate controls a tier.
Use this gate honestly before you commit to the redesign.
| Fit criterion | Fits this framework | Does not fit |
|---|---|---|
| Agent adoption | Agents are in the team's daily production workflow | Adoption is exploratory or blocked; redesign would price a hypothetical |
| Organization size | Enough engineers to hold a ladder (roughly ten and up) | A handful of founders and early hires; hiring is direct judgment |
| Regulatory posture | AI tooling is permitted in the development environment | Regulation or contract keeps AI away from the code entirely |
| Toolchain ownership | You own your toolchain and interview loop | Clients dictate tools per engagement, so no stable cockpit can exist |
| Time horizon | You intend to still be producing seniors in five years | Project-bounded team with no development mandate |
Two boundaries deserve emphasis. If agents are not yet in daily use, adopting this ladder first is theater; run the adoption, then re-price what it changed.
This quarter
01
Highlight every criterion denominated in production ("delivers features", "ships high volume", "writes high-quality code") and rewrite each as the custody, judgment or formation claim it was standing in for. This is a document edit, not a reorg, and it can be done in a week.
02
Build one prepared codebase, one seeded supervise-and-redirect exercise, and one manual station. Run it alongside the old loop for a hiring cycle and compare what each surfaces before cutting over.
03
Pick the percentage, name the mentors, define the first quarter's reps from your real backlog and incident history, and open an evidence file per apprentice.
04
Run signature, slop and reconstruction informally on the next promotion candidates before any titles change. The gaps you find are your formation curriculum.
The following two quarters
Move the remaining interview stages into the cockpit; make the evidence file a formal promotion input; put the builder floor on manager reviews; hand the ladder, the cockpit and the test bank to a named process-custody owner. By then the framework is self-maintaining: the tests generate the evidence, the evidence prices the levels, and the levels fund the pipeline.
Practitioner observation · Not an audited benchmark
Observations from engagements where we have run custody-based assessment, offered as practitioner observation rather than audited benchmark:
Observation 01
seeded slop-test exercises separating candidates that a production-style loop had scored identically
Observation 02
reconstruction tests revealing that agent-era output levels overstated formation by roughly a tier in a meaningful minority of cases, most often at the mid-level
Observation 03
and interview cockpits removing an entire category of candidate complaints about tool access at a per-cycle cost well below one bad hire
Your numbers will depend on how far agent adoption has already run ahead of your ladder, and on how honestly the signature test is answered in its first cycle.
Method
This guide generalizes IH Tech's assessment and team-formation practice for AI-native delivery into a leveling and hiring framework, and grounds the problem statement in reporting and survey data published by LeadDev and InfoQ in August 2026, each independently verified on 2026-08-16.
Assumptions
Agents are materially present in the organization's daily development workflow; the organization controls its own interview loop and toolchain; senior engineers exist in sufficient number to administer the tests and staff mentorship; leadership can protect a formation budget against quarterly delivery pressure.
Limitations
The 96% distrust figure and the 3.5-times onboarding satisfaction multiple are survey findings reported by the cited publications, not our measurements, and we have not audited their methodologies. Our own observations come from a small number of engagements and are not controlled studies. The specific budget percentages offered are starting points we tune per organization, not research-derived constants. Tool costs and access tiers cited in the sources are moving targets and will date faster than the structural argument. The framework prices individual contributors and their managers; executive leveling and non-engineering ladders are out of scope.
Corrections
If you believe something here is wrong, tell us. Corrections reach us through the contact route on ihtech.ba, and this guide's modified date changes when its content does.
Updated
All sources verified directly on 2026-08-16.
[1]
Your interview questions assume candidates can afford Claude Code Max
https://leaddev.com/ai/your-interview-questions-assume-candidates-can-afford-claude-code-max
Cited in Why the old ladder · The hiring loop
[2]
Engineering managers who build are pulling ahead
https://leaddev.com/management/engineering-managers-who-build-are-pulling-ahead
Cited in Why the old ladder · The formation budget
[3]
How Artificial Intelligence Disrupts Engineering Progression (reporting Alasdair Allan)
https://www.infoq.com/news/2026/08/AI-disrupts-engineering-progress/
Cited in Why the old ladder
[4]
Your junior engineers don't need an office. They need you
https://leaddev.com/management/your-junior-engineers-dont-need-an-office-they-need-you
Cited in Why the old ladder · The formation budget
[5]
InfoQ Culture and Methods Trends Report - 2026
https://www.infoq.com/articles/culture-trends-2026/
Cited in Why the old ladder