Two semi-structured interviews observe one typification loop at two points in its circulation. Read together they follow a single circulation rather than two separate populations. Neither interview names AI, judgment, proxies, criteria, or erosion.
Run with the frontline practitioner population the screener has identified. A single semi-structured session of about fifty minutes, run as a linear flow. The object throughout is the code the participant ships. Both screening axes (consequence exposure and AI engagement depth) are fixed before the session and never raised as topics, so the conversation does not tip the framework.
The interviewer confirms consent is in place (signed via the information sheet), confirms recording disclosure and agreement, and reads the framing:
Then a short warm-up (current role, current work, how long at this kind of work) re-confirms the screening facts in the participant's own words. The warm-up is not coded and does not ask about AI use, since it is already known from screening and raising it primes the frame. If the participant raises it, the interviewer follows naturally but does not pursue it.
"Walk me through a piece of work you finished recently that you were satisfied with. What made it good?"
Where the account stays thin, the interviewer draws it out: how they knew it was good, what would have made it not good, who else would have judged it and by what. The reading attends to whether the participant reaches for accountable criteria (the work held under real conditions, it solved the actual problem) or for proxies (it passed review, it shipped fast, the metrics were green). This is the unstressed baseline, the evaluative vocabulary before any failure framing stresses it.
(Amendment-pending, rationale to be filed with the REC.) The opener invites a real, specific piece of recent work to walk through in the moment, not a generalized account. Nothing is collected or retained: the artifact is used only as a concrete object to talk through, and the participant may keep it on their own screen or describe it without showing anything. Anchoring to real work raises the cost of a flattering reconstruction. Where a participant prefers not to anchor to a real example, the interviewer proceeds with the general version and notes the account is unanchored, which is itself a reading.
"Tell me about a time something you shipped passed every check you had, looked clean, and still turned out to be wrong."
Where the participant has a story, the interviewer asks how they realized, and what told them if the checks did not, which surfaces the criterion behind the proxies. Where the participant struggles, the interviewer offers a smaller version once ("maybe something where a green light didn't feel quite right"), without leading, and moves on if nothing comes. The difficulty is itself data. Where the "wrong" turns out to be a later check firing, the interviewer asks whether anything would have told them before that check existed.
Where the participant says it never happens, the interviewer does not push: "That's fine, plenty of people don't have one that jumps out, so when everything does pass and looks clean, how do you know it's actually right?" The interviewer notes whether the participant simply could not bring an instance to mind or asserted that nothing gets past their checks. The second is the stronger signal: treating the checks as complete is itself what it looks like when the checks are not anchored to any accountable criterion.
"When you finish something, how do you know it's actually done, and not just that everything passed?"
Where useful, the interviewer asks whether that has ever come apart (everything passing without it feeling done) and, with participants who have done the work long enough to have seen it change, whether how they know has shifted over the years. The reading attends to whether the participant has an evaluative standard the checks only stand in for, or whether "done" simply means everything passed, with nothing behind it.
The interview never asks whether the participant feels they have lost their edge or that their judgment has slipped. Such a question would prompt a performance rather than an account and would surface the very framing the design works to keep out. Whether an account reflects proxy seduction rather than Goodhart optimization, optimism bias, or skill atrophy is a question for analysis, settled against the transcript, not judged in the room.
The interviewer opens the floor for anything about how the participant evaluates work that the conversation did not reach, asks for referrals (which feed the screener), thanks the participant, and stops the recording.
Run with the boundary-activity population the screener has identified: engineering managers (who turn what a team builds into status reported upward), design leads (who set the review standards a team's craft is judged against), DevEx engineers (who build the internal tooling, pipelines, and metrics others are measured by), DevRel (who turn developer signal into usage and engagement metrics), and product managers (who turn the work into roadmap progress and success metrics). The object throughout is the metric the participant owns and passes upward, not the code they ship, since those doing boundary activity may not ship code directly.
Then a short warm-up (current role, what the team is working on, how long in this kind of role) re-confirms the screening facts. As with the frontline interview, the warm-up is not coded and does not raise AI use.
"Walk me through how you report your team's progress upward. What do you track, and what do you pass along?"
Where the two differ, the interviewer asks what makes something worth passing up. The reading attends to the gap between what the participant tracks and what they pass along, since what drops out between the two is the accountable criterion that has no legible channel upward. What the participant counts as worth passing up surfaces their sense of what the organization reads and rewards, without the interview naming legibility.
"Tell me about a time the numbers looked good, the team was hitting its targets, and something was wrong underneath that the numbers weren't showing."
Where the participant has a story, the interviewer asks how they knew despite the green numbers, and what the numbers had stopped capturing, which surfaces the accountable criterion the metric was meant to track. The fallback ("maybe a number that looked fine but you didn't fully trust") and the never-happens handling mirror the frontline probe. Treating the metrics as complete is the stronger signal.
The design logic is the same as the frontline reachability probe: the test is whether the accountable criterion is retrievable when a concrete case is supplied, not whether the participant feels its absence, since the mechanism predicts the absence will not be felt. The claim is about availability in the account, not deployment in practice.
"How do you know a team is actually doing good work, and not just that the metrics are green?"
The reading attends to whether the participant has an evaluative standard the metrics only stand in for, or whether good work simply means the metrics are green. Read against the track-versus-pass gap from Move 1, this is where RQ2 is settled: whether the criterion the metric was meant to track survives in the person who carries the metric upward, or whether only the metric remains.
The interviewer opens the floor for anything the conversation did not reach, asks for referrals, thanks the participant, and stops the recording.