← Back to Research Hub
Study 1 · Fieldwork Instrument

Interview Protocol

Two semi-structured interviews observe one typification loop at two points in its circulation. Read together they follow a single circulation rather than two separate populations. Neither interview names AI, judgment, proxies, criteria, or erosion.

Format: semi-structured, ~50 min, linear flow Moves: 3 per interview Screening axes: fixed before session, never raised Source: Study 1 Research Design
The loop, observed at two points
The boundary-activity interview catches the loop where situated practice aggregates into organizational templates: the moment a participant decides what is worth passing up is the moment activity is classified into the categories the organization recognizes, and what has no legible channel upward is lost rather than carried into the templates. The frontline interview catches the loop where those templates shape future practice: when a practitioner reports how they know work is done, they report the standard the loop has returned to them, and whether an accountable criterion still sits behind it or only the elevated metric remains under a persisting label.

The frontline interview

Run with the frontline practitioner population the screener has identified. A single semi-structured session of about fifty minutes, run as a linear flow. The object throughout is the code the participant ships. Both screening axes (consequence exposure and AI engagement depth) are fixed before the session and never raised as topics, so the conversation does not tip the framework.

Opening

The interviewer confirms consent is in place (signed via the information sheet), confirms recording disclosure and agreement, and reads the framing:

"The interest is in how developers work day-to-day and how they know when something is done, fifty minutes, real examples, free to skip anything, no right answers."

Then a short warm-up (current role, current work, how long at this kind of work) re-confirms the screening facts in the participant's own words. The warm-up is not coded and does not ask about AI use, since it is already known from screening and raising it primes the frame. If the participant raises it, the interviewer follows naturally but does not pursue it.

Move 1 · the satisfied-work opener

"Walk me through a piece of work you finished recently that you were satisfied with. What made it good?"

Where the account stays thin, the interviewer draws it out: how they knew it was good, what would have made it not good, who else would have judged it and by what. The reading attends to whether the participant reaches for accountable criteria (the work held under real conditions, it solved the actual problem) or for proxies (it passed review, it shipped fast, the metrics were green). This is the unstressed baseline, the evaluative vocabulary before any failure framing stresses it.

(Amendment-pending, rationale to be filed with the REC.) The opener invites a real, specific piece of recent work to walk through in the moment, not a generalized account. Nothing is collected or retained: the artifact is used only as a concrete object to talk through, and the participant may keep it on their own screen or describe it without showing anything. Anchoring to real work raises the cost of a flattering reconstruction. Where a participant prefers not to anchor to a real example, the interviewer proceeds with the general version and notes the account is unanchored, which is itself a reading.

Move 2 · the passed-every-check probe primary reachability probe

"Tell me about a time something you shipped passed every check you had, looked clean, and still turned out to be wrong."

Where the participant has a story, the interviewer asks how they realized, and what told them if the checks did not, which surfaces the criterion behind the proxies. Where the participant struggles, the interviewer offers a smaller version once ("maybe something where a green light didn't feel quite right"), without leading, and moves on if nothing comes. The difficulty is itself data. Where the "wrong" turns out to be a later check firing, the interviewer asks whether anything would have told them before that check existed.

Where the participant says it never happens, the interviewer does not push: "That's fine, plenty of people don't have one that jumps out, so when everything does pass and looks clean, how do you know it's actually right?" The interviewer notes whether the participant simply could not bring an instance to mind or asserted that nothing gets past their checks. The second is the stronger signal: treating the checks as complete is itself what it looks like when the checks are not anchored to any accountable criterion.

Why reachability, not spontaneous loss
Judgment stock is enacted, not possessed. A direct question would collect a performance of evaluative talk rather than the capacity itself, the way asking someone to describe their balance on a bicycle collects neither balance nor its absence. The probe reads proxy seduction from whether the accountable criterion becomes reachable when a concrete failing case is put in front of the practitioner, not from whether the practitioner reports something has been lost. Retrospective opacity holds that the evaluator who would register the loss has been transformed by the same engagement that produced it, so a spontaneous gap is exactly what the mechanism predicts will not appear. A probe built to wait for that gap would have been built to miss its object. The cost is that the probe leans on the quality of the confronted case, which is why the move follows the participant's own story before the interviewer supplies one.
Move 3 · the how-do-you-know-it's-done probe cross-check, every time

"When you finish something, how do you know it's actually done, and not just that everything passed?"

Where useful, the interviewer asks whether that has ever come apart (everything passing without it feeling done) and, with participants who have done the work long enough to have seen it change, whether how they know has shifted over the years. The reading attends to whether the participant has an evaluative standard the checks only stand in for, or whether "done" simply means everything passed, with nothing behind it.

What the interview does not ask

The interview never asks whether the participant feels they have lost their edge or that their judgment has slipped. Such a question would prompt a performance rather than an account and would surface the very framing the design works to keep out. Whether an account reflects proxy seduction rather than Goodhart optimization, optimism bias, or skill atrophy is a question for analysis, settled against the transcript, not judged in the room.

Closing

The interviewer opens the floor for anything about how the participant evaluates work that the conversation did not reach, asks for referrals (which feed the screener), thanks the participant, and stops the recording.

The boundary-activity interview

Run with the boundary-activity population the screener has identified: engineering managers (who turn what a team builds into status reported upward), design leads (who set the review standards a team's craft is judged against), DevEx engineers (who build the internal tooling, pipelines, and metrics others are measured by), DevRel (who turn developer signal into usage and engagement metrics), and product managers (who turn the work into roadmap progress and success metrics). The object throughout is the metric the participant owns and passes upward, not the code they ship, since those doing boundary activity may not ship code directly.

Opening

"The interest is in how teams' work gets assessed and reported, and how you know when things are going well, fifty minutes, real examples, free to skip anything, no right answers."

Then a short warm-up (current role, what the team is working on, how long in this kind of role) re-confirms the screening facts. As with the frontline interview, the warm-up is not coded and does not raise AI use.

Move 1 · the report-upward opener

"Walk me through how you report your team's progress upward. What do you track, and what do you pass along?"

Where the two differ, the interviewer asks what makes something worth passing up. The reading attends to the gap between what the participant tracks and what they pass along, since what drops out between the two is the accountable criterion that has no legible channel upward. What the participant counts as worth passing up surfaces their sense of what the organization reads and rewards, without the interview naming legibility.

Move 2 · the good-numbers probe primary reachability probe

"Tell me about a time the numbers looked good, the team was hitting its targets, and something was wrong underneath that the numbers weren't showing."

Where the participant has a story, the interviewer asks how they knew despite the green numbers, and what the numbers had stopped capturing, which surfaces the accountable criterion the metric was meant to track. The fallback ("maybe a number that looked fine but you didn't fully trust") and the never-happens handling mirror the frontline probe. Treating the metrics as complete is the stronger signal.

The design logic is the same as the frontline reachability probe: the test is whether the accountable criterion is retrievable when a concrete case is supplied, not whether the participant feels its absence, since the mechanism predicts the absence will not be felt. The claim is about availability in the account, not deployment in practice.

Move 3 · the how-do-you-know-good-work probe cross-check, every time

"How do you know a team is actually doing good work, and not just that the metrics are green?"

The reading attends to whether the participant has an evaluative standard the metrics only stand in for, or whether good work simply means the metrics are green. Read against the track-versus-pass gap from Move 1, this is where RQ2 is settled: whether the criterion the metric was meant to track survives in the person who carries the metric upward, or whether only the metric remains.

Closing

The interviewer opens the floor for anything the conversation did not reach, asks for referrals, thanks the participant, and stops the recording.

Derived from PSF Fieldwork: Study 1, Research Design and Interview Protocol. Amendment-pending items are flagged for filing with the REC. See the Coding Rubric for how these accounts are read in analysis, and the Study 1 Research Design for the sampling frame.