← Back to Research Hub
Study 1 · Master Document

Study 1: Research Design

The governing design document for Phase 1 fieldwork. Study 1 investigates how AI engagement reshapes the evidence professional software developers attend to when they evaluate code, their own and their team's, and looks for the observable signature of that displacement in how they describe what they attend to. It is diagnostic and comparative: it reads a signature already present, and does not claim to have caused or measured erosion within any individual.

Domain: professional software development Sample: ~50 interviews, 70/30 frontline to boundary Feeds: framework refinement → Paper 2 REC approval: 15 Apr 2026

The two screening axes

The study samples developers along two axes set at screening.

The cell structure: two frames, never collapsed

The two axes give four cells, read through two distinct frames that must not be collapsed. Keeping them apart is what stops the predicted reading from contaminating the sampling.

Screening frame (observable only)
What we sort on. Claims only when a developer started shipping accountable code and how much AI touches that code now. It says nothing about whether anyone holds accountable judgment. That is the study's question, not its starting point.
Coding frame (predicted reading, to be tested)
What PSF expects the interview to surface in each cell. Not shown to participants, not used to sort. Each cell is the hypothesis for that population, and the interview confirms or denies it. If the coding frame sorted participants, the study would recruit the people who confirm it, and the confirmation would be an artifact of selection rather than a finding.
CellCoding-frame role (predicted, tested)
Veteran, AI-heavyThe core, the active PSF site.
Newcomer, AI-heavyThe premature-arrest contrast (non-formation rather than erosion).
Veteran, AI-lightA fork: Condition 2 (the real baseline, accountable-criteria judgment under light AI) is hunted deliberately; Condition 1 (proxy-fluent under light AI) is allowed to arrive as a byproduct.
Newcomer, AI-lightExpected empty / off-mechanism, documented as structurally thin.

The cell populations are deliberately uneven, and the unevenness is theory-consistent rather than a design flaw. The separation is what lets a shortfall in a cell, especially Veteran AI-light, be reported as a finding about the population rather than a recruitment failure.

Research questions

RQ1 (judgment)

How does cumulative consequence exposure shape which dimensions practitioners have formed judgment for (whether tacit or codified into the tools and checks they have built), and on which of the dimensions the engagement now makes more legible can that judgment not be drawn? Constructs tested: judgment stock, detection, braking. RQ1 listens for whether consequence exposure lets a practitioner discriminate proxy metrics (speed, volume, apparent certainty) from accountable criteria (whether the code held under conditions the tests did not cover), and for the dimensions the engagement makes more legible, where even experienced practitioners cannot draw consequence-based judgment (the METR workflow-speed pattern).

RQ2 (propagation)

How do the metrics elevated through AI engagement move from individual practice into team, organizational, and field-level evaluation? Constructs tested: proxy elevation, the typification loop (Weber and Glynn 2006), adaptive instability (Gioia et al. 2000), comparative legibility. The loop operates on what is most legible, so engagement-elevated metrics travel through it and reinforce the templates that shape future practice, while craft judgments and tacit standards lose their feedback pathway. Through the loop the elevated metrics come to define what counts as good review, good analysis, or competent practice even while the labels persist.

Scope: which research directions Study 1 targets

Paper 1 names three research directions. Study 1 targets the first two as question arms and folds the third in as a coding lens, for reasons of method-fit. Study 1's findings feed the framework refinement that produces Paper 2.

Sample size, split, and allocation

Target approximately 50 semi-structured interviews, 70/30 frontline to boundary activity, frontline interviewed first so criteria shifts are documented before propagation is traced. The ratio is a judgment call, not a derivation: the frontline arm carries the judgment and reachability evidence available nowhere else and takes the larger share and the saturation room, while the boundary arm is sized to observe the handoff where a criterion survives or drops out, without needing parity. Both are open to revision against what saturation shows. Full cell targets are tabulated in the Empirical Phase Checklist.

Interview protocol and coding

Two interviews follow one typification loop at two points. The Interview Protocol gives both in full (three scripted moves each, neither naming AI or judgment). The Coding Rubric operationalizes the falsification spine at the transcript level, built around the Position 1/2/3 discrimination test that separates proxy seduction from skill atrophy.

Ethics and approvals

Cambridge Department of Engineering Research Ethics Committee approval granted 15 April 2026 (Chair: Dr. Robert Phaal). Consent, data handling, and anonymization follow the approved submission. The artifact-anchoring provisions in the protocol are marked amendment-pending and are not exercised until the REC amendment is approved. No artifact is collected, transmitted, or retained.

This is the master document for Study 1. The Interview Protocol, Coding Rubric, and Empirical Phase Checklist all derive from it. See the FYR Methodology for the reflexive account of the design choices.