← Back to Research Hub
First-Year Assessment

FY Report: Methodology Account

A reflexive account of the design choices that carry the most weight in Study 1. Study 1 is diagnostic and comparative: it reads the signature of eroded or eroding evaluative capacity from how developers describe what they attend to, and does not claim to have caused or measured that erosion within any individual. Each choice below is stated with the alternative it rejected.

For: First Year Report, Research Methods section Companion to: Study 1 Research Design Stance: choices, not quantities the framework fixed

1 · The discrimination test reads reachability, not spontaneous loss

The central probe confronts the practitioner with a concrete case (something that passed every check and still failed at its actual job) and reads whether the accountable criterion can be retrieved, rather than waiting for the practitioner to volunteer that their judgment has slipped.

Rejected alternative
Listening for an unprompted sense of loss was rejected because the mechanism predicts that loss will not be felt: the evaluator who would register it has been transformed by the same engagement that produced it. A probe built to wait for a spontaneous gap would have been built to miss its object. The cost is a heavy dependence on the specificity of the confronted case, which the protocol manages by following the participant's own story before supplying one.

2 · The pre-engagement baseline is reconstructed, not sampled

Almost no working veteran has stayed AI-light, so the baseline against which change is read is reconstructed from heavy-engagement veterans recalling earlier practice, and that recollection is filtered through the very transformation the study examines.

Rejected alternative
Claiming a clean live baseline from the AI-light cell was rejected as dishonest to the population, since a present-day light user is not a pre-engagement specimen. Recall is triangulated against datable artifacts rather than treated as transparent, and the residual limitation is attached narrowly to the within-person claim about criteria changing over time, not to the cross-sectional contrast the findings actually rest on.

3 · Sampling is sorted on observables, with the prediction held separate

Participants are placed by when they began shipping accountable code and how much AI touches that code now. What the framework expects each cell to surface is held in a separate analytical frame that never sorts a participant.

Rejected alternative
Collapsing the two frames would have produced cleaner cells and would have let the study recruit its own confirmation, which is why it was rejected. The separation is also what allows a thin cell to be reported as a finding about the population rather than a recruitment failure.

4 · The 70/30 frontline-to-boundary split is a judgment under uncertainty

The frontline arm carries the judgment and reachability evidence available nowhere else and takes the larger share. The boundary-activity arm carries the propagation evidence the frontline arm cannot reach and is sized to observe the handoff without needing parity. Frontline precedes boundary activity so that what is being propagated is documented before propagation is traced.

Open to revision
The ratio and the order are open to revision against what saturation shows, and are stated as choices rather than as quantities the framework fixed.
Derived from FYR Methodology: Reflexive Account of Design Choices. See the Study 1 Research Design for the design these choices defend, and the FY Report outline for where this sits in the report.