← Back to Research Hub
First-Year Assessment
FY Report: Methodology Account
A reflexive account of the design choices that carry the most weight in Study 1. Study 1 is
diagnostic and comparative: it reads the signature of eroded or eroding evaluative capacity from
how developers describe what they attend to, and does not claim to have caused or measured that
erosion within any individual. Each choice below is stated with the alternative it rejected.
For: First Year Report, Research Methods section
Companion to: Study 1 Research Design
Stance: choices, not quantities the framework fixed
1 · The discrimination test reads reachability, not spontaneous loss
The central probe confronts the practitioner with a concrete case (something that passed every
check and still failed at its actual job) and reads whether the accountable criterion can be
retrieved, rather than waiting for the practitioner to volunteer that their judgment has slipped.
Rejected alternative
Listening for an unprompted sense of loss was rejected because the mechanism predicts that loss
will not be felt: the evaluator who would register it has been transformed by the same engagement
that produced it. A probe built to wait for a spontaneous gap would have been built to miss its
object. The cost is a heavy dependence on the specificity of the confronted case, which the
protocol manages by following the participant's own story before supplying one.
2 · The pre-engagement baseline is reconstructed, not sampled
Almost no working veteran has stayed AI-light, so the baseline against which change is read is
reconstructed from heavy-engagement veterans recalling earlier practice, and that recollection is
filtered through the very transformation the study examines.
Rejected alternative
Claiming a clean live baseline from the AI-light cell was rejected as dishonest to the population,
since a present-day light user is not a pre-engagement specimen. Recall is triangulated against
datable artifacts rather than treated as transparent, and the residual limitation is attached
narrowly to the within-person claim about criteria changing over time, not to the cross-sectional
contrast the findings actually rest on.
3 · Sampling is sorted on observables, with the prediction held separate
Participants are placed by when they began shipping accountable code and how much AI touches that
code now. What the framework expects each cell to surface is held in a separate analytical frame
that never sorts a participant.
Rejected alternative
Collapsing the two frames would have produced cleaner cells and would have let the study recruit
its own confirmation, which is why it was rejected. The separation is also what allows a thin cell
to be reported as a finding about the population rather than a recruitment failure.
4 · The 70/30 frontline-to-boundary split is a judgment under uncertainty
The frontline arm carries the judgment and reachability evidence available nowhere else and takes
the larger share. The boundary-activity arm carries the propagation evidence the frontline arm
cannot reach and is sized to observe the handoff without needing parity. Frontline precedes
boundary activity so that what is being propagated is documented before propagation is traced.
Open to revision
The ratio and the order are open to revision against what saturation shows, and are stated as
choices rather than as quantities the framework fixed.