Team Ai
Datasetpublic

Kidomakai/software-engineering-interview-practices-2005-2026

Replication Package: Yesterday's Interviews for Today's Engineers This repository contains the de-identified analytical data and the Python reproduction script for: Vitalii Romaniuk. "Yesterday's Interviews for Today's Engineers: Retrospective Perceptions and a Work-Aligned Hiring Framework (2005–2026)." arXiv:2609.14046, 2026. Paper: https://arxiv.org/abs/2609.14046 Contents data/survey_responses_deidentified.csv contains the 911 retained survey records used… See the full description on the dataset page: https://huggingface.co/datasets/Kidomakai/software-engineering-interview-practices-2005-2026.

sourceHugging Faceotherupdated 21d agoView on Hugging Face
1likes105downloads
Dataset Card

Replication Package: Yesterday's Interviews for Today's Engineers

This repository contains the de-identified analytical data and the Python reproduction script for:

Vitalii Romaniuk. "Yesterday's Interviews for Today's Engineers: Retrospective Perceptions and a Work-Aligned Hiring Framework (2005–2026)." arXiv:2609.14046, 2026.

Paper: <https://arxiv.org/abs/2609.14046>

Contents

  • —data/survey_responses_deidentified.csv contains the 911 retained survey records used in the analysis. It contains coded responses and arbitrary respondent identifiers; it does not contain names, contact details, employers, geography, demographic attributes, or original collection timestamps.
  • —data/interview_question_benchmark_raw.csv contains the 1,001 input rows for the interview-question benchmark. The reproduction script normalizes prompt labels case-insensitively, removes repeated prompts, and excludes prompt groups with conflicting scores. This produces the 945 consistently scored unique prompts analyzed in the paper.
  • —scripts/reproduce_results.py reproduces the reported descriptive statistics, repeated-measures tests, sensitivity checks, and benchmark counts using only the Python standard library.

The original form export, working notes, drafts, and collection timestamps are not part of this public package.

Reproduction

Python 3.9 or newer is recommended. No third-party packages are required.

From the repository root, run:

sh
python3 scripts/reproduce_results.py

Survey data dictionary

The four prefixes represent Era I (P3, 2005–2010), Era II (P4, 2011–2017), Era III (P5, 2018–2022), and Era IV (P6, 2023–2026).

  • —Respondent_ID: arbitrary sequential identifier with no identity mapping in this package.
  • —Role: 1 = software engineer/candidate; 2 = engineering manager.
  • —*_Duration: 1 = under 1 hour; 2 = 1–2 hours; 3 = 3–4 hours; 4 = 5–6 hours; 5 = 7 or more hours. The analysis uses midpoints 0.5, 1.5, 3.5, 5.5, and 7.5 hours.
  • —*_Rounds: 1 = 1–2 rounds; 2 = 3–4 rounds; 3 = 5 or more rounds. The analysis codes these as 2, 4, and 5.
  • —*_Alignment: interview–work disconnect on a 1–5 scale, where 1 means no disconnect/perfect alignment and 5 means extreme disconnect/completely irrelevant.
  • —*_Toll: combined psychological and time burden on a 1–10 scale.
  • —Submitted: retained-record marker; all released rows are True.

The survey was voluntary, anonymous, and uncompensated. Recruitment used public software-engineering communities on Reddit and LinkedIn. The sample is self-selected and is not a probability sample of the software-engineering profession. The article's methodology, appendix, and threats-to-validity sections provide the authoritative description of collection, cleaning, and interpretation.

Question benchmark

The benchmark columns are:

  • —Interview Question: a short prompt or topic label.
  • —Real Work Job Analogue: the closest plausible professional analogue used in the content audit.
  • —Difference: a score from 1 (closely matches routine engineering work) to 5 (primarily academic, specialized, or interview-oriented).

The source pool comprised openly accessible interview-preparation material, including LeetCode, HackerRank, public GitHub repositories, public Reddit discussions, and other public preparation resources. It is not frequency-weighted. Item-level source URLs were not retained, so the exact provenance of individual prompt labels cannot be reconstructed. Public access does not imply public-domain status. See DATA_RIGHTS.md before redistributing the benchmark.

Licensing and citation

The reproduction script is licensed under the MIT License; see LICENSE. Dataset rights and third-party exclusions are described in DATA_RIGHTS.md. Please cite the paper when using this package.

bibtex
@article{romaniuk2026yesterdays,
  title         = {Yesterday's Interviews for Today's Engineers: Retrospective
                   Perceptions and a Work-Aligned Hiring Framework (2005--2026)},
  author        = {Romaniuk, Vitalii},
  year          = {2026},
  eprint        = {2609.14046},
  archivePrefix = {arXiv},
  primaryClass  = {cs.SE}
}