Security Hiring Research: What Job Postings Say About Security Work
Written entirely by agents in the Research Factory, within guardrails and controls set by our team. Figures as of 1 October 2026.
- 01 Read about 200,000 postings a day from about 2,200 job boards done
- 02 Keep about 17,000 security and engineering postings since July done
- 03 Score labels against one to three frontier model families done
- 04 Recompute every figure from source before publishing done
- 05 Publish only when every control passes done
A job posting is the most detailed public statement a company makes about the security work it needs done. It names the tools, the level of hands-on work, the certifications, the reporting line and, more and more, what the hire is expected to do with AI. Read one posting and you learn about one team. Read 2,194 job boards every day and changes in the work show up in the text.
This series does that. Each piece takes one question about security hiring and answers it from the data, with the numbers, a confidence rating, a representativeness rating and the tests the answer survived.
What we collect
The daily read
Every day, agents read 2,194 job boards across 16 applicant tracking systems, including Greenhouse, Ashby, Lever and Workday. On 1 October 2026 that pass read 220,747 open postings. Collection began on 31 July 2026.
The company list follows employers that hire security and AI engineering talent. It grew from public company lists, a cohort of AI-first companies and daily sweeps for new boards.
The security and engineering set
From that daily read, the agents keep the full text of every security posting and a sample of engineering postings for comparison. When a posting changes, the earlier version is kept too. Since 31 July the set has collected 15,163 security postings and 2,886 engineering postings from 2,537 companies.
Public baselines for comparison
Two months of collection cannot show a trend on its own. For questions about change over time, each piece pairs the current data with a public dataset from an earlier year, run through exactly the same rules. The first baseline is the LinkedIn Job Postings dataset published on Hugging Face (datastax/linkedin_job_listings), covering US postings listed between 5 and 19 April 2024.
One population for every comparison
Every piece that compares 2024 with 2026 reads one population, built the same way for both years. A posting is in when its title passes the series' security title rules and it carries a description. A local model then reads the title and the opening of each posting and rules whether the job is security work at all, which removes postings such as financial auditors, guards, lawyers and alarm installers. The same model, prompt and settings read both years. Last, each company counts each title once. Physical security, fraud and trust and safety roles are out of scope in both years, and so are job aggregators, which repost other employers' jobs.
The 2026 side holds postings open on or after 23 September 2026. Until the evening of 21 September the job-board collector skipped titles containing analyst, auditor, consultant, CISO, director, head of, specialist and "manager, security", so postings that closed before then lean toward engineer titles. The window leaves 4,587 postings at 831 US-headquartered companies, against 1,038 US postings in spring 2024.
How a question gets answered
Most measures are written rules applied identically to every posting: the level implied by a title, the specialty, whether the description asks for a programming language, automation or a certification. Rules are checked by having a frontier model, one of the most capable AI models available, read a sample without seeing the rule's answers. Title level agreed with Claude Opus at a Cohen's kappa of 0.975, specialty at 0.76. Kappa runs from 0, no better than chance, to 1, perfect agreement.
Some measures need judgement, such as whether the hire manages people or personally does the technical work. For those, a local model reads every posting, and its answers are checked against the answers of one to three frontier models from different AI companies. In Q1 to Q4, three AI models each read the postings, GPT, Grok and the local model, and the pieces use the answer at least two agree on. In Q5, GPT and Grok check the local model's answers on US postings. In Q6 the three models are Claude Opus, GPT and Grok. Each piece reports the agreement it relies on.
How every answer is rated
A posting is what a company says it wants. It is not a headcount, and it is not proof that anyone was hired. Every answer in this series is about stated expectations, and each one carries two separate ratings.
Each figure comes with a range of likely values. We get it by redrawing the companies in our data at random many times and recomputing the figure each time, so the range shows how much the answer could move if a different set of companies had been collected. The range we show holds 95% of those recomputed values. If that range includes zero, we can't tell the change from no change.
Confidence is how likely the number is right for the postings we read.
Three rules sit under the scale, and the pieces use a few terms for them. A change is a clear rise or a clear fall when its whole range of likely values sits on one side of zero, so the range excludes no change. That makes it a real rise or fall, too large to be explained by which postings happened to be collected. When the range includes zero, the change is too small to tell apart from no change, and the pieces call it not measurable or no clear change.
A finding that there is no difference is rated by how narrow its range is: High when the whole range sits within 5 percentage points of zero, Medium within 10, and otherwise the difference is not measurable. A comparison with fewer than 30 postings or 10 companies on either side is rated Medium at most.
Each range is computed several times, each time from a different random starting point. A range counts as borderline when it excludes zero but ends within 1 percentage point of it, or within 5% of the comparison group's value for a ratio. It also counts as borderline when it excludes zero in some of those computations but not in all. A borderline range that excludes zero is called a likely but borderline rise or fall. The change is probably real, but only just. A borderline range that reaches zero or sits just past it is called no clear change, but only just. It includes no change, by a small margin. The pieces call both kinds borderline.
One more rule covers time. A High rating for a difference needs the difference to have been clear in the previous run on the same rules as well. Without an earlier run on the same definitions, a finding is rated Medium at most, whether it is a difference or a finding of no difference, and the 30 September run was the first under the rules above. This release, on 1 October, is the second. A rating can be overridden downward, with a stated reason, but never upward, and an override that sets High stops the run.
Representativeness is how far the answer generalises to security hiring as a whole. The company list leans toward technology companies, US headquarters and employers that post on modern applicant tracking systems. It under-covers government hiring, staffing and managed-security firms, and companies that post only on LinkedIn or Indeed. An answer can be high on confidence and low on representativeness, and several are.
Representativeness has three levels. High means the answer holds across the companies the list covers well and the ones it under-covers. Medium means the answer rests on security postings broadly, and the list's lean could change its size more than its direction. Low means the answer rests on one level, specialty, industry or segment of postings, or on a group the list under-covers, or part of it comes from which companies happen to be in each year's data.
How every answer is stress-tested
Every answer goes through the same checks before it is published. Each piece reports what the checks found, including the ones that changed the answer.
- Repeat checks. Every comparison is repeated on the same companies in both years where possible, outside technology companies, on postings of similar length, and with each company counted once. Outside technology companies means 2026 employers in healthcare, retail and e-commerce, manufacturing, defense and government, staffing and consulting, media and entertainment, consumer goods and financial services.
- Measure checks. In every period compared, keyword matches are read by hand to see how many are genuine. Title rules and model labels are compared with frontier models' answers, without either seeing the other's.
- Scope checks. A local model, run on our own hardware, reads every posting in both years and rules whether it is a security role. Before use it was scored against Claude Opus on every 2024 posting, and a person read the postings where the two disagreed. That score was measured on the same 2024 postings used to tune the prompt, so it may read higher than it would on new postings. The prompt was amended once after the first pass over the 1,444 2024 postings, the same prompt then ran on both years, and Claude Opus did not read the 2026 postings.
- Ranges of likely values. Every change carries a range of likely values, built by redrawing whole companies rather than single postings. So one large employer with dozens of similar postings cannot carry a result on its own.
- A fresh run. Every figure is recomputed from source data before a piece is published or updated, and the date at the top of the piece says when. Each release pins its figures to one day's collection. Version 1.0.0 rests on the 1 October 2026 collection. Figures change only in a new release, and the changelog says what moved and why. A finding that no longer holds comes down. On 24 September a refresh showed that one of the four findings on our homepage no longer cleared zero, and it was replaced the same evening.
Each piece closes with the tests that would strengthen its answer and why they cannot run yet.
The postings quoted in this series are public. The company list and the labels behind the numbers are not published.
Some of these checks were added when a problem turned up, and each now applies to every piece. The tables below list every check the series has passed, grouped by what it looks at: the data, the measures, the ratings, reproducibility, the text and review. Each row gives what the check found, what changed and the effect on the findings.
Q1 to Q10 are the ten questions listed at the end of this page. Three kinds of review appear in the rows:
- A cold-read review is a read of the finished pieces by a reviewer who had not worked on them.
- An independent review is a blind review, by Claude Opus or OpenAI's Codex, that recomputed the figures from the data.
- A cross-family read is an AI model from a different company reading a piece for claims that go past the data.
"At that stage" marks figures from an intermediate run; each piece carries the current figures. "Standing rule from before the series" marks a rule adopted before the first piece was drafted.
Data and population
| Check | What it found | What changed | Effect on the findings |
|---|---|---|---|
| Posting counts over time | Posting counts grew from 2,462 to 6,432 in 38 days, almost all of it from adding job boards | No growth claim rests on posting counts. Pieces compare shares against a baseline. | Standing rule from before the series |
| Posting dates | Posting dates come from the applicant tracking system and favour postings still open | Posting age is never read as a trend | Standing rule from before the series |
| Samples read from every specialty and level, both years | 2024 carried guards, trades advertised "with security clearance" and psychiatric nurses. 2026 carried chip-design "SoC" titles. | Removed by title rule, in both years | In this release the title rules remove 414 2024 postings and 127 2026 postings at US-headquartered companies before the scope check: guard, physical-security and similar titles (405 and 108), chip-design and clinical titles (8 and 19), and one 2024 title that read as security only for its clearance wording |
| Duplicate postings, both years | 2024 carried repeat postings, the same title at the same company, that the 2026 side had already collapsed | One posting per company and title, in both years | In this release 107 repeats leave the 2024 comparison set and none leave 2026. Kept, they would put the 2024 manager share at 8.8% rather than 7.7%, engineer titles at 34.5% rather than 34.0%, and language asks at 22.5% rather than 22.0%. |
| Scope check of the 2024 set: Claude Opus read all 1,608 2024 postings with a security title | 27% were not security roles: financial auditors, physical security, alarm installers, a disability attorney. Among leadership titles it was 41%. | Removed from the 2024 set. Later replaced by one scope check run on both years (see "Same rules in both years: the 2024 side"). | 2024 detection and response automation rose from 28% to 41%, and the certification drops in detection and response and in GRC became visible |
| Scope check of Q6's 2024 leadership set: Opus read all 245 leadership-titled 2024 postings | 100 were not security roles | Q6's 2024 comparison uses in-scope people managers only | 72 people managers in this release, after the series-wide scope check and the aggregator exclusion |
| Same rules in both years: the 2026 side, compared line by line with 2024 | 2026 postings counted only if the job-board collector had tagged the title with a security function. The tag missed titles such as "Information Security Manager", spelled-out CISO titles and "…Engineering Manager". 2,610 stored postings had no tag, and 2,071 of them were in scope; about 27 of 30 sampled were real security jobs. A further 2,074 untagged postings first seen 31 July to 21 September were dropped: 88% engineer titles, 81% closed. | 2026 is no longer filtered by the tag. Both years pass the same title rules. | The manager-share fall in an early-access draft of Q2 (8.8% to 3.5%) came mainly from this side. At that stage, on one population, the manager share was 7.6% in 2024 and 6.4% in 2026, a change of −1.2 points (−2.5 to +0.4), with no measurable fall. Q2's title changed. Engineer titles (Q9) went from 34.5% to 61.7%, and programming-language asks (Q1) from 22.1% to 42.7%. |
| Same rules in both years: the 2024 side | An Opus scope check removed 438 of the 2024 postings (26%), and nothing like it ran on 2026 | One scope check now runs on both years, by a local model with the same prompt and settings. The 2024-only check was removed. | The 2024-only check had lowered the 2024 manager share, so it worked against the fall, not toward it. At that stage: 9.8% with no scope check, 7.8% with the Opus check, 7.6% with the local check. In this release: 10.0% without a scope check, 7.7% with it. |
| A first fix, reviewed before merge | The first fix narrowed 2024 to the 2026 tag. The independent review found the work correct, with every headline recomputing, but the approach wrong: the tag drops real security titles, and narrowing 2024 copies those misses into both years. | Not merged. Replaced by widening 2026 and applying one rule set and one scope check to both years, then reviewed again before merge. | Widening rather than narrowing, per the review: manager share 7.8% to 7.4%, engineer titles 33.7% to 57.2%, language asks 21.6% to 39.6%. Every direction holds. |
| Q9 held to the same title rule and de-duplication in both years | Q9 had not applied the same title rule and de-duplication to both years | Both years pass the same rule and de-duplication | Engineer titles 34% to 63% became 39% to 63%. On the one population that followed, 34% to 62%, and 34% to 58% in this release. |
| Every piece rerun on the one population | Findings that rested on populations built differently in each year | Q1 to Q10 recomputed, and their text and ratings updated | At that stage: Q1 application security titles, certifications and cloud titles are no longer measurable. Q2: the manager fall does not survive. Q3: the fall in the entry-level share is a clear fall, the CISSP fall is High, and the operations lead over privacy is no longer measurable. Q4: AI expectations are rated High, executives' lead over managers on AI governance is clear, and the manager tool-use gap is borderline. Q5: industry table re-ranked, overall AI gap by size rated Medium. Q6: the executive deciding gap is no longer measurable. Q8: posted pay +27.2%, with a level-mix check; leadership and managers a likely but borderline rise; individual contributors High. Q9: engineer titles 34% to 62%, and manager titles rise. Q10: findings and ratings unchanged. |
| Collection history, upstream of the series' own rules (independent review) | Before the evening of 21 September the 2026 collector did not collect titles containing analyst, auditor, consultant, CISO, "manager, security", director, "head of" or specialist. 904 postings from before that change were still in the population, 81% with engineer titles. Postings collected before the change were 82.5% engineer titles and 0.5% executive; postings collected after it, 35.0% and 9.8%. The series' own rules cannot see this. | 2026 is limited to postings open on or after 23 September 2026. Both views are shown (see "Both views side by side"). | On the 30 September data, when the review found this, the window left 4,524 2026 postings at 830 US-headquartered companies, from 5,339 at 861. On it at that stage, with aggregators out: engineer titles 34% to 58% (+20 to +29 points), language asks 22% to 43%, the manager share 7.7% to 6.7% (−3.1 to +1.0) and the executive share 3.8% to 5.3% (−0.4 to +3.7). Every direction holds. The 2026 engineer-title share is 4 points lower than on the full set. The pieces carry the 1 October figures: 4,587 postings, the manager share 6.8% (−3.0 to +1.0) and the executive share −0.4 to +3.8. |
| Company concentration (independent review) | Companies with 20 or more postings made up 39.5% of 2026 and 2.0% of 2024. The largest 2026 "company" was Jobgether, a job aggregator, with 229 templated postings and a language-ask share of 3.1%. | Job aggregators, which repost other employers' jobs, are left out in both years: Jobgether, Dice, ClearanceJobs, Jobs via eFinancialCareers, Talentify.io and The Job Network | 33 postings leave 2024 (21 from Dice, 8 from ClearanceJobs) and 229 leave 2026, all Jobgether. On the full 2026 set, language asks move from 43% to 44%. Q5's startup degree gap moves from clear to likely but borderline, and its headline from "about half as often" to "about 60% as often". Q1's architecture decision role now shows no clear change, and Q9's executive engineer-title rise is now clear. Q6's fixed panel loses two Jobgether postings, and one that closed before 23 September; with both rules its headline moves from 76% against 44% to 78% against 43%. Q7's own snapshot loses 568 security postings and one job board, and AI engineering openings per security opening at other companies move from 0.68 to 0.64 on the 30 September data (0.65 on 1 October, as Q7 reports). |
| Both views side by side: every 2026 posting against postings open on or after 23 September, both without aggregators | Whether the window changes any headline | Both views computed in full. The series publishes the window. | On the 30 September data: Full set, then window. Q1 language asks 22% to 44%, and 22% to 43%. Q2 manager share 7.7% to 6.4%, and 7.7% to 6.7%. Q3 share of postings open to zero to two years' experience 18% to 13% in both. Q4 postings with any AI expectation 3% to 37%, and 3% to 36%. Q5 startup degree share 9.8% against 16.0% at enterprises, and 9.9% against 16.3%. Q6 hands-on work in manager postings that expect AI 78% against 44% without, and 78% against 43%. Q7 reads its own snapshot and is the same in both. Q8 posted pay +27.2%, and +28.3%. Q9 engineer titles 34% to 62%, and 34% to 58%. Q10 scam warnings 0.7% to 11.8%, and 0.7% to 11.4%. Every direction holds and every size is within 4 points, Q9 the largest. Four secondary findings move: Q2 people management (borderline to no clear change), the Q4 manager tool-use gap (borderline to no clear change), the Q5 AI gap by size (no clear change to borderline) and Q1 vulnerability-management engineer titles (clear to borderline). |
Measures and labels
| Check | What it found | What changed | Effect on the findings |
|---|---|---|---|
| Labels scored against another model family | A model label agreed with itself but scored a kappa of 0.18 against Claude | Labels are scored against a different model family, never against themselves | Standing rule from before the series |
| Keyword hits read by hand | Keyword matches on company boilerplate, such as vendor marketing copy | Keyword precision is read by hand before a measure is used | Standing rule from before the series |
| The agent, LLM and MCP keyword count, read by hand | Counted over the whole posting, 13 matches in 20 were genuine, most of the rest company boilerplate | Counted only in the responsibilities and requirements sections (19 in 20), marked in Q7 as added after the run | The figure on our homepage and in our company post went from 49% against 27% to 26% against 16%. Q7 shows 26% against 17% in this release. In Q7 the scale gave High, and the rating is set to Medium: the measure was added after the run and combines securing AI with using AI for security work. |
| Title rules against Claude Opus, 192 titles labelled blind across both years | Level agreement was a kappa of 0.94. The misses showed that "Directory Services" matched "director", so engineers were counted as executives. | Rule fixed | Level agreement rose to 0.975 |
| The dataset's own seniority tag, audited | Titles naming a domain called "management" counted as managers. 184 of 495 "manager" postings were engineers ("Security Engineer, Vulnerability Management"). | Fixed and applied to every stored posting. The pieces use their own title rule. | Q5 management share by company size went from 14.0 / 8.4 / 5.9% to 9.8 / 6.4 / 4.6%, and the gap between sizes held. The Q4 manager rate for securing AI went from 4.6% to 3.9%. The Q6 panel moved by at most a point. |
| Q7 domain labels against a keyword rule, then against a blind Claude Opus read of 100 postings | Model against rule: a kappa of 0.52, below the 0.6 bar | An Opus check was added and marked as added after the run. Domain shares come from the model. | Model against Opus 0.75, rule against Opus 0.51 |
| Q7 reporting lines against Opus, 65 postings | In 11 of the 20 cases where only the model saw a reporting line, the line was not there | Only lines that both the model and a keyword rule find are counted | Q7 reports no measurable difference in reporting lines (−9.8 to +17.5 points), on lines both methods find |
| The local labeller on both years, scored against frontier labels | It agreed poorly on hands-on work. Tuning raised its people-management agreement from 0.69 to 0.82 on a held-out set, but not its hands-on agreement. After relabelling both years on 28 September, it agrees with GPT at 0.56 on hands-on work. | The hands-on and architecture findings rest on GPT and Grok labels | The Q1 answer moved from "more hands-on" to "the kind of hands-on work changed" |
| Choice of scope-check model: two local models scored against Claude Opus on a weighted sample of 200 2024 postings, with finding out-of-scope postings as the task | The faster model missed out-of-scope postings. It caught 79.6% of them (recall 0.796), and 93.5% of the postings it flagged were truly out of scope (precision 0.935, kappa 0.812). It gave no answer on 7 postings | Rejected for low recall. The chosen model scored precision 0.920, recall 0.940 and kappa 0.904 on the sample. | Used for every scope ruling in both years |
| Scope-check prompt, disagreements read by hand | The model ruled IT SOX auditors (for example "IT SOX Sr. Auditor, Internal Audit") to be financial audit | The prompt now says that IT audit (IT general controls, IT SOX controls, technology audit) and technology or IT risk count as security work. All of 2024 was re-read, and both years ran on the amended prompt. The prompt was tuned on 2024 only, which the independent review recorded as a mild bias. | Precision against Opus on 2024 rose from 0.88 to 0.936 |
| Scope-check calibration on all 1,444 2024 postings, recomputed by the independent review | Precision 0.936, recall 0.902, kappa 0.889. It rightly flagged 349 postings as out of scope, wrongly flagged 24 and missed 38. 3.5% of kept 2024 postings were out of scope by Opus, and 6.4% of dropped ones were in. A hand read of 40 disagreements, 20 each way, sided with the local model in about 29. The calibration is in-sample: the prompt was amended once after the first pass over 2024, and there is no Opus reference for 2026. | Adopted as the scope check for both years | On its own reading of 2026, the review found 40 of 40 kept postings were real security roles and about 20 of 22 drops correct. The model drops 25.8% of 2024 and 10.0% of 2026; the difference fits the 2026 collector, which already keeps only security postings. |
| Label versions across years, recomputed by the independent review | The same versions label both years: the local model on the 28 September prompt, GPT and Grok at fixed settings. On 30 September, 2 of 4,524 2026 postings lacked a local label. On 1 October, 1 of 4,587 did. 6 postings with two labels on file take either one; each pair agrees. | None needed | None |
| Q8 pay parser against an independent re-parse (independent review) | A rough independent parse gave +21.1% (+12 to +30), or +17.5% at the 2024 level mix, against the piece's +27.2% and +23.2%. An earlier check against a blind Claude Opus read of 300 postings had passed: 99.0% precision, 97.5% recall, 194 of 198 ranges within 2%. | The gap was traced posting by posting, and the parser stays as it is. The rough parse found pay in about 1,500 2026 postings, the parser in 2,272. About 800 of the ranges it missed sit in job-board pay-range markup, where the dash between the two amounts is a character code, and they run higher (median $192k). It also counted about 30 ranges the parser leaves out on purpose: Canadian dollars, on-target earnings, and ranges whose top is more than three times the bottom. The parser misses 4 US ranges the rough parse reads. | The gap adds up: +21.1% becomes +26.0% with the missed ranges, +26.5% taking the highest tier where a posting lists several, +28.7% with pay the job board returned for 248 postings, and +27.2% with the 2024 ranges read from the posting text. No figure changed. Under the 23 September window the piece reports +28.3% (+19 to +37), and +23.6% at the 2024 level mix. |
| Q8 pay coverage | The collector did not keep the Ashby and Lever pay field, so pay for postings that closed before 30 September could not be read | Collection keeps the field, and open postings were filled in | Without that field the rise was +19.2% (+13 to +30) at the time. In this release it is +25.1% (+17 to +34) without the field, against +27.4% with it. |
Statistics and ratings
| Check | What it found | What changed | Effect on the findings |
|---|---|---|---|
| Distinctiveness against a control | A "this is unusual" finding, close to publication, failed when the same measure was run on the comparison group | Every claim that something is unusual is first run on the comparison group | The finding was not published |
| The confidence scale, applied by code | Ratings had been set by judgement. Some Highs rested on borderline ranges, or on controls whose range included no change. | Code applies the scale to every rating. A control whose range includes zero drops High to Medium, and so does a kappa under 0.7 or a spread of more than 5 points between methods. Small bases are capped at Medium. No-difference findings are rated by the width of the likely range. | At that stage: the Q3 entry-level share moved from High to Medium (borderline), Q4 AI expectations from High to Medium, the Q1 amount of hands-on work from High to Low (not measurable), and Q2 executives from Low to Medium |
| How the rules read the end of a range (cold-read review) | The rules compared the unrounded end of the range, so an end shown as "+1" passed as a clear change. Q1's architecture decision role showed "+1 to +26" and was rated High. Four other ranges in Q1, Q3 and Q4 had the same gap. | The rules read each end of a range as the piece prints it. A range that excludes zero with a printed end inside the borderline margin is described as likely but borderline, and the wording and ratings follow. | The five ranges are described as likely but borderline |
| Stability across runs (cold-read review) | Q7's security-duty finding failed on 23 September, passed on 25 September and was rated High. It had replaced a homepage card that failed the same way the day before. | High needs the main range to show a clear change in the previous recorded run too | Q7's security-duty finding passed the 25, 28 and 30 September runs. It was held at Medium on 30 September by the first-run rule, and is High again in this release after passing on 1 October. |
| Run history for the stability rule (independent review of the first fix) | Earlier runs were recorded by data version only, so a run under different definitions still counted | Runs are recorded by data version and definitions | Runs on other definitions no longer count |
| The first-run rule (independent review) | With no earlier run on the same definitions, the stability rule passed automatically. No recorded run carried its definitions, so every piece passed. | With no earlier run on the same definitions, a rating is capped at Medium, for a difference and for a finding of no difference | The 30 September run was the first on its definitions, so every High moved to Medium: 27 ratings across all ten pieces (Q1 3, Q2 2, Q3 5, Q4 3, Q5 3, Q6 1, Q7 3, Q8 1, Q9 5, Q10 1), including Q2's finding that the generic leadership title held. The 1 October run, the second under the same rules, restored all 27. |
| Overrides of the scale | An override that set a rating higher than the scale gives was reported but did not stop the run | An override can lower a rating, with a stated reason, and never raise it. One that sets High fails the run. | No rating in this release is raised by an override |
| Q10 no-change finding against every test (cross-family read) | Background checks were rated High as no change, on a headline likely range within 5 points of zero. At that stage the each-employer-once test showed a clear rise (+0.2 to +5.3), and the same-company test ran to −7.5. | Rating overridden | Background checks went from High to Medium. In this release they stay at Medium: counted once per employer the rate shows a likely but borderline rise (+0.3 to +5.4), and the test on 2024 postings applied through a company job board runs to −6.3. |
| Q5 rating of a level, not a direction | The scale tests direction against zero, which a level of 61% passes trivially. Defense's clearance share rests on 37 defense and government companies, with a likely range of 38% to 85%, and the middle of the ranking has no adjacent-rank test. | Rating set to Medium | Held at Medium, not High |
| Rating levels (cold-read review) | Pieces used in-between levels with no definition, such as Medium-low and Low-medium | Both scales use only High, Medium, Low and Not answerable | 11 representativeness ratings in Q1 and Q3 to Q7, each Medium-low or Low-medium, became Low. No confidence rating had used an in-between level. |
| Fresh run before publishing | On 24 September a refresh showed that one of the four homepage findings no longer showed a clear change | Replaced the same evening | The finding came down |
Reproducibility
| Check | What it found | What changed | Effect on the findings |
|---|---|---|---|
| Saved queries | Q3's table query had never been saved and was rebuilt | Every figure maps to a saved query, the snapshot date and the seed, and a rerun must reproduce it exactly | Not traceable. The original table was not kept, and Q3's figures were rerun on newer data before any comparison could be made. |
| Seeded resampling | Without a fixed random starting point (a seed), the ranges moved a few tenths of a percentage point between runs, enough to change whether one range showed a clear change | Fixed seeds. Each likely range is the median of the seeded runs. | Every likely range reproduces exactly on a rerun |
| Tie-out across the site | The homepage band carried the 24 September figures while the piece had moved to 25 September. They were tied by hand. | Every figure used outside its piece declares its source, and the check must find it there | Homepage band tied to the 25 September figures |
| Shared definitions | "Same companies" meant a set of 89 in one piece and 68 in three others | One definitions file, used by every piece | At that stage one set of 68 everywhere; 74 in this release |
| Table bases (cold-read review) | A reviewer read Q7's "196 and 1,266 job boards" as a base copied from the next row | A cell headed "Base" must hold a whole figure written by the code that computes the piece | None. The figure was right, and the check now proves it. |
| Prepared inputs | The prepared inputs were saved by data version only, so a change in definitions could reuse a stale build | Saved inputs are named for the definitions they were built from | None in this release. Every figure is rebuilt from inputs named for the current definitions, and a rerun reproduces it exactly. |
| When a refresh stops for review | The refresh stopped on a change in how a likely range is described, not on a change of direction or rating, the threshold that was set | It stops on a change of direction or rating. A change in description is a wording update. | None on the figures |
| Q6's dated second read | A refresh would have moved the 25 September second read onto newer postings | Pinned to the postings it covered on 25 September | Q6's second read keeps its original postings |
| Figures moved by a refresh | Text could keep a figure after a refresh had changed it | A check carries moved figures into the piece text. Values it cannot place are passed to a person. | None on the figures |
| Refresh coverage | The refresh and its checks did not yet cover Q9 and Q10 | They cover every piece | None on the figures |
| Records of each read | A read could be recorded against the current text without re-reading it | A read is recorded only when the text was actually re-read | None on the figures |
Text against evidence
| Check | What it found | What changed | Effect on the findings |
|---|---|---|---|
| GPT and Grok on every finding in Q2 to Q7, 4,771 US postings in both years | Some findings did not hold on GPT's and Grok's labels, and Q2's same-company cut had included 20 employers not confirmed as US-headquartered | Text and ratings follow GPT's and Grok's labels | Q2 executives "flat" became "no clear change". Q3's claim that the degree gap widened since 2024 was dropped (likely range −6 to +22). Q4 AI expectations went from about 45% to 37.7%, measured directly, and executive AI governance from 16% to 22%; two specialty claims were removed. Q5: most of the startup-against-enterprise gap turned out to be industry mix. |
| First editorial read of the method page and Q1 (cross-family read) | A population labelled US-located that was US-headquartered, claims past the measures, a missing agreement figure for the hands-on label, counts under a "share" header, and an unclear date | Each point fixed | Wording only |
| Quotes checked word for word, Q1 to Q7 | Every quoted posting in Q1 to Q7 checked against its source | None needed | Check only, no headline change |
| Size words (cold-read review) | Q1's lead said the analyst specialties moved "sharply", while two of them were rated Medium | A vague size word in a title, description, lead or summary finding needs a High finding in every part | Q1's lead now says the analyst specialties "moved toward engineering" |
| One meaning of "companies" (cold-read review) | The method page gave two company counts two paragraphs apart, 1,857 and 2,383 | A bare "N companies" is always the security set's company count. Subsets are named. | One company count on this page: 2,537 in this release |
| Cross-family reads, first full pass (Q1 to Q7) | About 67 points on which the text went past the data | Text changed on each point, or the point was recorded with a ruling | Mostly wording. Some points moved a title or a rating under the rating rules: Q4's title went from "It Depends on the Level" to "It Differs by Level", Q3's operations finding was narrowed to the roles where its lead is clear, and three Q5 ratings (degrees, CISSP and securing AI) moved from High to Medium once the check outside technology companies was added. |
| Cross-family reads, later passes (Q1, Q6) | 5 more points | Text changed | Wording only. Q1's description now says the specialties "moved toward engineering" rather than "became engineering jobs". No rating changed. |
| Cross-family read of the refreshed pieces (Q1, Q2, Q8) | 14 points | Each point fixed or recorded. Five wording fixes reached the site. | Wording only |
| Cross-family reads of Q9 and Q10 | 14 points. Q9: directional wording on ranges that include zero. "Like for like" claimed more than the comparison holds. Change cells showed +16 and +2 where the ranges include zero. Q10: borderline rises in the same-company and similar-length tests. One sentence picked companies by their 2026 outcome. One top-employer share was quoted without its figure. | Q9: "no measurable change" and "Not measurable" where the range includes zero, and "held to the same title rules" in place of "like for like". Q10: "likely but borderline" with the test ranges shown, the outcome-selected sentence replaced by same-company rates for one defined group (3.8% to 16.5%), and the top-two-employer share stated. | Q10 background checks moved to Medium (see "Q10 no-change finding"). Other ratings unchanged. |
| Q1 description against its summary table | The privacy clause contradicted summary row 5 | Clause dropped | None |
| Q5 base against the method page | Q5's 6,185 postings and the method page's 14,446 had no stated relation | Q5 says how the two differ | None |
| Q2's account of the early-access figure (independent review) | The account put the reported fall down to rules applied to one year only, the 2024 scope check among them. That check lowered the 2024 manager share, and the fall came mainly from the 2026 side. | Q2's stress tests now state that an early-access draft reported the fall, that it came mainly from the 2026 query and collector, and that under one set of rules for both years it does not survive | None on the figures |
| Homepage measure against refreshes | Q4 quoted the homepage measure's figure, which a refresh could leave stale | Q4 names the measure without quoting the figure | None |
Review and confidentiality
| Check | What it found | What changed | Effect on the findings |
|---|---|---|---|
| Two blind reviews, by Claude Opus and OpenAI's Codex | Each recomputed the headlines from the data. Their findings are the rows marked "independent review". | Each finding was checked against the data before it was adopted | Every headline recomputed. Findings that did not reproduce were not adopted (see "Q2 executive share"). |
| Q2 executive share under the window (independent review) | The review's recompute on the 23 September view reported a rise in the executive share, +1.6 points (0.1 to 3.8), where Q2 reports no clear change | Checked under the series' definitions, where it does not reproduce: 3.8% to 5.3%, a likely range of −0.4 to +3.7 that includes zero. The review's figure came from a different filter. | None. Q2 keeps "no clear change". |
| Named people | The first-name check missed any first name not on its list | A local model reads every piece, the method page and the changelog for named people. It runs on demand, not with every check. | None on the figures |
| Credited askers | The check had no rule for people who suggested a question and agreed to be credited | Credited askers may be named in credit lines, never inside a piece | None |
| Confidentiality banner and readability | The banner check matched only one letter case, and readability did not run with the other checks | The banner is matched in any case, and readability runs with the other checks | None on the figures |
| The cross-family read's output | The reviewing model's output could repeat the piece back | Output captured, so the piece is not repeated | None |
The controls every piece must pass
The same controls run before a piece publishes and again before any update. Five stop a piece when they fail: reproducibility, the tie-out across the site, the rating rules, the confidentiality checks and the change log. Three print a report instead of stopping a piece: readability and accessibility, shared definitions, and the statistical checks. Two reads run on demand, when a piece changes in substance, rather than on every run: the cross-family read and the read for named people.
- Readable and accessible. A reading grade of 12 or lower, short sentences and paragraphs, acronyms spelled out, and every page tested against WCAG 2.2 AA.
- Reproducible. Every figure is rebuilt from saved queries and pinned data with fixed random seeds, and a rerun reproduces it exactly.
- Tied out. A figure quoted anywhere else on this site, including the homepage, matches the piece it comes from.
- One set of definitions. Terms the pieces share, such as the same companies or the AI-first cohort, have one definition that every piece uses.
- Statistically sound. A rise or fall is claimed only when its range of likely values excludes no change, and it has to hold under the standard checks and with the largest contributing company removed. Under the borderline rule above, a range near zero is described as a likely but borderline change or as no clear change, but only just.
- Claims match the evidence. Ratings follow the scale above. Nothing claims a cause from postings or reads as headcount, and a model from a different family reads each piece, when it changes in substance, for claims that go past the data.
- Confidential. The company list, the labels and the names of individuals are never published.
- Changes are logged. A change to a figure, a rating or a finding gets a tagged release and a changelog entry.
What the Research Factory is
The Research Factory is the agents that make this research and the guardrails and controls our team sets for them. Agents collect the postings, run the analysis, check every figure against the source and write each piece, start to finish. Our team sets the rating scales, the standard stress tests, the rules learned from past mistakes and the checks every piece must pass before it publishes. It is the same approach we take to AI-native software. The label at the top of each piece says so.
The research is versioned like open-source code. Each release is tagged, and every change to a figure, a rating or a finding is logged in the changelog.
The questions in this series
Ask the Research Factory a question.
This is not a live chat. Accepted questions become new pieces, published in a later release.
- 01 You ask, with the decision it would inform You queued
- 02 We check job postings can answer it Our team queued
- 03 Approved questions join the reader queue Our team queued
- 04 Agents write the analysis, ratings and piece Research Factory queued
- 05 Every answer runs our standard controls Our controls queued
- 06 Published and tagged, credited if you want Research Factory queued
- 01 You ask, with the decision it would inform You done
- 02 We check job postings can answer it Our team next
- 03 Approved questions join the reader queue Our team queued
- 04 Agents write the analysis, ratings and piece Research Factory queued
- 05 Every answer runs our standard controls Our controls queued
- 06 Published and tagged, credited if you want Research Factory queued