All research

Security Hiring Research: What Job Postings Say About Security Work

Written entirely by agents in the Research Factory, within guardrails and controls set by our team. Figures as of 1 October 2026.

How each run worksFall 2026
  1. 01 Read about 200,000 postings a day from about 2,200 job boards done
  2. 02 Keep about 17,000 security and engineering postings since July done
  3. 03 Score labels against one to three frontier model families done
  4. 04 Recompute every figure from source before publishing done
  5. 05 Publish only when every control passes done

A job posting is the most detailed public statement a company makes about the security work it needs done. It names the tools, the level of hands-on work, the certifications, the reporting line and, more and more, what the hire is expected to do with AI. Read one posting and you learn about one team. Read 2,194 job boards every day and changes in the work show up in the text.

This series does that. Each piece takes one question about security hiring and answers it from the data, with the numbers, a confidence rating, a representativeness rating and the tests the answer survived.

What we collect

The daily read

Every day, agents read 2,194 job boards across 16 applicant tracking systems, including Greenhouse, Ashby, Lever and Workday. On 1 October 2026 that pass read 220,747 open postings. Collection began on 31 July 2026.

The company list follows employers that hire security and AI engineering talent. It grew from public company lists, a cohort of AI-first companies and daily sweeps for new boards.

The security and engineering set

From that daily read, the agents keep the full text of every security posting and a sample of engineering postings for comparison. When a posting changes, the earlier version is kept too. Since 31 July the set has collected 15,163 security postings and 2,886 engineering postings from 2,537 companies.

Public baselines for comparison

Two months of collection cannot show a trend on its own. For questions about change over time, each piece pairs the current data with a public dataset from an earlier year, run through exactly the same rules. The first baseline is the LinkedIn Job Postings dataset published on Hugging Face (datastax/linkedin_job_listings), covering US postings listed between 5 and 19 April 2024.

One population for every comparison

Every piece that compares 2024 with 2026 reads one population, built the same way for both years. A posting is in when its title passes the series' security title rules and it carries a description. A local model then reads the title and the opening of each posting and rules whether the job is security work at all, which removes postings such as financial auditors, guards, lawyers and alarm installers. The same model, prompt and settings read both years. Last, each company counts each title once. Physical security, fraud and trust and safety roles are out of scope in both years, and so are job aggregators, which repost other employers' jobs.

The 2026 side holds postings open on or after 23 September 2026. Until the evening of 21 September the job-board collector skipped titles containing analyst, auditor, consultant, CISO, director, head of, specialist and "manager, security", so postings that closed before then lean toward engineer titles. The window leaves 4,587 postings at 831 US-headquartered companies, against 1,038 US postings in spring 2024.

How a question gets answered

Most measures are written rules applied identically to every posting: the level implied by a title, the specialty, whether the description asks for a programming language, automation or a certification. Rules are checked by having a frontier model, one of the most capable AI models available, read a sample without seeing the rule's answers. Title level agreed with Claude Opus at a Cohen's kappa of 0.975, specialty at 0.76. Kappa runs from 0, no better than chance, to 1, perfect agreement.

Some measures need judgement, such as whether the hire manages people or personally does the technical work. For those, a local model reads every posting, and its answers are checked against the answers of one to three frontier models from different AI companies. In Q1 to Q4, three AI models each read the postings, GPT, Grok and the local model, and the pieces use the answer at least two agree on. In Q5, GPT and Grok check the local model's answers on US postings. In Q6 the three models are Claude Opus, GPT and Grok. Each piece reports the agreement it relies on.

How every answer is rated

A posting is what a company says it wants. It is not a headcount, and it is not proof that anyone was hired. Every answer in this series is about stated expectations, and each one carries two separate ratings.

Each figure comes with a range of likely values. We get it by redrawing the companies in our data at random many times and recomputing the figure each time, so the range shows how much the answer could move if a different set of companies had been collected. The range we show holds 95% of those recomputed values. If that range includes zero, we can't tell the change from no change.

Confidence is how likely the number is right for the postings we read.

High The measure agrees with frontier models at a kappa of 0.7 or better, and the result holds with duplicate postings removed, with each company counted once, and across its range of likely values.
Medium The direction holds in every check, but the size moves by more than a few percentage points between methods, or the measure agrees only moderately with frontier models.
Low One method supports it, or its range of likely values includes zero, so the change is too small to tell apart from no change.
Not answerable The data cannot speak to the question yet, and the piece says so instead of guessing.

Three rules sit under the scale, and the pieces use a few terms for them. A change is a clear rise or a clear fall when its whole range of likely values sits on one side of zero, so the range excludes no change. That makes it a real rise or fall, too large to be explained by which postings happened to be collected. When the range includes zero, the change is too small to tell apart from no change, and the pieces call it not measurable or no clear change.

A finding that there is no difference is rated by how narrow its range is: High when the whole range sits within 5 percentage points of zero, Medium within 10, and otherwise the difference is not measurable. A comparison with fewer than 30 postings or 10 companies on either side is rated Medium at most.

Each range is computed several times, each time from a different random starting point. A range counts as borderline when it excludes zero but ends within 1 percentage point of it, or within 5% of the comparison group's value for a ratio. It also counts as borderline when it excludes zero in some of those computations but not in all. A borderline range that excludes zero is called a likely but borderline rise or fall. The change is probably real, but only just. A borderline range that reaches zero or sits just past it is called no clear change, but only just. It includes no change, by a small margin. The pieces call both kinds borderline.

One more rule covers time. A High rating for a difference needs the difference to have been clear in the previous run on the same rules as well. Without an earlier run on the same definitions, a finding is rated Medium at most, whether it is a difference or a finding of no difference, and the 30 September run was the first under the rules above. This release, on 1 October, is the second. A rating can be overridden downward, with a stated reason, but never upward, and an override that sets High stops the run.

Representativeness is how far the answer generalises to security hiring as a whole. The company list leans toward technology companies, US headquarters and employers that post on modern applicant tracking systems. It under-covers government hiring, staffing and managed-security firms, and companies that post only on LinkedIn or Indeed. An answer can be high on confidence and low on representativeness, and several are.

Representativeness has three levels. High means the answer holds across the companies the list covers well and the ones it under-covers. Medium means the answer rests on security postings broadly, and the list's lean could change its size more than its direction. Low means the answer rests on one level, specialty, industry or segment of postings, or on a group the list under-covers, or part of it comes from which companies happen to be in each year's data.

How every answer is stress-tested

Every answer goes through the same checks before it is published. Each piece reports what the checks found, including the ones that changed the answer.

  • Repeat checks. Every comparison is repeated on the same companies in both years where possible, outside technology companies, on postings of similar length, and with each company counted once. Outside technology companies means 2026 employers in healthcare, retail and e-commerce, manufacturing, defense and government, staffing and consulting, media and entertainment, consumer goods and financial services.
  • Measure checks. In every period compared, keyword matches are read by hand to see how many are genuine. Title rules and model labels are compared with frontier models' answers, without either seeing the other's.
  • Scope checks. A local model, run on our own hardware, reads every posting in both years and rules whether it is a security role. Before use it was scored against Claude Opus on every 2024 posting, and a person read the postings where the two disagreed. That score was measured on the same 2024 postings used to tune the prompt, so it may read higher than it would on new postings. The prompt was amended once after the first pass over the 1,444 2024 postings, the same prompt then ran on both years, and Claude Opus did not read the 2026 postings.
  • Ranges of likely values. Every change carries a range of likely values, built by redrawing whole companies rather than single postings. So one large employer with dozens of similar postings cannot carry a result on its own.
  • A fresh run. Every figure is recomputed from source data before a piece is published or updated, and the date at the top of the piece says when. Each release pins its figures to one day's collection. Version 1.0.0 rests on the 1 October 2026 collection. Figures change only in a new release, and the changelog says what moved and why. A finding that no longer holds comes down. On 24 September a refresh showed that one of the four findings on our homepage no longer cleared zero, and it was replaced the same evening.

Each piece closes with the tests that would strengthen its answer and why they cannot run yet.

The postings quoted in this series are public. The company list and the labels behind the numbers are not published.

Some of these checks were added when a problem turned up, and each now applies to every piece. The tables below list every check the series has passed, grouped by what it looks at: the data, the measures, the ratings, reproducibility, the text and review. Each row gives what the check found, what changed and the effect on the findings.

Q1 to Q10 are the ten questions listed at the end of this page. Three kinds of review appear in the rows:

  • A cold-read review is a read of the finished pieces by a reviewer who had not worked on them.
  • An independent review is a blind review, by Claude Opus or OpenAI's Codex, that recomputed the figures from the data.
  • A cross-family read is an AI model from a different company reading a piece for claims that go past the data.

"At that stage" marks figures from an intermediate run; each piece carries the current figures. "Standing rule from before the series" marks a rule adopted before the first piece was drafted.

Data and population

CheckWhat it foundWhat changedEffect on the findings
Posting counts over timePosting counts grew from 2,462 to 6,432 in 38 days, almost all of it from adding job boardsNo growth claim rests on posting counts. Pieces compare shares against a baseline.Standing rule from before the series
Posting datesPosting dates come from the applicant tracking system and favour postings still openPosting age is never read as a trendStanding rule from before the series
Samples read from every specialty and level, both years2024 carried guards, trades advertised "with security clearance" and psychiatric nurses. 2026 carried chip-design "SoC" titles.Removed by title rule, in both yearsIn this release the title rules remove 414 2024 postings and 127 2026 postings at US-headquartered companies before the scope check: guard, physical-security and similar titles (405 and 108), chip-design and clinical titles (8 and 19), and one 2024 title that read as security only for its clearance wording
Duplicate postings, both years2024 carried repeat postings, the same title at the same company, that the 2026 side had already collapsedOne posting per company and title, in both yearsIn this release 107 repeats leave the 2024 comparison set and none leave 2026. Kept, they would put the 2024 manager share at 8.8% rather than 7.7%, engineer titles at 34.5% rather than 34.0%, and language asks at 22.5% rather than 22.0%.
Scope check of the 2024 set: Claude Opus read all 1,608 2024 postings with a security title27% were not security roles: financial auditors, physical security, alarm installers, a disability attorney. Among leadership titles it was 41%.Removed from the 2024 set. Later replaced by one scope check run on both years (see "Same rules in both years: the 2024 side").2024 detection and response automation rose from 28% to 41%, and the certification drops in detection and response and in GRC became visible
Scope check of Q6's 2024 leadership set: Opus read all 245 leadership-titled 2024 postings100 were not security rolesQ6's 2024 comparison uses in-scope people managers only72 people managers in this release, after the series-wide scope check and the aggregator exclusion
Same rules in both years: the 2026 side, compared line by line with 20242026 postings counted only if the job-board collector had tagged the title with a security function. The tag missed titles such as "Information Security Manager", spelled-out CISO titles and "…Engineering Manager". 2,610 stored postings had no tag, and 2,071 of them were in scope; about 27 of 30 sampled were real security jobs. A further 2,074 untagged postings first seen 31 July to 21 September were dropped: 88% engineer titles, 81% closed.2026 is no longer filtered by the tag. Both years pass the same title rules.The manager-share fall in an early-access draft of Q2 (8.8% to 3.5%) came mainly from this side. At that stage, on one population, the manager share was 7.6% in 2024 and 6.4% in 2026, a change of −1.2 points (−2.5 to +0.4), with no measurable fall. Q2's title changed. Engineer titles (Q9) went from 34.5% to 61.7%, and programming-language asks (Q1) from 22.1% to 42.7%.
Same rules in both years: the 2024 sideAn Opus scope check removed 438 of the 2024 postings (26%), and nothing like it ran on 2026One scope check now runs on both years, by a local model with the same prompt and settings. The 2024-only check was removed.The 2024-only check had lowered the 2024 manager share, so it worked against the fall, not toward it. At that stage: 9.8% with no scope check, 7.8% with the Opus check, 7.6% with the local check. In this release: 10.0% without a scope check, 7.7% with it.
A first fix, reviewed before mergeThe first fix narrowed 2024 to the 2026 tag. The independent review found the work correct, with every headline recomputing, but the approach wrong: the tag drops real security titles, and narrowing 2024 copies those misses into both years.Not merged. Replaced by widening 2026 and applying one rule set and one scope check to both years, then reviewed again before merge.Widening rather than narrowing, per the review: manager share 7.8% to 7.4%, engineer titles 33.7% to 57.2%, language asks 21.6% to 39.6%. Every direction holds.
Q9 held to the same title rule and de-duplication in both yearsQ9 had not applied the same title rule and de-duplication to both yearsBoth years pass the same rule and de-duplicationEngineer titles 34% to 63% became 39% to 63%. On the one population that followed, 34% to 62%, and 34% to 58% in this release.
Every piece rerun on the one populationFindings that rested on populations built differently in each yearQ1 to Q10 recomputed, and their text and ratings updatedAt that stage: Q1 application security titles, certifications and cloud titles are no longer measurable. Q2: the manager fall does not survive. Q3: the fall in the entry-level share is a clear fall, the CISSP fall is High, and the operations lead over privacy is no longer measurable. Q4: AI expectations are rated High, executives' lead over managers on AI governance is clear, and the manager tool-use gap is borderline. Q5: industry table re-ranked, overall AI gap by size rated Medium. Q6: the executive deciding gap is no longer measurable. Q8: posted pay +27.2%, with a level-mix check; leadership and managers a likely but borderline rise; individual contributors High. Q9: engineer titles 34% to 62%, and manager titles rise. Q10: findings and ratings unchanged.
Collection history, upstream of the series' own rules (independent review)Before the evening of 21 September the 2026 collector did not collect titles containing analyst, auditor, consultant, CISO, "manager, security", director, "head of" or specialist. 904 postings from before that change were still in the population, 81% with engineer titles. Postings collected before the change were 82.5% engineer titles and 0.5% executive; postings collected after it, 35.0% and 9.8%. The series' own rules cannot see this.2026 is limited to postings open on or after 23 September 2026. Both views are shown (see "Both views side by side").On the 30 September data, when the review found this, the window left 4,524 2026 postings at 830 US-headquartered companies, from 5,339 at 861. On it at that stage, with aggregators out: engineer titles 34% to 58% (+20 to +29 points), language asks 22% to 43%, the manager share 7.7% to 6.7% (−3.1 to +1.0) and the executive share 3.8% to 5.3% (−0.4 to +3.7). Every direction holds. The 2026 engineer-title share is 4 points lower than on the full set. The pieces carry the 1 October figures: 4,587 postings, the manager share 6.8% (−3.0 to +1.0) and the executive share −0.4 to +3.8.
Company concentration (independent review)Companies with 20 or more postings made up 39.5% of 2026 and 2.0% of 2024. The largest 2026 "company" was Jobgether, a job aggregator, with 229 templated postings and a language-ask share of 3.1%.Job aggregators, which repost other employers' jobs, are left out in both years: Jobgether, Dice, ClearanceJobs, Jobs via eFinancialCareers, Talentify.io and The Job Network33 postings leave 2024 (21 from Dice, 8 from ClearanceJobs) and 229 leave 2026, all Jobgether. On the full 2026 set, language asks move from 43% to 44%. Q5's startup degree gap moves from clear to likely but borderline, and its headline from "about half as often" to "about 60% as often". Q1's architecture decision role now shows no clear change, and Q9's executive engineer-title rise is now clear. Q6's fixed panel loses two Jobgether postings, and one that closed before 23 September; with both rules its headline moves from 76% against 44% to 78% against 43%. Q7's own snapshot loses 568 security postings and one job board, and AI engineering openings per security opening at other companies move from 0.68 to 0.64 on the 30 September data (0.65 on 1 October, as Q7 reports).
Both views side by side: every 2026 posting against postings open on or after 23 September, both without aggregatorsWhether the window changes any headlineBoth views computed in full. The series publishes the window.On the 30 September data: Full set, then window. Q1 language asks 22% to 44%, and 22% to 43%. Q2 manager share 7.7% to 6.4%, and 7.7% to 6.7%. Q3 share of postings open to zero to two years' experience 18% to 13% in both. Q4 postings with any AI expectation 3% to 37%, and 3% to 36%. Q5 startup degree share 9.8% against 16.0% at enterprises, and 9.9% against 16.3%. Q6 hands-on work in manager postings that expect AI 78% against 44% without, and 78% against 43%. Q7 reads its own snapshot and is the same in both. Q8 posted pay +27.2%, and +28.3%. Q9 engineer titles 34% to 62%, and 34% to 58%. Q10 scam warnings 0.7% to 11.8%, and 0.7% to 11.4%. Every direction holds and every size is within 4 points, Q9 the largest. Four secondary findings move: Q2 people management (borderline to no clear change), the Q4 manager tool-use gap (borderline to no clear change), the Q5 AI gap by size (no clear change to borderline) and Q1 vulnerability-management engineer titles (clear to borderline).

Measures and labels

CheckWhat it foundWhat changedEffect on the findings
Labels scored against another model familyA model label agreed with itself but scored a kappa of 0.18 against ClaudeLabels are scored against a different model family, never against themselvesStanding rule from before the series
Keyword hits read by handKeyword matches on company boilerplate, such as vendor marketing copyKeyword precision is read by hand before a measure is usedStanding rule from before the series
The agent, LLM and MCP keyword count, read by handCounted over the whole posting, 13 matches in 20 were genuine, most of the rest company boilerplateCounted only in the responsibilities and requirements sections (19 in 20), marked in Q7 as added after the runThe figure on our homepage and in our company post went from 49% against 27% to 26% against 16%. Q7 shows 26% against 17% in this release. In Q7 the scale gave High, and the rating is set to Medium: the measure was added after the run and combines securing AI with using AI for security work.
Title rules against Claude Opus, 192 titles labelled blind across both yearsLevel agreement was a kappa of 0.94. The misses showed that "Directory Services" matched "director", so engineers were counted as executives.Rule fixedLevel agreement rose to 0.975
The dataset's own seniority tag, auditedTitles naming a domain called "management" counted as managers. 184 of 495 "manager" postings were engineers ("Security Engineer, Vulnerability Management").Fixed and applied to every stored posting. The pieces use their own title rule.Q5 management share by company size went from 14.0 / 8.4 / 5.9% to 9.8 / 6.4 / 4.6%, and the gap between sizes held. The Q4 manager rate for securing AI went from 4.6% to 3.9%. The Q6 panel moved by at most a point.
Q7 domain labels against a keyword rule, then against a blind Claude Opus read of 100 postingsModel against rule: a kappa of 0.52, below the 0.6 barAn Opus check was added and marked as added after the run. Domain shares come from the model.Model against Opus 0.75, rule against Opus 0.51
Q7 reporting lines against Opus, 65 postingsIn 11 of the 20 cases where only the model saw a reporting line, the line was not thereOnly lines that both the model and a keyword rule find are countedQ7 reports no measurable difference in reporting lines (−9.8 to +17.5 points), on lines both methods find
The local labeller on both years, scored against frontier labelsIt agreed poorly on hands-on work. Tuning raised its people-management agreement from 0.69 to 0.82 on a held-out set, but not its hands-on agreement. After relabelling both years on 28 September, it agrees with GPT at 0.56 on hands-on work.The hands-on and architecture findings rest on GPT and Grok labelsThe Q1 answer moved from "more hands-on" to "the kind of hands-on work changed"
Choice of scope-check model: two local models scored against Claude Opus on a weighted sample of 200 2024 postings, with finding out-of-scope postings as the taskThe faster model missed out-of-scope postings. It caught 79.6% of them (recall 0.796), and 93.5% of the postings it flagged were truly out of scope (precision 0.935, kappa 0.812). It gave no answer on 7 postingsRejected for low recall. The chosen model scored precision 0.920, recall 0.940 and kappa 0.904 on the sample.Used for every scope ruling in both years
Scope-check prompt, disagreements read by handThe model ruled IT SOX auditors (for example "IT SOX Sr. Auditor, Internal Audit") to be financial auditThe prompt now says that IT audit (IT general controls, IT SOX controls, technology audit) and technology or IT risk count as security work. All of 2024 was re-read, and both years ran on the amended prompt. The prompt was tuned on 2024 only, which the independent review recorded as a mild bias.Precision against Opus on 2024 rose from 0.88 to 0.936
Scope-check calibration on all 1,444 2024 postings, recomputed by the independent reviewPrecision 0.936, recall 0.902, kappa 0.889. It rightly flagged 349 postings as out of scope, wrongly flagged 24 and missed 38. 3.5% of kept 2024 postings were out of scope by Opus, and 6.4% of dropped ones were in. A hand read of 40 disagreements, 20 each way, sided with the local model in about 29. The calibration is in-sample: the prompt was amended once after the first pass over 2024, and there is no Opus reference for 2026.Adopted as the scope check for both yearsOn its own reading of 2026, the review found 40 of 40 kept postings were real security roles and about 20 of 22 drops correct. The model drops 25.8% of 2024 and 10.0% of 2026; the difference fits the 2026 collector, which already keeps only security postings.
Label versions across years, recomputed by the independent reviewThe same versions label both years: the local model on the 28 September prompt, GPT and Grok at fixed settings. On 30 September, 2 of 4,524 2026 postings lacked a local label. On 1 October, 1 of 4,587 did. 6 postings with two labels on file take either one; each pair agrees.None neededNone
Q8 pay parser against an independent re-parse (independent review)A rough independent parse gave +21.1% (+12 to +30), or +17.5% at the 2024 level mix, against the piece's +27.2% and +23.2%. An earlier check against a blind Claude Opus read of 300 postings had passed: 99.0% precision, 97.5% recall, 194 of 198 ranges within 2%.The gap was traced posting by posting, and the parser stays as it is. The rough parse found pay in about 1,500 2026 postings, the parser in 2,272. About 800 of the ranges it missed sit in job-board pay-range markup, where the dash between the two amounts is a character code, and they run higher (median $192k). It also counted about 30 ranges the parser leaves out on purpose: Canadian dollars, on-target earnings, and ranges whose top is more than three times the bottom. The parser misses 4 US ranges the rough parse reads.The gap adds up: +21.1% becomes +26.0% with the missed ranges, +26.5% taking the highest tier where a posting lists several, +28.7% with pay the job board returned for 248 postings, and +27.2% with the 2024 ranges read from the posting text. No figure changed. Under the 23 September window the piece reports +28.3% (+19 to +37), and +23.6% at the 2024 level mix.
Q8 pay coverageThe collector did not keep the Ashby and Lever pay field, so pay for postings that closed before 30 September could not be readCollection keeps the field, and open postings were filled inWithout that field the rise was +19.2% (+13 to +30) at the time. In this release it is +25.1% (+17 to +34) without the field, against +27.4% with it.

Statistics and ratings

CheckWhat it foundWhat changedEffect on the findings
Distinctiveness against a controlA "this is unusual" finding, close to publication, failed when the same measure was run on the comparison groupEvery claim that something is unusual is first run on the comparison groupThe finding was not published
The confidence scale, applied by codeRatings had been set by judgement. Some Highs rested on borderline ranges, or on controls whose range included no change.Code applies the scale to every rating. A control whose range includes zero drops High to Medium, and so does a kappa under 0.7 or a spread of more than 5 points between methods. Small bases are capped at Medium. No-difference findings are rated by the width of the likely range.At that stage: the Q3 entry-level share moved from High to Medium (borderline), Q4 AI expectations from High to Medium, the Q1 amount of hands-on work from High to Low (not measurable), and Q2 executives from Low to Medium
How the rules read the end of a range (cold-read review)The rules compared the unrounded end of the range, so an end shown as "+1" passed as a clear change. Q1's architecture decision role showed "+1 to +26" and was rated High. Four other ranges in Q1, Q3 and Q4 had the same gap.The rules read each end of a range as the piece prints it. A range that excludes zero with a printed end inside the borderline margin is described as likely but borderline, and the wording and ratings follow.The five ranges are described as likely but borderline
Stability across runs (cold-read review)Q7's security-duty finding failed on 23 September, passed on 25 September and was rated High. It had replaced a homepage card that failed the same way the day before.High needs the main range to show a clear change in the previous recorded run tooQ7's security-duty finding passed the 25, 28 and 30 September runs. It was held at Medium on 30 September by the first-run rule, and is High again in this release after passing on 1 October.
Run history for the stability rule (independent review of the first fix)Earlier runs were recorded by data version only, so a run under different definitions still countedRuns are recorded by data version and definitionsRuns on other definitions no longer count
The first-run rule (independent review)With no earlier run on the same definitions, the stability rule passed automatically. No recorded run carried its definitions, so every piece passed.With no earlier run on the same definitions, a rating is capped at Medium, for a difference and for a finding of no differenceThe 30 September run was the first on its definitions, so every High moved to Medium: 27 ratings across all ten pieces (Q1 3, Q2 2, Q3 5, Q4 3, Q5 3, Q6 1, Q7 3, Q8 1, Q9 5, Q10 1), including Q2's finding that the generic leadership title held. The 1 October run, the second under the same rules, restored all 27.
Overrides of the scaleAn override that set a rating higher than the scale gives was reported but did not stop the runAn override can lower a rating, with a stated reason, and never raise it. One that sets High fails the run.No rating in this release is raised by an override
Q10 no-change finding against every test (cross-family read)Background checks were rated High as no change, on a headline likely range within 5 points of zero. At that stage the each-employer-once test showed a clear rise (+0.2 to +5.3), and the same-company test ran to −7.5.Rating overriddenBackground checks went from High to Medium. In this release they stay at Medium: counted once per employer the rate shows a likely but borderline rise (+0.3 to +5.4), and the test on 2024 postings applied through a company job board runs to −6.3.
Q5 rating of a level, not a directionThe scale tests direction against zero, which a level of 61% passes trivially. Defense's clearance share rests on 37 defense and government companies, with a likely range of 38% to 85%, and the middle of the ranking has no adjacent-rank test.Rating set to MediumHeld at Medium, not High
Rating levels (cold-read review)Pieces used in-between levels with no definition, such as Medium-low and Low-mediumBoth scales use only High, Medium, Low and Not answerable11 representativeness ratings in Q1 and Q3 to Q7, each Medium-low or Low-medium, became Low. No confidence rating had used an in-between level.
Fresh run before publishingOn 24 September a refresh showed that one of the four homepage findings no longer showed a clear changeReplaced the same eveningThe finding came down

Reproducibility

CheckWhat it foundWhat changedEffect on the findings
Saved queriesQ3's table query had never been saved and was rebuiltEvery figure maps to a saved query, the snapshot date and the seed, and a rerun must reproduce it exactlyNot traceable. The original table was not kept, and Q3's figures were rerun on newer data before any comparison could be made.
Seeded resamplingWithout a fixed random starting point (a seed), the ranges moved a few tenths of a percentage point between runs, enough to change whether one range showed a clear changeFixed seeds. Each likely range is the median of the seeded runs.Every likely range reproduces exactly on a rerun
Tie-out across the siteThe homepage band carried the 24 September figures while the piece had moved to 25 September. They were tied by hand.Every figure used outside its piece declares its source, and the check must find it thereHomepage band tied to the 25 September figures
Shared definitions"Same companies" meant a set of 89 in one piece and 68 in three othersOne definitions file, used by every pieceAt that stage one set of 68 everywhere; 74 in this release
Table bases (cold-read review)A reviewer read Q7's "196 and 1,266 job boards" as a base copied from the next rowA cell headed "Base" must hold a whole figure written by the code that computes the pieceNone. The figure was right, and the check now proves it.
Prepared inputsThe prepared inputs were saved by data version only, so a change in definitions could reuse a stale buildSaved inputs are named for the definitions they were built fromNone in this release. Every figure is rebuilt from inputs named for the current definitions, and a rerun reproduces it exactly.
When a refresh stops for reviewThe refresh stopped on a change in how a likely range is described, not on a change of direction or rating, the threshold that was setIt stops on a change of direction or rating. A change in description is a wording update.None on the figures
Q6's dated second readA refresh would have moved the 25 September second read onto newer postingsPinned to the postings it covered on 25 SeptemberQ6's second read keeps its original postings
Figures moved by a refreshText could keep a figure after a refresh had changed itA check carries moved figures into the piece text. Values it cannot place are passed to a person.None on the figures
Refresh coverageThe refresh and its checks did not yet cover Q9 and Q10They cover every pieceNone on the figures
Records of each readA read could be recorded against the current text without re-reading itA read is recorded only when the text was actually re-readNone on the figures

Text against evidence

CheckWhat it foundWhat changedEffect on the findings
GPT and Grok on every finding in Q2 to Q7, 4,771 US postings in both yearsSome findings did not hold on GPT's and Grok's labels, and Q2's same-company cut had included 20 employers not confirmed as US-headquarteredText and ratings follow GPT's and Grok's labelsQ2 executives "flat" became "no clear change". Q3's claim that the degree gap widened since 2024 was dropped (likely range −6 to +22). Q4 AI expectations went from about 45% to 37.7%, measured directly, and executive AI governance from 16% to 22%; two specialty claims were removed. Q5: most of the startup-against-enterprise gap turned out to be industry mix.
First editorial read of the method page and Q1 (cross-family read)A population labelled US-located that was US-headquartered, claims past the measures, a missing agreement figure for the hands-on label, counts under a "share" header, and an unclear dateEach point fixedWording only
Quotes checked word for word, Q1 to Q7Every quoted posting in Q1 to Q7 checked against its sourceNone neededCheck only, no headline change
Size words (cold-read review)Q1's lead said the analyst specialties moved "sharply", while two of them were rated MediumA vague size word in a title, description, lead or summary finding needs a High finding in every partQ1's lead now says the analyst specialties "moved toward engineering"
One meaning of "companies" (cold-read review)The method page gave two company counts two paragraphs apart, 1,857 and 2,383A bare "N companies" is always the security set's company count. Subsets are named.One company count on this page: 2,537 in this release
Cross-family reads, first full pass (Q1 to Q7)About 67 points on which the text went past the dataText changed on each point, or the point was recorded with a rulingMostly wording. Some points moved a title or a rating under the rating rules: Q4's title went from "It Depends on the Level" to "It Differs by Level", Q3's operations finding was narrowed to the roles where its lead is clear, and three Q5 ratings (degrees, CISSP and securing AI) moved from High to Medium once the check outside technology companies was added.
Cross-family reads, later passes (Q1, Q6)5 more pointsText changedWording only. Q1's description now says the specialties "moved toward engineering" rather than "became engineering jobs". No rating changed.
Cross-family read of the refreshed pieces (Q1, Q2, Q8)14 pointsEach point fixed or recorded. Five wording fixes reached the site.Wording only
Cross-family reads of Q9 and Q1014 points. Q9: directional wording on ranges that include zero. "Like for like" claimed more than the comparison holds. Change cells showed +16 and +2 where the ranges include zero. Q10: borderline rises in the same-company and similar-length tests. One sentence picked companies by their 2026 outcome. One top-employer share was quoted without its figure.Q9: "no measurable change" and "Not measurable" where the range includes zero, and "held to the same title rules" in place of "like for like". Q10: "likely but borderline" with the test ranges shown, the outcome-selected sentence replaced by same-company rates for one defined group (3.8% to 16.5%), and the top-two-employer share stated.Q10 background checks moved to Medium (see "Q10 no-change finding"). Other ratings unchanged.
Q1 description against its summary tableThe privacy clause contradicted summary row 5Clause droppedNone
Q5 base against the method pageQ5's 6,185 postings and the method page's 14,446 had no stated relationQ5 says how the two differNone
Q2's account of the early-access figure (independent review)The account put the reported fall down to rules applied to one year only, the 2024 scope check among them. That check lowered the 2024 manager share, and the fall came mainly from the 2026 side.Q2's stress tests now state that an early-access draft reported the fall, that it came mainly from the 2026 query and collector, and that under one set of rules for both years it does not surviveNone on the figures
Homepage measure against refreshesQ4 quoted the homepage measure's figure, which a refresh could leave staleQ4 names the measure without quoting the figureNone

Review and confidentiality

CheckWhat it foundWhat changedEffect on the findings
Two blind reviews, by Claude Opus and OpenAI's CodexEach recomputed the headlines from the data. Their findings are the rows marked "independent review".Each finding was checked against the data before it was adoptedEvery headline recomputed. Findings that did not reproduce were not adopted (see "Q2 executive share").
Q2 executive share under the window (independent review)The review's recompute on the 23 September view reported a rise in the executive share, +1.6 points (0.1 to 3.8), where Q2 reports no clear changeChecked under the series' definitions, where it does not reproduce: 3.8% to 5.3%, a likely range of −0.4 to +3.7 that includes zero. The review's figure came from a different filter.None. Q2 keeps "no clear change".
Named peopleThe first-name check missed any first name not on its listA local model reads every piece, the method page and the changelog for named people. It runs on demand, not with every check.None on the figures
Credited askersThe check had no rule for people who suggested a question and agreed to be creditedCredited askers may be named in credit lines, never inside a pieceNone
Confidentiality banner and readabilityThe banner check matched only one letter case, and readability did not run with the other checksThe banner is matched in any case, and readability runs with the other checksNone on the figures
The cross-family read's outputThe reviewing model's output could repeat the piece backOutput captured, so the piece is not repeatedNone

The controls every piece must pass

The same controls run before a piece publishes and again before any update. Five stop a piece when they fail: reproducibility, the tie-out across the site, the rating rules, the confidentiality checks and the change log. Three print a report instead of stopping a piece: readability and accessibility, shared definitions, and the statistical checks. Two reads run on demand, when a piece changes in substance, rather than on every run: the cross-family read and the read for named people.

  • Readable and accessible. A reading grade of 12 or lower, short sentences and paragraphs, acronyms spelled out, and every page tested against WCAG 2.2 AA.
  • Reproducible. Every figure is rebuilt from saved queries and pinned data with fixed random seeds, and a rerun reproduces it exactly.
  • Tied out. A figure quoted anywhere else on this site, including the homepage, matches the piece it comes from.
  • One set of definitions. Terms the pieces share, such as the same companies or the AI-first cohort, have one definition that every piece uses.
  • Statistically sound. A rise or fall is claimed only when its range of likely values excludes no change, and it has to hold under the standard checks and with the largest contributing company removed. Under the borderline rule above, a range near zero is described as a likely but borderline change or as no clear change, but only just.
  • Claims match the evidence. Ratings follow the scale above. Nothing claims a cause from postings or reads as headcount, and a model from a different family reads each piece, when it changes in substance, for claims that go past the data.
  • Confidential. The company list, the labels and the names of individuals are never published.
  • Changes are logged. A change to a figure, a rating or a finding gets a tagged release and a changelog entry.

What the Research Factory is

The Research Factory is the agents that make this research and the guardrails and controls our team sets for them. Agents collect the postings, run the analysis, check every figure against the source and write each piece, start to finish. Our team sets the rating scales, the standard stress tests, the rules learned from past mistakes and the checks every piece must pass before it publishes. It is the same approach we take to AI-native software. The label at the top of each piece says so.

The research is versioned like open-source code. Each release is tagged, and every change to a figure, a rating or a finding is logged in the changelog.

The questions in this series

Ask the Research Factory a question.

This is not a live chat. Accepted questions become new pieces, published in a later release.

  1. 01 You ask, with the decision it would inform You
  2. 02 We check job postings can answer it Our team
  3. 03 Approved questions join the reader queue Our team
  4. 04 Agents write the analysis, ratings and piece Research Factory
  5. 05 Every answer runs our standard controls Our controls
  6. 06 Published and tagged, credited if you want Research Factory
Ask the Research Factory a question
Answers come as published pieces, not replies. Not every question can be answered from job postings.