Ghost Job
pre-registeredA posting that remains open, or is re-listed under the same deduplication key, for at least 90 days without disappearing from its source.
Observation window required: 90 days
How Reqbeat measures hiring demand — the sample, the deduplication rule, the known biases, and the metric definitions we have frozen in advance.
Version 1.12.0 · Last generated September 14, 2026
We track 4,781,303 job postings drawn from 55 distinct sources. Both of those figures describe what we ingest: postings admitted to the corpus, counted one row per posting as crawled, across the whole history we hold rather than a rolling window, and inclusive of postings sourced from LinkedIn. They state how wide the crawl reaches. They are deliberately not counted on the surface the market pages on this site read — the deduplicated, rolling 30-day view described under Deduplication and Vintage below, which is narrower in both respects.
What we measure is the volume of online job postings we observe, deduplicated to one row per company, job title and country.
Repeat and re-posted advertisements collapse to a single row per (company, job title, country). Where several postings share that key, we keep the most complete one — the row that resolves to a named company with an industry, headcount and location — and break remaining ties by the most recent posting date. Note the key is country, not city: two postings for the same role in two cities of one country count once, and the same role in two countries counts twice.
Partition key
lower(COALESCE(cp.name, s.company_name, s.url)), lower(s.title), s.country
Winner within each partition
completeness DESC (company name, industry, headcount, location present), posted_at DESC, id DESC
The corpus surface these pages read retains a rolling 30-day window on posting date. That window bounds what the market pages count, not what the corpus holds. The two breadth figures under The sample above are counted before it is applied, on the ingest population described there. Any metric requiring a longer observation span than that is listed below as pre-registered and is not published until the retention exists to support it.
These definitions are frozen and dated before any figure is computed under them, so a threshold can never be tuned after seeing a result. A metric marked pre-registered has no published figure — the reason is stated with it.
A posting that remains open, or is re-listed under the same deduplication key, for at least 90 days without disappearing from its source.
Observation window required: 90 days
The median number of days a requisition stays live, from the first time we observe it to the point it disappears from its source.
Observation window required: unbounded
For postings we observe on two or more sources, the median gap in hours between the first time the earliest source shows it to us and the first time the latest source does.
Observation window required: 30 days
A monthly index of advertised hiring demand. Each edition's observation is the count of deduplicated postings — one row per company, job title and country, staffing agencies excluded — posted in the 28 days ending on that edition's publication slot. The index expresses that observation as a percentage of the base edition's observation, which is fixed at 100.0. The base edition is the first edition published under this definition.
Observation window required: 28 days
For each reference month in which both series have a print, the Hiring-Demand Index's month-over-month percentage change minus the official series' month-over-month percentage change for that month, in percentage points, together with the share of those months in which the two moved in the same direction. An edition's reference month is the calendar month holding the majority of the 28 days its observation was counted over. The comparison is made against the official series' first published print for the month, never a later revision. Movement is compared, never level: this index counts advertised vacancies in a sample of public postings and does not estimate the vacancy stock an official survey measures.
Observation window required: unbounded
The whole-month shift — searched over 0 to 3 months — at which the Hiring-Demand Index's month-over-month direction agrees most often with the official series', reported together with the agreement at every shift tested so a shift that wins by one month is visible as winning by one month.
Observation window required: unbounded
Some figures on this site are readings rather than metrics: one measurement of a corpus that keeps moving, taken on a stated day. Nothing recomputes them, so each is published here with the population it describes, how it was derived, when it was taken and when it expires — 90 days later. Past that date our build fails until the figure is re-measured or the claim that quotes it is withdrawn.
The detection-latency readings below are published in full, with their derivations and their cohorts, at https://reqbeat.com/detection-latency.
2.1 hours
Population: postings read directly from an employer's applicant-tracking system, over the 14-day ingest window the whole latency cut was taken from.
Derivation: Median of `created_at - posted_at` over the ATS-direct cohort of a 14-day ingest window (n=2,441,587 across all cohorts). p95 on the same cohort is 5.7 days. 7.94% of corpus rows carry `posted_at` stamped equal to `created_at` at ingest (signalsapi-4450), which reads as zero latency; the cut is not filtered for them, so this is a floor rather than a point estimate.
Computed by: The 2026-08-08 measurement pack (`docs/gtm-cofounder/founder-brief.md`, §Measured facts, finding 2), cut directly against the corpus database upstream. This repo holds no read path that can reproduce it.
Expires 2026-11-06.
9 hours
Population: every non-synthetic posting in the same 14-day ingest window.
Derivation: Median of `created_at - posted_at` over the whole window (n=2,441,587); 9.0 hours unrounded. The class of rows with a trustworthy timestamp is 92.3% `li_jobs`, so this figure is substantially a statement about one source. 7.94% of corpus rows carry `posted_at` stamped equal to `created_at` at ingest (signalsapi-4450), which reads as zero latency; the cut is not filtered for them, so this is a floor rather than a point estimate.
Computed by: The 2026-08-08 measurement pack (`docs/gtm-cofounder/founder-brief.md`, §Measured facts, finding 2), cut directly against the corpus database upstream. This repo holds no read path that can reproduce it.
Expires 2026-11-06.
48 hours
Population: every non-synthetic posting in the same 14-day ingest window.
Derivation: 95th percentile of `created_at - posted_at` over the same population as the corpus-wide median above; 48.0 hours unrounded. 7.94% of corpus rows carry `posted_at` stamped equal to `created_at` at ingest (signalsapi-4450), which reads as zero latency; the cut is not filtered for them, so this is a floor rather than a point estimate.
Computed by: The 2026-08-08 measurement pack (`docs/gtm-cofounder/founder-brief.md`, §Measured facts, finding 2), cut directly against the corpus database upstream. This repo holds no read path that can reproduce it.
Expires 2026-11-06.
12 min
Population: postings whose source is Lever, in the same 14-day ingest window.
Derivation: Median of `created_at - posted_at` restricted to the Lever source; 12.1 minutes unrounded. 7.94% of corpus rows carry `posted_at` stamped equal to `created_at` at ingest (signalsapi-4450), which reads as zero latency; the cut is not filtered for them, so this is a floor rather than a point estimate.
Computed by: The 2026-08-08 measurement pack (`docs/gtm-cofounder/founder-brief.md`, §Measured facts, finding 2), cut directly against the corpus database upstream. This repo holds no read path that can reproduce it.
Expires 2026-11-06.
1.7 h
Population: postings whose source is Greenhouse, in the same 14-day ingest window.
Derivation: Median of `created_at - posted_at` restricted to the Greenhouse source. The unrounded figure is 99.3 minutes; the published form is its rounding to hours. 7.94% of corpus rows carry `posted_at` stamped equal to `created_at` at ingest (signalsapi-4450), which reads as zero latency; the cut is not filtered for them, so this is a floor rather than a point estimate.
Computed by: The 2026-08-08 measurement pack (`docs/gtm-cofounder/founder-brief.md`, §Measured facts, finding 2), cut directly against the corpus database upstream. This repo holds no read path that can reproduce it.
Expires 2026-11-06.
3.07 days
Population: postings whose source is Workday, in the same 14-day ingest window.
Derivation: Median of `created_at - posted_at` restricted to the Workday source, with 0.05% of that cohort inside an hour. Registered although no page quotes it: it is two orders of magnitude off the two vendors that are quoted, so it is what makes a pooled median under a sentence naming only the fast vendors a mixing error rather than a simplification. 7.94% of corpus rows carry `posted_at` stamped equal to `created_at` at ingest (signalsapi-4450), which reads as zero latency; the cut is not filtered for them, so this is a floor rather than a point estimate.
Computed by: The 2026-08-08 measurement pack (`docs/gtm-cofounder/founder-brief.md`, §Measured facts, finding 2), cut directly against the corpus database upstream. This repo holds no read path that can reproduce it.
Expires 2026-11-06.
6.29%
Population: live non-synthetic requisitions in the corpus on the measurement date.
Derivation: 44,125 of 700,999 live non-synthetic reqs name an applicant-tracking system as their source; the other 93.7% carry a job-board or aggregator URL. This is the figure behind every statement about where rows come from, and the reason no page may promise a source-ATS apply link.
Computed by: The 2026-08-08 measurement pack (`docs/gtm-cofounder/founder-brief.md`, §Measured facts, finding 1), cut directly against the corpus database upstream. This repo holds no read path that can reproduce it.
Expires 2026-11-06.
16.7%
Population: a stratified sample of 500 posting URLs the corpus still carries in its 14-day frame, 250 per stratum, drawn on a deterministic seed against the corpus rebuilt on 2026-08-18 — a different population from the 2026-08-08 reading this supersedes, which was taken over a 31-day frame that no longer exists.
Derivation: Re-fetch of each sampled URL, weighted back to corpus volume: 18.1% of 232 decidable ATS-direct URLs and 5.8% of 243 decidable aggregator URLs were dead or expired, giving a volume-weighted 16.7% (95% CI 12.3-21.2%) over a 14-day frame. This is not a like-for-like refresh of the 49.8% (95% CI 37.7-61.8%) read on 2026-08-08 and must not be published as an improvement: the corpus was rebuilt, its oldest record is 15 days old, so the 31-day frame the earlier figure was taken on returns nothing, and the rate is strongly age-dependent — signalsapi-4458 measured ATS-direct at 4.2%, 14.2%, 31.5% and 45.4% at 3, 7, 14 and 31 days. Most of the fall is the shorter frame. Two things the earlier entry said no longer hold. The strata are no longer indistinguishable — their confidence intervals are disjoint at both 7 and 14 days (14-day ATS-direct 13.7-23.6% against aggregator 3.5-9.4%) — so the earlier reading's conclusion that scoping a liveness claim to ATS-direct sources buys nothing is withdrawn, and the sign is the opposite of what it implied: ATS-direct is the worse stratum, and the aggregator stratum is the one a scoped claim would favour. And the ATS-direct stratum had to be framed as `lifecycle_status IS NULL` rather than active, because the lifecycle re-check has never run against that cohort at all (signalsapi-5032). Supersedes the 37.5% (n=48) figure the front door was never permitted to publish. This is the measured size of the gap behind the PRE_REGISTERED Ghost Job metric above — it is not a figure published under that definition, which stays blocked on a retention surface longer than its own 90-day window.
Computed by: signalsapi-5031, a re-run of `signals/scripts/measure_ghost_rate.py` on seed 5031 drawn 2026-09-02, against the plane. It replaces the 2026-08-08 sweep of signalsapi-4453 (closed), whose design cannot be re-run on today's corpus. This repo holds no read path that can reproduce it.
Expires 2026-12-01.
Some studies count postings by matching their titles against a fixed list of terms. Those lists are published here in full and dated, so the window they define cannot be widened or narrowed after a result is seen without that change appearing in the changelog below.
Terms are matched against a posting's job title, case-insensitively and on a word boundary — so “llm” matches “LLM Engineer” and not “fulfillment”. They are not matched against the description of the posting: the corpus surface these pages are built from carries a posting's title, company, location and salary but not its body text, so a posting whose body mentions a term under an unrelated title is not counted. Every list below therefore understates its term's true incidence, in the same direction and for the same reason.
Titles that name the model-building work itself. Deliberately excludes “ai” alone, which matches unrelated titles on a word boundary, and the vendor and framework names that trend and fade faster than a monthly study's cadence can honestly track.
Titles for building and moving the data the models are trained and served on — the pipeline half of the same team, kept separate so growth in one is not read as growth in the other.
Titles for analysis and experimentation rather than production systems. Narrow on purpose: broadening it toward “analyst” would absorb business reporting roles and inflate the cut.
Titles for the runtime the above is deployed onto. Included as the comparison set: it is the specialism whose demand would have to fall if model work were displacing infrastructure work rather than adding to it.
Titles for securing the systems above. A second comparison set, and the one least coupled to model adoption — a control against reading a market-wide hiring swing as a specialism-specific one.
Some studies rank demand across sectors. The sectors are published here in full and dated, so a taxonomy cannot be redrawn around a result after the result is seen without that change appearing in the changelog below.
Industry labelling covers 67.22% of the deduplicated postings these pages render (measured 2026-07-22). That coverage is not evenly distributed: it is 88.81% on our largest single source and 31.55% across every other source combined, and two national corpora carry almost no industry labels at all. So counting postings per industry and ranking the industries would largely rank which sectors our best-labelled source happens to cover. Any figure we publish by industry is therefore a rate measured inside a sector — the same labelled population in the numerator and the denominator — never a share of the whole corpus, and it excludes every posting whose employer carries no industry label.
Industry groups are matched against the industry label the source itself publishes for the employer, case-insensitively and as a substring — so “software” matches “Software Development”. Substring matching is looser than the word-boundary rule used for job titles, which is why every term below is a phrase specific enough that no unrelated label contains it. A label may match more than one group; each group is measured against its own postings, so an employer counted in two groups is counted correctly in both rather than twice in one.
The sector that builds software as its product. Expected to lead any engineering-demand ranking, and included precisely so that expectation is measured against the others rather than assumed.
Regulated, data-heavy, and an early adopter of model-driven work outside the technology sector — the clearest test of whether AI hiring has spread beyond employers whose product is software.
Research-intensive and slow-moving on hiring, so a rate that rises here is harder to explain as a labelling artefact than one that rises in software.
The largest employer group in several of the national corpora we read, and the one whose postings least often reach an English-language ATS — a control against reading source coverage as sector demand.
High posting volume, low engineering density. Included as the low end of the expected range: a ranking with no low end is a ranking whose scale a reader cannot judge.
Recommendation and content-generation work sits here, so it is where applied model work would appear first outside the technology sector proper.
Public-sector and academic postings follow a hiring calendar rather than a market, so this group is expected to move differently from the rest and is kept separate rather than folded into professional services.
Firms that sell expertise rather than a product. Their postings are often for client projects, which is stated here because it is the group whose demand is least attributable to the employer's own operations.
v1.0.0 — 2026-07-18
Initial publication. Sample definition, deduplication rule and vintage window stated; three metric definitions pre-registered. No figure is published under any of them yet.
Publishes no figure — no review quorum required.
v1.1.0 — 2026-07-20
Pre-registered a fourth metric, Cross-Source Lead Time, ahead of the status page that reports on it. No figure is published under it; the entry exists so the figure cannot be published without passing this page's gates first.
Publishes no figure — no review quorum required.
v1.3.0 — 2026-07-22
Pre-registered the Hiring-Demand Index and fixed its base to the first edition published under this definition. The index was originally specified against a January 2026 base; the render surface retains only a rolling recent window (stated below), so no observation of that month exists or can be recovered, and back-casting one would be the invented vintage this policy forbids. No figure is published under the index; its observations begin accumulating with the first edition.
Publishes no figure — no review quorum required.
v1.4.0 — 2026-07-22
Pre-registered the five term lists the AI-engineering demand study is cut from, and stated how they are matched. The study has published counts under these exact terms since it launched, but the terms themselves lived in the generator where no reader could see them and no date fixed them — a figure published under a definition this page had not published. The lists are unchanged; what changed is that they are now on this page, dated, and are the same object the query is compiled from, so editing one is a revision of this page. The match rule is also stated for the first time, including the part that narrows it: terms are matched against a posting's title, not its description, which we do not hold.
Publishes no figure — no review quorum required.
v1.5.0 — 2026-07-22
Pre-registered how the Hiring-Demand Index will be graded against official vacancy statistics — the error measure, the reference-month alignment, the lead-lag search, and the rule that a comparison is made against the official series' first published print rather than a later revision. No comparison has been made and no figure is published under either definition; the accrual toward the first one is published at /labs/scorecard. Registering the protocol while nothing can be compared is the point: a scorecard whose rules were written after the results were in is a rationalisation, not a scorecard.
Publishes no figure — no review quorum required.
v1.7.0 — 2026-07-22
Stated which of our own series the accuracy comparison grades, and corrected the reason it is not yet published. The scorecard is graded against a geography-matched sub-series counted on the same slot, window and deduplication rule as the global index and differing from it only in a country filter — never against the global index itself, which spans every market the corpus covers and could not be compared like for like with a national statistic. One such series is now recorded per benchmark, and the stated blocker no longer says none exists. No definition changed and no figure is published under either scorecard metric: what changed is that the comparable side is being recorded, which cannot be back-filled and so had to start before there was anything to compare it to.
Publishes no figure — no review quorum required.
v1.8.0 — 2026-07-22
Pre-registered the eight-group industry taxonomy the AI-engineering demand study is ranked across, stated how industry terms are matched (a substring of the source's own label, which is looser than the word-boundary rule used for job titles), and published the measured coverage of that label. The coverage measurement is the reason the taxonomy is being registered the way it is: industry labelling reaches 67.22% of the postings we render, but 88.81% on our largest source against 31.55% across all others, so a ranking of postings-per-sector would largely rank which sectors that source covers. Any figure published by industry is therefore defined here as a rate measured inside a sector, never a share of the whole corpus. No figure is published under the taxonomy yet; the first ranking appears at the next monthly edition, after these definitions.
Publishes no figure — no review quorum required.
v1.9.0 — 2026-08-10
Corrected a factual claim this page has carried since v1.4.0 and widened the field list on /data-ethics from 13 fields to the full served set. Both had the same cause: the pages generated their statement of what the product holds from the projection these pages render, which is not what the product holds. The API serves a posting's full advertised body text, so v1.4.0's parenthetical — that a posting's description is something "we do not hold" — was wrong, and is superseded by this entry rather than edited out of it. What is unchanged is the method it justified: term lists are still matched against titles alone, because the corpus surface these pages read carries no body text, so every count still understates its term's true incidence in the direction already stated. No figure changes, and no figure is published under a new definition.
Publishes no figure — no review quorum required.
v1.10.0 — 2026-08-10
Registered the dated measurements this site publishes outside the metric definitions above, and gave each one an expiry. The registry until now held only ongoing metrics — a definition plus a status — so a figure that is a one-off reading of a moving corpus had nowhere to live and was written as a literal into the copy that quoted it. That is how "refreshed every 3 hours" came to ship in 24 places while being true of nothing; the 2026-08-08 detection-latency and corpus-composition readings were on the same path. Each is now published below with the population it describes, the derivation that produced it, the date it was taken and the date it expires, and the build fails once one is past its expiry. No figure changes value, and none is published under a metric definition.
Publishes no figure — no review quorum required.
v1.11.0 — 2026-09-01
Renamed the index and the two scorecard metrics graded on it, dropping the product name from all three: they are now the Hiring-Demand Index, its Accuracy vs Official Vacancy Statistics, and its Lead Over Official Vacancy Statistics. Nothing about any definition changed — the sample, the window, the base, the alignment rule and the blockers are the same objects with the same dates, and the earlier entries above name them by their current spelling so the record reads as one metric rather than two. The reason is that these two pages are now rendered for two hosts under two product names, and a pre-registered name is the handle a citation quotes: a name that reads differently depending on which host it was fetched from is not frozen, and a figure published under one spelling could not be checked against a scorecard graded under the other. A brand-free name is frozen once and true on both. No figure is published under any of the three, so no published figure changes its governing definition, and no figure changes value.
Publishes no figure — no review quorum required.
v1.12.0 — 2026-09-02
Re-measured the sampled ghost rate and withdrew a conclusion the earlier reading carried. This is the first entry under which a published figure changes value, so it is written as a supersession rather than an edit: the 2026-08-08 reading was 49.8% volume-weighted over a 31-day frame, and the 2026-09-02 re-run is 16.7% over a 14-day one. The two are different quantities, not a before and after. The corpus was rebuilt on 2026-08-18, its oldest record is 15 days old, and the 31-day frame the first figure was taken on now returns nothing at all — while the rate itself rises steeply with age, so most of the fall is the shorter window rather than a cleaner corpus. The entry says so in its own cohort and derivation, because a reader who meets the smaller number without that sentence reads an improvement we did not measure. Withdrawn with it: the earlier reading's finding that the two source strata were statistically indistinguishable, and therefore that scoping a liveness claim to ATS-direct sources buys nothing. On this cut the intervals are disjoint and the sign is inverted — postings read directly from an employer's applicant-tracking system are the worse stratum, not the safer one. No metric definition changed and no metric moved off PRE_REGISTERED; what changed is one dated reading and the claim that rested on it.
Publishes no figure — no review quorum required.
See also Data ethics — what we hold, and what we will not publish.