Guide

Job posting data: sources, freshness and deduplication

The data layer under every hiring signal. Where it comes from, how it goes bad, and the four questions that expose a weak provider.

Where job posting data comes from

Every job posting dataset starts from the same public surfaces. There are three kinds, and each behaves differently.

Job boards

Boards are where companies pay to advertise a role. They are broad and easy to observe, but noisy. The same role appears on several boards at once. Paid postings can linger after the role is filled, because the advertiser bought a listing period, not a live mirror of the req.

ATS career pages

Most companies run hiring through an applicant tracking system, and most ATS products publish a public careers page or board tenant. These pages sit closest to the source of truth. The company maintains them as part of its own hiring workflow, so a posting that disappears there usually means the req closed. A link like boards.greenhouse.io/… is an ATS link. It is also the best link to put in an email: public, direct, and no login wall.

Aggregators

Aggregators collect postings from boards and career pages and republish them. They add reach, and they add lag and duplication. A posting can take time to surface on an aggregator, and it can remain there after the source removed it.

The three surfaces disagree with each other all the time. A role can be live on an ATS page, expired on one board and still cached on an aggregator. Any dataset that merges them without recording the source cannot resolve those conflicts, and you inherit the confusion.

A serious provider observes a defined set of these surfaces and tells you which ones. In Reqbeat, every result is a real posting tagged with the board it was detected on, so each row names its own source and conflicts stay visible.

Why freshness decays fast

Job posting data ages worse than most B2B data. A company's name and domain hold for years. A posting is a live object. The company edits it, pauses it, fills the role and removes it. Each dataset is a photograph of a moving scene, and the scene keeps moving after the shutter clicks.

For outreach, the decay is sharper still. The value of a hiring signal is timing: budget was just committed, and the pain was just written down. The hiring signals guide covers why. A req you learn about weeks later is a fact, not a trigger. Worse, outreach built on a closed posting reads as spam. "Saw you're hiring" about a role that was filled last month teaches the prospect to ignore you.

Decay hurts aggregates too, not just single rows. A count of "open reqs" computed on stale data mixes live roles with filled ones. Any velocity or trend built on that count drifts away from reality in the same direction. Fresh inputs are the floor under every derived number.

So the first property to check in any dataset is the gap between the posting going live and the row reaching you. A low price does not compensate for a long gap. The gap decides whether the data supports a timing play at all, or only market research.

Why duplicates and repostings corrupt counts

A requisition rarely appears once. The company publishes it on its ATS page, syndicates it to boards, and aggregators copy it again. One role becomes many rows.

Reposting adds a second layer. Advertisers refresh a posting to restart its board ranking, or relist a role that did not fill. The old req reappears with a new date and looks like a new signal.

Left uncorrected, the damage spreads through everything built on the data:

  • Counts inflate. "Open reqs" overstates real hiring, and every score built on the count inherits the error.
  • Trends lie. A reposting wave looks like hiring growth. Velocity and direction computed on raw rows measure advertising behavior, not headcount plans.
  • Outreach repeats. Two rows for one role means two emails to the same hiring manager. Nothing burns a domain faster than looking automated.

Deduplication is not a nice-to-have. It decides whether the numbers mean anything. And the rule must be explicit, because any dedup rule makes trade-offs you need to know about.

Four questions to ask any provider

Before you buy job posting data, or build on a free source, get written answers to four questions. A good provider answers all four on a public page.

1. What is the sample?

Which surfaces do you observe: which ATS products, which boards, which aggregators? Is LinkedIn included? What segments does the sample miss, and what known biases does that create? "We cover millions of postings" is not an answer. A named sample with named gaps is.

2. What is the refresh cadence?

How often is each source re-observed? Is delivery push or poll? What is the typical gap between a posting going live and the row reaching my endpoint? Cadence decides whether the data can drive a same-day play.

3. What is the dedup rule?

What is the exact partition that collapses duplicates? When two rows collapse, which one survives? How are repostings handled, and do they fire events? Vague answers here mean the counts are soft.

4. What does absence mean?

If a company returns no postings, does that mean "not hiring" or "not covered"? The two are different facts, and a provider that cannot tell them apart will poison any scoring model you build. Absence of coverage must be labeled, not silent.

One follow-up applies to all four: where is this written down? An answer that lives in a sales deck can change after you sign. An answer published on a method page is a commitment you can hold the provider to. Prefer providers that publish.

How Reqbeat answers each one

Sample

The data comes from public job boards and ATS pages we observe directly. LinkedIn is not included: API results cover ATS tenants plus public job boards and aggregators, so every posting is one you can link in an email without sending your prospect to a login wall. Each row's boards field names its source. The method page states the sample and the known biases, and data ethics states what we hold and what we do not publish.

Refresh cadence

The corpus refreshes every 3 hours across 7 regions. Delivery is push: a webhook fires when a new matching posting is detected, typically hours after it goes live. There is no poll gap to add on top.

Dedup rule

Repeat and re-posted advertisements collapse to one row per company, job title and country. The most complete row is kept. The exact partition rule is published on the method page. Billing follows the same logic: a watch bills only when your endpoint accepts a delivery, and duplicates cost nothing.

Coverage semantics

A company with no ATS or board coverage answers with coverage_status: "no_ats_signal". That means "unknown", not "not hiring". Payloads say so explicitly, so your scoring model never mistakes a blind spot for a negative signal.

You can check all of this against live data in minutes. Sign up, the API key is on screen, no card, and the quickstart starts with the first call. The API docs hold the full field reference.

Keep reading

Ask us the four questions. The API answers most of them on the free tier.