Audit of our own data
How we checked
Before showing this data to anyone outside, we tested it against the job boards themselves. We drew a random sample, published it before looking at a single answer, and then checked every posting by hand-built code that shares nothing with our collector.
This page reports what we found, including the three things we got wrong.
Audit run 1 September 2026. Every figure below is frozen at that date — this is a report on a check, not a live counter.
What did we actually check?
We drew 375 job postings at random: 200 that our data said had been open 90 days or longer, 100 that our data said had been removed, and 25 each from three smaller platforms that a purely proportional sample would have almost missed — a flat draw of 200 landed only four Lever postings and no Ashby postings at all.
For each one we visited the posting's own page on the employer's job board and read the date the board publishes, then compared it to ours to the day.
The reader that did this is a second implementation. It shares no code with the collector that gathers our data, reaches the boards by a different route, and parses a different document. Otherwise the check would only be our collector agreeing with itself.
Why did we publish the sample before checking it?
A recorded random seed proves a sample can be reproduced. It does not prove it was drawn honestly — anyone can try seeds until one flatters them, then publish the winning seed, and the result is indistinguishable from an honest draw.
So we drew the sample, committed it publicly with every result column empty, and only then began checking. The public record shows the list existed before any posting in it had an answer. The seed was zeltru-gt-2026-09-01 and the draw is recomputable by anyone.
What was the match rate?
Of the 200 long-open postings, 188 could be read from the board at all. 181 matched exactly — 96.3% (95% confidence interval 92.5–98.2%). There were no near-misses: every disagreement was larger than three days, so these are not rounding or time-zone errors.
| Open 90+ days (200 sampled) | Count |
|---|---|
| Matched the board exactly | 181 |
| Could not be read | 12 |
| Did not match | 7 |
Of the 100 postings our data said had been removed, 92 gave a clear answer: 85 were genuinely gone, and 7 were still live. That is 7.6% (95% confidence interval 3.7–14.9%) — roughly one in fourteen postings we mark as removed is still up. It is a real error rate on our side and we would rather state it than round it away.
We use a specific definition here, because a loose one would have flattered us. A posting is removed when it no longer appears in the board's own listing — not when its web address stops working. Boards routinely keep a job's direct link serving long after the job leaves the listing, so judging by the link alone would let us count delisted jobs as removed and never notice.
So we checked all seven against the listing rather than the link. All seven were in it. They are live postings we wrongly marked removed, and the rate stands.
What that rate implies at scale, and we would rather do this arithmetic ourselves than have a reader do it: applied naively to the 185,659 removals in our most recent published edition, 7.6% is roughly 14,100 postings counted as removed that were not — somewhere between 6,900 and 27,700 across the confidence interval. The interval is that wide because the check rests on 92 decidable postings, not because the estimate is careless.
It also nudges our published median. A posting we wrongly mark as removed enters the statistics with its life cut short, and six of the seven had a true age exceeding what we recorded by 3 to 36 days. Their recorded durations sit just below the 21-day median, so correcting them would move that median up — we understate slightly. By how much we genuinely cannot say from seven cases; one to two days is the most the arithmetic supports as a ceiling.
We could not explain it, and we looked. We had a plausible culprit: a defect where an employer with two job boards could have each board retire the other's postings, fixed on 31 August. All seven removals predate that fix, which on the dates alone would have let us file this under a bug we had already closed. The mechanism does not allow it — it needs an employer with more than one board, exactly one such employer exists, and none of these seven is it. So the rate is unexplained, and we are publishing it as unexplained rather than attributing it to something convenient.
On the three smaller platforms, every posting we could read matched: Lever 25 of 25, Ashby 24 of 24, Greenhouse 14 of 14. These are reported on their own and are deliberately not blended into the headline figure — mixing an oversample into an average would quietly distort it.
What did the misses look like?
All seven pointed the same way. The board showed a date newer than ours, by 95 to 227 days. Not one pointed the other way.
| Our date | Board's date | Difference |
|---|---|---|
| 13 Jan 2026 | 28 Aug 2026 | +227 days |
| 2 Feb 2026 | 29 Aug 2026 | +208 days |
| 4 May 2026 | 27 Aug 2026 | +115 days |
| 19 May 2026 | 1 Sep 2026 | +105 days |
| 14 May 2026 | 24 Aug 2026 | +102 days |
| 5 May 2026 | 12 Aug 2026 | +99 days |
| 28 May 2026 | 31 Aug 2026 | +95 days |
The cause is a deliberate rule in our collector: once we record a posting's date, we never overwrite it. That rule protects against a board handing us a worse date later. But when an employer re-dates a posting, we keep the older date — so the posting stays in our long-open counts carrying a date the board itself no longer claims.
This means we overstate the age of those postings, not understate it. At seven in 188, that is about 3.7% of the long-open population.
The seven postings we wrongly marked as removed have a separate cause, and we traced five of them. Those five sit on job boards that will only show us the first 2,000 postings regardless of the board's full inventory. A posting outside the served window can be missed repeatedly. Individual audit examples are withheld; the measured count of five cases and the mechanism remain part of the aggregate finding.
The other two we could not explain. Their boards did not share the enumeration limit, so those two cases remain unexplained rather than folded in with the five. Individual record details are withheld.
A warning about our removal figures for this period
This one is not an error we found in a sample. It is a property of the period itself, and it is why the figures above lead with the cautious reading.
Between late July and the end of August, the stored records for a large part of our dataset were replaced — rewritten in place across two days, 10 and 11 August. This touched 223 employers, holding 23.2% of the postings we track, and 109,739 records were written carrying history from before the record itself existed — in some cases nineteen days of it.
So the removal accounting for this window rests on records that were rewritten during it. 58.9% of the removals we recorded in the period are at those employers, against their 23.2% share of postings — two and a half times their weight. A quarter of our data produced well over half the removals we are reporting.
We know what did it, and it was us. When we published this warning we said we could not establish the cause. We can now: a data-cleanup job of ours. We had been storing some job postings under two different internal identities, and in August we merged each pair into one record, keeping the earlier of the two first-seen dates so no history would be lost.
That merge was deliberate, careful, and correct on its own terms. What nobody connected at the time is that rewriting the provenance of 100,000-odd records changes what every statistic computed over that period means. The cleanup did not corrupt anything; it made the removal accounting for those weeks unreadable as a description of employer behaviour, and it did so silently.
So the conclusion is unchanged and better founded: removal figures for this window cannot be read as a clean description of what employers did.
We are not publishing a corrected count. We have no way to separate the affected removals from ordinary ones, and a number that looks cleaned up without being cleaned up is worse than the warning.
Two things we can say precisely. The rebuilds kept the underlying observation history rather than rewriting it — records dated in July were written in July, and nothing was deleted — so what was damaged is the accounting, not the observations. And we now check for this every night: a daily tripwire watches for records being written with history they did not earn, so a third rebuild day cannot pass unnoticed the way the first two did.
Two of the claims above are better established than the third, and we would rather say which. That it was those two days and no others was worked out twice over, by two methods with nothing in common — one reading when each record was written, the other reading the activity log and the server logs from the night itself. Neither could inherit the other's mistake, and they agree.
The employer count and the percentages — 223, 23.2%, 58.9% — rest on the first method alone. Nobody re-derived them a second way. They are our best measurement and we believe them; they simply do not carry the same weight as the part that was checked twice, and presenting the whole thing as equally confirmed would overstate what we did.
This warning has been corrected twice, and the sequence is worth stating. We first described a single day affecting 272 employers. We then corrected that to five days and 570 employers — and the correction was wrong, because the test we used also caught an innocent case: a nightly job that runs past midnight makes a record look one day older than itself. Building the automatic check forced us to define it properly, and the answer came back to two days and 223 employers. Our first figures were approximately right, our correction made them worse, and the tool we built to watch for the next occurrence is what settled it.
Why we are not publishing anything about re-posting
When a posting we had marked as gone turns up again, our records call it a reappearance. It is tempting to read those as employers re-posting the same job, and we are not going to.
We have 10,123 of them. The typical gap between a posting being marked gone and turning up again is four days; 85.6% are inside a week, and not one exceeds thirty days. Marking a posting gone already takes three days by design. An employer re-posting a job weeks later would leave a long tail in that distribution, and there is none — the shape is what our own detection flickering produces.
Three specific reasons it cannot carry a claim about employer behaviour:
One. Our records only span 40 days, so a gap longer than that could not appear even if it were common.
Two. Stored records for much of the dataset were replaced mid-period (see the warning above). A posting whose record is replaced cannot reappear under its old identity — it enters as new instead. 77% of these reappearances are at affected employers. The underlying observations survived the replacement intact, so this is a bounded problem rather than a fabricated one, but it is not nothing.
Three. The 2,000-posting board limit described earlier manufactures reappearances outright: a posting past the limit is retired without going anywhere, then reappears when the board's window shifts. No employer did anything. Around 8,800 postings currently sit past that limit.
We publish the distribution as an observation and stop there.
What did we get wrong?
Three things, all found during this audit, all corrected here.
One. We had reported that four of the five job-board platforms never change a posting's date — zero changes across 136,169 postings. That was not a fact about those platforms. It was a description of our own code, which never updates a date once it has one. If every one of those boards re-dated every posting nightly, our data would still have shown zero. We have since instrumented the collector to record the disagreement instead of discarding it.
Two. We had concluded that our stale dates would cause us to undercount long-open postings. The seven misses above show the opposite, for the reason given in the previous section. We had the direction of our own error backwards.
Three. We found a date change on one platform by picking a posting to look at, and reported it. It did not reappear in 24 random draws from that platform. A result from a chosen example is not a rate, and we should not have written it as one.
What are the limits of this check?
Our reader is independent in route, not in source. It reaches each board differently from our collector and parses a different document, but on every platform both ultimately read the same underlying field. So this check catches our own staleness, storage and parsing errors — which is what it found — and it cannot certify that a platform's idea of “posted” is true.
No job board shows a posted date to a human being. We looked. One platform's posting page contains the word “Posted” exactly once in 731,567 characters, and not as a date. These dates exist as metadata for search engines, not as something an applicant sees on the page.
About 28% of our 90-days-or-longer count sits in bulk-dated groups. We call a posting bulk-dated when 50 or more postings from the same employer share one identical date. 76 such groups exist; the largest is 9,363 postings on a single date. Thousands of jobs are not posted on one day and then each stay open the same number of days — that is one bulk upload carrying one date. The boards genuinely publish these dates and we report them faithfully, but the figure deserves to be stated both ways.
The threshold of 50 is a judgement call, so here is how little it matters.
| Group size | Excluded | Postings remaining |
|---|---|---|
| 25 or more | 32.1% | 68,671 |
| 50 or more (adopted) | 28.2% | 72,577 |
| 100 or more | 25.9% | 74,889 |
| 500 or more | 19.4% | 81,434 |
There is no natural break in that curve, which is the useful part: a quarter to a third of the long-open tail is bulk-dated whichever line you draw. We adopted 50 because ten postings on one day is ordinary behaviour, 500 misses obvious bulk uploads, and at 50 there are few enough groups that every one can be inspected by hand.
So the long-open figure is stated as a pair, never as one number. Measured on 1 September 2026: 101,077 postings open 90 days or longer as measured, or 72,577 excluding bulk-dated groups — a 28.2% reduction. From our next edition on, every 90-days-or-longer figure is published both ways, on the same line.
What have we not checked yet?
We asked the Internet Archive about our 60 oldest postings. Nothing contradicted us, and nothing corroborated the dates either. 28 had captures, 20 had none, and 12 could not be checked because the archive was returning errors that day. Where captures existed, the earliest was a median of 2,264 days — more than six years — after the date we claim. So the archive confirms these postings existed recently. It says nothing about when they were first posted, and no capture predated any of our dates.
We are recording that as a null result rather than as support. For the oldest rows a contradiction would need a capture from before the claimed date, and for a 2009 claim that is close to impossible to find — so —not contradicted— here is far weaker than it sounds, and we would rather say so than bank it.
Our single oldest active posting is dated December 2009 — three years before the job-board platform that issued it existed — so it is flagged and excluded rather than displayed.
A re-check of the two platforms where our collector never re-reads a posting's date once it has one. Those cover most of our data, and the sample above is currently the only instrument that can see errors there.
A review by an outside labour economist or data journalist. When that happens, their questions and our answers go on this page — including the ones we would rather they had not asked.
If you find something here that looks wrong, we want to know. See our methodology for how the numbers are built in the first place, or get in touch.