Method
What four new polls did to the forecast
We thought we had nearly doubled the 2026 poll count. A third of the growth turned out to be the same polls twice. Here is what the real ones moved, and how the duplicates were caught.
This week the 2026 poll corpus went from 191 polls to 243, run by nine pollsters. Two things caused it: four genuinely new polls filed between 17 and 19 August, and a second source we had not been reading — the Wikipedia compilation for this cycle, which turns out to publish a sample-size column no earlier page has.
The count read 290 for a day. Forty-seven of those were polls we already had, and the section below is about how they got in.
The honest question is what the real ones bought. Here is the whole of it.
| 191 polls | 243 polls | |
|---|---|---|
| Mean 90% interval | 7.67 seats | 7.44 seats |
| Likud | 24.25 | 23.82 |
| Netanyahu bloc | 53.1 | 52.6 |
| Netanyahu P(≥61) | 7.1% | 6.1% |
| Opposition P(≥61) | 17.6% | 15.7% |
Intervals tightened by 3%. Likud lost about half a seat. The Netanyahu bloc’s path to a majority narrowed from 7.1% to 6.1%.
Why this is not a bigger number
More data tightens a forecast roughly as the square root of how much you add, and only when the new data is independent of what you already had. A good deal of the Wikipedia material describes the same fieldwork the compilation already recorded, so the effective gain is smaller than “99 more polls” suggests.
It is also worth saying what did not change: nothing at all in the backtest. The five settled elections gained no polls, so every historical score is identical to six decimal places. New data landed entirely on an election that has not happened yet, which means we cannot score it until 27 October. Anyone telling you that more data has made their model measurably better, without a held-out election to prove it on, is telling you about their intentions.
Written on 21 August 2026, and overtaken the next day. The corpus went back to February 2009, which added three elections rather than more polls for the ones already in it — and that did move the historical scores, because the error model is fitted on the elections that came before each forecast. Point accuracy is unchanged; CRPS, the interval widths and the bloc probabilities are not. The paragraph above is right about the four polls it was written about and wrong as a general claim.
The forty-seven polls we counted twice
Reading a second source is how you find out whether your idea of “the same poll” survives contact with someone else’s spelling. Ours did not.
A poll’s identity here is deliberately not its publication: it is the pollster, the fieldwork window, the commissioner and the ballot scenario, so that two outlets reporting one sample collapse to one record. The rule was sound and the input was not — it used the pollster’s name exactly as the source wrote it. The compilation writes Midgam R&C; the Wikipedia table writes Midgam. Same firm, same fieldwork, same Channel 12, same 501 respondents, same seats for every list — and two records, each carrying full weight.
Twenty-six of Channel 12’s weekly polls were in the corpus twice, and eight more of Maagar Mohot’s under a transliteration. The effect was small but it was in a specific direction: the opposition bloc’s chance of reaching 61 read 19.1% when the duplicates were counted and 15.7% when they were not.
The fix is that identity now resolves the pollster through the alias table before hashing, so the two spellings produce one id. What it must not do is merge polls that only look alike: April 2019’s page has no commissioner column and puts the outlet inside the pollster cell, where “Midgam/ Yedioth Ahronoth” and “Midgam/Channel 12” are one firm and two different polls. That distinction is read back out rather than discarded.
The alias table cannot reach every case. It resolves two spellings of one firm; it cannot resolve two sources that credit one poll to different firms, and they do. Channel 13’s weekly poll arrives as Midgam Project from one source and HaMadad from the other, on the same ten dates with the same numbers. Left alone, that channel carried 21% of the forecast’s weight against the 12% it had earned.
So the last rule takes identity from the numbers and needs no view about who runs what: same election, same fieldwork window, same commissioner, and the same seats for every list that won any. Two firms polling one outlet on the same days to identical results across fifteen lists is not a thing that happens. It requires a fieldwork window to fire, because an absence is not a window and two undated polls sharing a number are a guess rather than a deduction.
What that rule deliberately does not settle is whose poll it is. The surviving record keeps the first source’s attribution, and whether Channel 13’s pollster should be called Midgam Project or HaMadad is still an open question for someone who knows.
The correction we nearly missed
The compilation does not only grow — it revises. Sixty-two lines of what we had already committed were rewritten, including one Filber poll from 25 June that had been recorded with only 115 of 120 seats allocated and now carries the missing Ra’am row.
That is why every snapshot we hold is dated rather than overwritten. Had we simply refetched the file, the forecast would have silently changed and the version any earlier forecast was built from would be gone.