Ofqual's "comparable outcomes" approach exists specifically to prevent competition between exam boards turning into a race to lower standards. Grade boundaries in England aren't fixed at a set mark year to year — they're adjusted using statistical evidence about how each cohort has performed in the past, so that a grade from AQA and the same grade from Pearson Edexcel are intended to represent the same standard. Whether this fully removes every incentive a board might have to be perceived as "easier" is genuinely debated, and this article won't claim more certainty than the evidence supports.

This is one of the oldest anxieties parents raise about the English system: if four different companies are all competing for the same schools' business, doesn't at least one of them have an incentive to make their exam a bit easier, to win market share? It's a fair question to ask. It's also one Ofqual was specifically built to answer — the mechanism is worth understanding on its own terms before drawing a conclusion.

The concern, stated plainly

Schools choose which board to use for each subject. If grades were simply a matter of "how many marks out of the total paper," a board whose exam consistently produced higher grades for the same underlying ability might attract more schools — not because its teaching or content is better, but because its outcomes look better on a school's results table. Left unchecked, that dynamic could, in theory, pull standards down across all boards as each competes not to be the "harder" option.

This is a real and reasonable concern to have about a multi-board system, and it's exactly the concern comparable outcomes was designed to remove.

How comparable outcomes actually works

Rather than fixing grade boundaries at a specific raw mark score each year, Ofqual requires boards to set boundaries using statistical evidence — principally, the prior attainment of the cohort sitting the exam (for GCSEs, this typically draws on students' Key Stage 2 results as a predictor). The aim is that, for a cohort of similar prior attainment, roughly the same proportion of students achieve each grade as in previous years, regardless of whether that year's paper turned out to be harder or easier than intended, and regardless of which board set it.

Ofqual's own technical guide to comparable outcomes sets this out formally. The key point for the "do multiple boards cause inflation" question is this: because grade boundaries are anchored to cohort-level statistical evidence rather than to a fixed mark, a board cannot simply choose to award more top grades to look more attractive — its boundaries are constrained by the same evidence base every other board is held to.

This doesn't mean individual results are decided in advance. A specific student's grade still depends entirely on how they perform on the day; what comparable outcomes constrains is the overall distribution across the whole cohort sitting that subject with that board, so it tracks what similar cohorts have achieved historically rather than drifting upward or downward independent of ability.

Has grade inflation happened anyway?

Yes — but the clearest documented episode wasn't caused by competition between boards. During the 2020 and 2021 pandemic years, GCSE and A-level exams were not sat at all; grades were instead based on teacher and centre assessment (Teacher Assessed Grades and Centre Assessed Grades). Overall grade profiles rose significantly in both years compared with pre-pandemic exam-based results — a pattern Ofqual has documented and discussed at length in its own collection on awarding qualifications in 2020 and 2021. Grades were subsequently brought back down over the following exam series as boundaries reverted toward pre-pandemic comparable-outcomes norms.

That episode is a genuine, well-evidenced case of grade inflation in the English system. It's important to be precise about its cause: it happened because exams themselves were replaced by a different assessment method under exceptional circumstances, not because four boards were competing against each other under normal exam conditions. Attributing it to "having multiple exam boards" would be inaccurate — the same pandemic disruption would have produced a similar effect under a single-board system too, since the mechanism (no external exam, assessment by the people who teach the student) is the source of the risk, not board competition.

A worked example: what a boundary shift actually looks like

The mechanism is easier to see with a concrete illustration. Suppose a GCSE Maths paper set by one board in a given summer turns out, once results come in, to have been harder than the board intended — perhaps a particular question was ambiguously worded, or the paper as a whole demanded more time pressure than previous years' papers in that subject. Under a raw-mark system, this would simply produce lower results that year, purely as an artefact of the paper's difficulty rather than any real change in students' maths ability. Under comparable outcomes, the board (with Ofqual oversight) instead lowers the raw mark required for each grade boundary — so that, say, 62 out of 100 might be needed for a grade 7 that year rather than 68, because statistical evidence about the cohort's prior attainment (drawn substantially from Key Stage 2 results) indicates the proportion of students capable of grade-7-standard maths hasn't actually changed, only the paper's difficulty has. The student sees a grade, not a raw mark, so this recalibration is invisible to them — but it is precisely the mechanism that stops "this year's AQA paper was unusually hard" from becoming "AQA is a harder board to be graded by."

What the genuine debate looks like

Comparable outcomes is not free of criticism, and a fair article on this topic should say so rather than present the system as beyond question. Researchers and commentators have raised points including:

  • Whether anchoring outcomes too rigidly to prior cohorts' results can make it hard to recognise genuine improvements in teaching or curriculum over time, since the target distribution is partly historical.
  • Whether the specific statistical models used to predict cohort performance are transparent and robust enough for the weight placed on them.
  • Whether comparable outcomes fully neutralises every incentive for a board to be perceived as offering an easier route through subtler routes than raw grade inflation — for example, through specification design, question style, or the availability of coursework, discussed in board-shopping: do schools pick the easiest board.

This is genuinely contested territory among people who study assessment policy, and it would be dishonest to present a settled verdict either way. What can be said with confidence, because it's a matter of published regulatory design rather than opinion, is that the mechanism comparable outcomes uses is specifically built to prevent the straightforward "board lowers its bar to win market share" scenario — whether it succeeds completely is a separate, harder question.

Does "board shopping" achieve what grade competition can't?

Given that comparable outcomes closes off the direct route to competing on grade generosity, it's worth asking a sharper follow-up question this article hasn't answered elsewhere: could a school achieve something similar through choice of specification rather than choice of grade boundary? This is a distinct question from grade inflation, and is addressed at length in board-shopping: do schools pick the easiest board — but the short version worth flagging here is that specification design (question style, topic choice, how a syllabus is structured) is not the same lever as grade boundaries, and comparable outcomes doesn't reach into specification design at all. A school could plausibly find one board's specification style suits its students' strengths better than another's without either board's grades being inflated relative to the other — which is a genuinely different phenomenon from the inflation this article focuses on, and one worth keeping conceptually separate when reading claims about boards being "easier."

The bottom line

Having multiple exam boards creates a theoretical incentive for competition on perceived ease, and Ofqual's comparable outcomes approach exists specifically to close that door by anchoring grade boundaries to statistical evidence rather than letting any one board set its own bar. Grade inflation has happened in England, most clearly during the pandemic exam-cancellation years — but that was a function of not sitting exams, not of board competition under normal conditions. Whether comparable outcomes removes every subtler incentive to compete on ease is a live and reasonable question that the evidence available doesn't let us settle definitively here.