QueryPort

The palette was chosen on fills. Three of its six colours failed as text.

A palette gets approved by looking at swatches and buttons, and then the same colours are used to set words. Three of six measured 1.70, 2.50 and 2.83 against the paper, under the 3.0 that even large text needs. What to measure, and why darkening everything is the wrong repair.

A palette gets approved the way palettes are always approved: someone looks at swatches, buttons and filled panels, and picks. Then the same values go into the stylesheet, and some of them end up setting words. A colour that is excellent as a fill and unusable as text is not an unusual colour. It is the normal case, and nothing on the comparison sheet warns you about it.

The measurement

A palette we were choosing was presented properly: five variants on one screen, so the decision was a comparison rather than a description. The choice was correct, too, and it fixed a real problem: in the previous range three colours read as one brown.

The palette was six colours. It was chosen on fills. On the site, one of those colours sets a name, which is the largest word on the result screen and on the image people share. Measured against that site's paper immediately after installation, three of the six came in at 1.70, 2.50 and 2.83 — against 3.0, which is the threshold for large text, the easy one.

Why the comparison sheet could not have shown it

The sheet that won the decision had swatches, strips and a colour wheel on it. It did not have a single caption set in the palette colours. So the failure was not missed through carelessness: it was not on the sheet to be missed. A colour does at least three jobs — it fills, it outlines, and it sets type — and a sheet showing one of the three collects an opinion about one of the three.

Two rules came out of that, and they are the whole of the process part. When you put a palette up for a decision, show every job of the colour: the fill, the outline, and a line of type set in it. And once a choice is made, measure the contrast without asking — approval is a decision about appearance, not a statement about legibility, and the person approving is not claiming otherwise.

Two repairs that sound right and are not

  • Take a darker version of each colour. This destroys exactly the property the palette was chosen for. The previous range had been abandoned because three of its colours read as one brown; darkening a bright set walks straight back into that, and the second decision quietly cancels the first.
  • Set the text in black instead. This works for legibility and breaks identification. Where the colour is the only thing tying a name to its slice in a chart, black type means the reader has to look the connection up rather than see it. You have moved the cost from the eye to the memory, and called it a fix.

What we did instead: split the colour by job

One key in the data, two values out of it. The chosen colour keeps doing the fills exactly as approved, and the shade used for type is computed from it rather than stored beside it. There is no second palette to keep in step, and no opportunity for the two to drift.

The one detail that matters in the computation: darken in steps until the threshold is met, not by a fixed fraction. A fixed fraction is a single number applied to colours that start at different lightnesses, so it overshoots the dark ones, drying them out for no gain, and never reaches the light ones, which are the ones that failed in the first place. Stepping to a threshold gives every colour the smallest change that makes it readable, which is also the smallest change to the look that was approved.

And the formula lives in one place. Pages, share images and anything drawn on a canvas take the finished value from the same computation. A second implementation of the same formula is two answers to one question, and they diverge on the day someone fixes a rounding bug in one of them.

The same failure in our own palette

Our own site had a version of this that had nothing to do with taste. A muted grey from a stock scale, text-stone-500, was used for small print. It measures 3.90:1 on our dark background and exactly 4.50 on the light one — one class, two backgrounds, one pass and one fail. It had spread to 15 files, and every one of them looked fine to the person who added it, because on a light screen it is fine. A test now refuses the muted grey scale outright, because the judgement that put it there will be made again.

The palette we settled on was measured before it was used: 16 pairs, minimum 5.32:1. That is a line in our release notes rather than a claim about our taste, and it is the only form in which this can be stated.

Contrast is not only about type, and the cheapest reminder of that came from our own brand mark. In one of the two arrangements its tail vanished at 16 px: the dark green and the graphite sit too close in lightness, and at small sizes a thin shape has nothing left to stand on. Nobody reported it as low contrast. It was reported as a missing tail.

The tool below measures one pair. It runs in your browser, sends nothing anywhere, and names the consequence rather than a grade.

Measure one pair

Type the two colours as HEX, or pick them. The arithmetic happens in your browser and nothing is sent anywhere — there is no send code on this page.

— the ratio does not change, the sample does.

Large text, 24 px bold

Body text at the size most of a page is actually set in. This is the line to judge by: if reading it is work, the number above has already told you so.

Small print, footnotes, form hints and the caption under a chart.

Border and icon, non-text

Where this measurement stops being reliable

  • It says nothing about colour blindness. The ratio compares lightness only. Two hues can sit far apart on this scale and still be the same colour to someone who cannot separate them, so a passing pair can still be a chart in which two series are indistinguishable.
  • It ignores weight above the large-text threshold. Once type counts as large, the ratio treats a heavy face and a hairline face identically. A thin 24 px face at just over 3.0 is legible on the designer's screen and gone on a phone in daylight.
  • It does not survive text over a photograph. There is no single background to enter: it changes under the letters, and the worst spot under one word decides the outcome. A plate, a scrim or an outline is the fix, not a better number for the average pixel.
  • It assumes both colours are opaque. Any transparency, and the pair being measured is not the pair on screen — what is actually behind the element is. Compose it first, then measure what came out.

Read next

The transferable part is not about colour. An approval is a decision about one property, and it arrives looking like a decision about all of them. Ask what jobs the thing you just approved will be doing, and put every one of those jobs on the sheet before you ask.

The second half is cheaper still: after the decision, measure without asking. Nobody approving a palette is claiming the colours are legible, so nothing is being overruled when you check — and a number, unlike an opinion, can be put in a test that keeps the next person from making the same call again.

How our gateway handles input and refusals · Tell us where our own contrast is wrong