We stopped grading the engine on pages and started grading it on CSS
Every way this engine had ever been measured started from a page: our own test pages, then twelve real websites. That biases the work toward whatever those pages happen to use, and it goes completely silent wherever the engine has never been asked a question. A site corpus can only find a fault some site in it triggers. It cannot tell you that a width given in centimetres has never once been rendered. So this release generates the questions instead: 373 of them, each a single CSS declaration in a context where it does something visible, rendered by Floati and by the reference browser and compared. Nothing in it comes from a real page. The first run put 37 genuine faults on the board, and the name of each case was the diagnosis -> no hunting required. Five of the six absolute length units did not exist. Only points had ever been implemented, at seven separate places in the code, so inches, picas, centimetres, millimetres and quarter-millimetres were not recognised as lengths at all, which for a width means the box simply filled its container. A column layout never stretched its items: the code demanded that the container's width be written down explicitly, which is the right rule for a row's height and the wrong one for a column's width, so the default behaviour did nothing and every vertical stack shrank to its text. That is the shape of most sidebars, card stacks and dialog bodies on the web. Preformatted text that is supposed to wrap never wrapped. Text set to keep its line breaks but collapse its spaces did neither. A percentage margin on a box that shares its margin with its parent measured against the grandparent instead. The three individual transform properties -> the ones modern stylesheets prefer because each can be animated on its own -> were absent entirely, and angles written in gradians were silently discarded because the word ends in the same three letters as radians. The largest finding was not a fault but a shape. Grid had THREE separate implementations of one algorithm: a real one for columns, a stub for rows that reported four of the six track types as zero, and a third copy for column-flow grids with its own placement, sizing and alignment. That is why the defects kept arriving one at a time -> they were not separate defects, they were one missing abstraction seen from different angles. There is now one track sizing algorithm for both axes, the third copy is deleted, and column flow is what it always was: a placement order, nothing more. Two faults closed together rather than one at a time, and a real error in the specification's stretch rule fell out of it -> we took the larger of the content and an equal share where the standard says content PLUS an equal share, so any grid wider than its content rendered as even columns. One more worth naming because it is invisible until it is not: a shorthand written with a variable in it lost the cascade. A shorthand is supposed to set everything it controls, but ours could not while the value was still an unresolved variable, so the pieces kept whatever an earlier declaration had left there and the later rule quietly lost. On a themed site that is most of the styling. The generated sweep now stands at 365 of 373 cases within their measurement floor. What remains is named honestly: vertical writing modes on a plain block of text are still a third of a page different from the reference browser, the first-letter pseudo-element is not there, and the last percent is sub-pixel rounding. The twelve real sites improved as a side effect, without any of them being opened -> which is the whole argument for measuring the feature space instead.