10 / 12  ·  Instrument sampling 0 frames

Chapter 10 · measurement as exhibit

INSTRUMENT

Eleven chapters in this portfolio made claims. A frame budget, a body count, a contrast ratio, a millisecond. Claims are cheap. Measurement is not — it costs a number you are willing to be caught by.

This is the chapter that closes the loop. Every figure on this page was computed by the code you are reading, in the tab you are reading it in, and every figure can be recomputed from the raw sample it came from. Open the exhibits, break the page on purpose, and watch the instruments say so.

01 / The claim

A number on screen is a sentence somebody wrote.

60 fps. 28.8 ms per frame. AA contrast throughout. Sixty-five thousand bodies. Each of those is a sentence, and every sentence was written by the same person who benefits from it being believed. That is not an accusation. It is the ordinary condition of making things, and the only available response is to make the number checkable.

A measurement is a claim with a second claim attached: this is how it was measured, and you can run that method yourself. The second claim is the expensive one. It is why a portfolio that cannot be checked is worth exactly as much as its loudest label, and it is why the last chapter of this one is not a summary. It is an apparatus.

The failure modes are not subtle, and this project has already produced most of them. A suite reports all green while being unable to detect a defect deliberately injected into the file it was supposed to be reading. A frame-time readout quietly reuses the solver's stability clamp, so the number on screen under-reports by a factor of nearly four and looks plausible the entire time. A masthead says 65,536 while a telemetry strip nine hundred pixels away says 6,553, permanently, because something upstream changed and nobody rewrote the label. A caption asserts that a quantity equals itself. None of these are caught by looking hard. They are caught by breaking the thing on purpose and watching what notices.

A suite that cannot fail is not evidence. It is a decoration shaped like evidence, which is worse than no suite, because it spends the reader's trust.

So the rule for this chapter is narrow and absolute: every number printed below is a value the running code computed, from a sample the reader can pull apart, at the moment it is read. No constant stands in for a measurement. No display value is clamped, smoothed into compliance, or passed through a solver's stability limit on its way to the screen. Where an API is missing, the page says the API is missing — it does not print a zero, and it does not print a plausible number it cannot defend.

02 / The trace

Every frame, kept.

One array: the raw inter-frame delta of every animation frame since load, in milliseconds, stored unmodified. The line above is a direct projection of that array — same samples, same order, no smoothing, no averaging. The table below it reports percentiles of the same array, computed by a nearest-rank rule pinned on this page. Change the window and both change together, because they are not two features; they are one number and two ways of looking at it.

Frame-interval trace Line plot of the interval between consecutive animation frames, most recent sample at the right. Values above the plot ceiling are drawn as flat notches on the ceiling rail and counted in the caption; the numeric percentiles below the plot are exact and are never clipped.
window
band

samples
0
p50
—ms
p90
—ms
p99
—ms
max
—ms
over 50ms
0

03 / Method

Four collectors, and what each one is allowed to claim.

The instruments on this page read from four sources, and each source is trustworthy in a different way. A requestAnimationFrame loop gives exact frame intervals with no clamping, but it only samples while the tab is foregrounded, so it is a statement about the frames you actually saw. The PerformanceObserver gives long tasks, layout shifts and paint timing straight from the browser's own instrumentation, which is more authoritative than anything this page could compute about itself and, for the same reason, worth less than a claim it does not make — it reports what the engine recorded, not what the engine ought to have recorded.

Contrast is the third, and it is the one most often faked. It is computed from the browser's resolved colour of a real element, not from the token that was written in the stylesheet, and not from the hex in a comment. That distinction turned out to matter: the brand file annotates oklch(65% 0.25 350) as #FF1F78, and those are not the same colour. One is rgb(243,42,164); the other is rgb(255,31,120). An instrument that read the comment would report a contrast figure for a colour this page never paints.

The fourth is the one with no API at all: claim auditing. Several things on this page assert something about itself — that the worst small-text pair on the page clears a threshold, that the harness is armed, that the budgets can fail. Those are claims, and claims are what this chapter is about, so they get the same treatment. EXHIBIT 04 holds a set of defects that can be switched on one at a time; each one leaves a specific claim behind, and the audit finds it.

What follows, then, is not a scorecard. It is four instruments, one claim audit, and a self-check that is expected to fail when it should.

Exhibit 01 Frame timing, in distribution rAF · longtask · layout-shift · paint · heap closed

Percentiles, not an average. An average of frame times is the single most common way to make a janky page look smooth, because one 900 ms stall in four hundred frames moves a mean by two milliseconds and moves a p99 by eight hundred. The rule used here is pinned and printed: nearest rank on the sorted sample, sorted[ceil(q · n) − 1], clamped to the array. Both the display path and the audit implement that rule separately and are compared against each other on every run.

Nothing here is clamped. A page that animates needs a stability clamp on its integration step — a 900 ms tab-switch must not teleport a spring — and that clamp is a solver concern, not a measurement. The only ceiling on this page is the plot ceiling, which is a drawing decision and is labelled on the axis; the numbers underneath it are exact at any magnitude. Press STALL 400 ms and watch a single figure leave the chart and still be reported in full.

samples kept
0
p50
—ms
p90
—ms
p99
—ms
max
—ms
frames > 50ms
0
long tasks
0
longest task
—ms
blocking time
—ms
shift events
0
CLS
—
worst shift
—
Signal Value Source Scope
First paintfirst contentful style/layout — PerformanceObserver “paint” this document
First contentful paint — PerformanceObserver “paint” this document
Frames committedanimation frames the browser presented 0 requestAnimationFrame while foregrounded
JS heap in usepage-isolated, quantised by the browser — — this document
Sub-resource transfersanything fetched after the document 0 Resource Timing this document

STALL 400 ms blocks the main thread for four hundred milliseconds on purpose. It is a real stall, produced the way stalls are produced, so the long-task observer sees a real long task and the frame sampler sees a real four-hundred-millisecond gap. Nothing about it is simulated, and the page does not try to hide it afterwards — which is why the budgets in EXHIBIT 02 turn red when you press it. That is the intended demonstration, not a defect.

Exhibit 02 Budgets with published thresholds nine gates · press stall to watch them fail closed

Every row below is a threshold with a stated provenance: PUBLISHED rows are industry figures (Core Web Vitals), HOUSE rows are this page's own, written down here so they cannot be moved later without leaving a visible diff. A budget that cannot fail is a decoration, and a decoration in a table of numbers is the exact failure this chapter exists to correct — so the honest version of the claim is a budget that goes red.

Right now this page should pass most of them, because it is a document with a small script on it. Press STALL 400 ms in Exhibit 01 and three of these turn red within a second and stay red, because the stall is real and nobody is quietly resetting the counter. The claim this page is making is not “we pass every budget.” It is: these are the budgets, this is the measurement, and here is the measurement failing.

Budget Threshold Measured Provenance State
Summaryrecomputed from the rows above on every render — —
Worst margindistance to the nearest threshold — —

How to falsify this panel. The comparison function used by every row is a single function, cmp(value, budget), evaluated in one place. A budget that cannot fail is one where that function is rigged. EXHIBIT 04 has an INJECT control that does exactly that, and the audit catches it by calling the comparator directly on a value that is known to be out of bounds — it does not infer failure from the panel's own rendering, which would be trusting the thing under test to grade itself.

Exhibit 03 Accessibility, measured not asserted contrast · focus order · keyboard path · ax tree closed

Contrast is computed from what the browser resolved, not from what the stylesheet said. For each pair below, the foreground comes from getComputedStyle(el).color on a real element, the background is the ancestor chain's background colours composited in sRGB alpha until opaque, and the ratio is WCAG 2.x relative luminance. The colour string is then handed to a canvas, so the value used is the browser's own gamut-mapped result — which is how oklch(65% .25 350) arrives as rgb(243,42,164) rather than the #FF1F78 written beside it in the token file.

The brand rule, tested. The house note in gmp-tokens.css says pink and green “are not for interface — they fail contrast on black below 18 px.” The measurement does not support that sentence as written, and the honest thing to do is print the measurement and let the rule be argued with. What follows is the real matrix. Where the rule is stricter than the standard, it is kept, and the row says so; where the standard is stricter than the rule, the standard wins. Neither is allowed to hide behind the other.

Rendered pair Resolved Ratio Required State
live pairs sampled
0
worst small text
—:1
text below 24px
0
tab stops
0
order = DOM
—
positive tabindex
0

—

Computed accessibility tree

A partial tree built from the DOM at read time: explicit role, or the implicit role implied by the element and its attributes, plus the accessible name the HTML-AAM algorithm would compute from aria-label, aria-labelledby or subtree text, plus the live state. The browser's internal accessibility tree is not exposed to script on any engine, so this is a faithful reconstruction from the same inputs, not a read of the real tree — labelled that way because the difference matters to anyone deciding whether to trust it.

Exhibit 04 Break the page on purpose five defect classes · closed closed

The thesis, made tactile. Five defect classes from this project's own history are wired to INJECT buttons. Each one breaks something real, and each one leaves a specific claim on this page behind. Press RUN SELF-CHECK afterwards and the audit names the claim, the value it displayed, and the value it recomputed — not “test failed.”

Two properties keep this from being theatre. First, every injection is verified to have landed: the harness fingerprints the page before and after and reports INJECTION FAILED rather than CAUGHT if nothing changed. A harness that reports a catch it did not earn is the specific catastrophe this project already produced once. Second, the registry count and the audit's own verdict are printed outside the registry, so neutering the checks cannot hide the neutering.

Defect What it breaks State
Audit verdictprinted outside the registry — cannot be edited by neutering it — —
Check registrydeclared in source vs registered in memory — —

Self-check report

Keyboard path

A scripted walk through the real controls using real key events dispatched on the real handlers, logged as it goes. This is a transcript, not a summary.

How a number on this page gets to the screen

Frame interval. A requestAnimationFrame loop stores t − t_prev into one array in milliseconds. That is the whole pipeline. No smoothing window, no outlier rejection, no clamp. Percentiles use nearest rank on a sorted copy: sorted[ceil(q · n) − 1], index bounded to the array.

Long tasks and blocking time. The browser's longtask observer, with 50 ms subtracted per task before summing, which is the standard definition of total blocking time. The subtraction is the definition, not a fudge: a task that takes exactly 50 ms blocks nothing.

Layout shift. The layout-shift observer, summed across the standard session window — a new window starts after a gap of more than one second or five seconds of accumulation — and entries with recent input excluded. CLS is the largest such window, not the lifetime sum.

What the audit is allowed to trust

The self-check in Exhibit 04 never asks a display function what it displayed. It reads the text content of the DOM and compares it against a percentile it recomputes from the raw array using its own implementation of the pinned rule. That is the only reason it can catch a clamped readout, and it is also the only reason it would catch one introduced by accident.

A green audit is a claim, not a verdict, and the file-level harness in qa/ch10/mutate.js is the part that holds it to that standard: it re-reads each mutant from disk, checks the byte hash moved and the marker is present, asks the suite which absolute path it actually opened, and reports a surviving mutant as a surviving mutant.

Known limits

Frame sampling stops when the tab is backgrounded, so a long idle gap is absent from the distribution rather than recorded as a spike. Heap figures use performance.memory, a non-standard Chromium surface quantised to 100 kB; on any engine that omits it the row says so rather than reporting zero. The computed accessibility tree is a reconstruction from DOM inputs, not the engine's own tree. And the audit can only catch a defect that changes something it reads — a defect in code no instrument touches is invisible to it, which is why the file-level mutants matter and are not a formality.

10 — Instrument ↑ Top Exhibits Commission it one file · no dependencies · runs from file://
← Lab