Eleven Seconds on a Six-Inch Screen
The replay was pulled a couple of days after the respondent submitted, which is late enough that nobody could pretend the behaviour was a fluke of a slow server. A 12-row satisfaction grid loaded on a 6-inch portrait handset. The respondent pinched to zoom, scanned nothing in particular, then tapped straight down a single column and hit next. The on-page timer for that screen read eleven seconds.
In the desktop authoring pane, the same question looked disciplined: one screen, twelve attributes, a shared five-point scale, no scrolling. Compact. Efficient, apparently.
What the handset rendered was a different instrument entirely. And the arithmetic is worth saying out loud, because it gets lost in dashboards: that grid returned twelve numbers per respondent, and almost none of them carried information about the twelve attributes we had named.
The Working-Memory Bill a Matrix Sends
Decompose the task the way a respondent actually experiences it. Hold the scale definition in working memory. Read the row stem. Map that stem onto the scale. Locate the correct radio button in a narrow column. Repeat, without losing your place in the row you were on.
The scale header appears once, at the top. Every row after the first still demands a fresh stem-to-scale mapping, now performed from memory rather than from the page. Compactness serves the person writing the questionnaire; it is an authoring convenience that arrives on the respondent's side as a dozen sequential judgements stacked into one visual block.
The load is not uniform. A five-row grid on a scale the audience already knows behaves very differently from a twenty-row battery dropped in late, after several long modules have already drained attention. And there is a second failure hiding inside the compression: row stems get squeezed into two-to-four-word fragments that no longer name the object, the timeframe, or the comparison. "Value for money" is not a question. It is a label waiting for the respondent to invent the missing half.
When effort climbs past what the answer feels worth, people stop answering the question and start producing something that looks like an acceptable answer. That shift is the practical core of Jon Krosnick's research on satisficing in attitude measures, and a matrix is close to a purpose-built invitation to it.
Three Spreadsheet Checks That Expose Flat Rows
Straight-lining is identical or near-identical selections across every row of a battery, usually the leftmost, the middle, or the top option. You can find it in a raw export with no special tooling.
- Within-respondent standard deviation across the battery rows. Zero variance means one selection repeated twelve times.
- Time-on-page for the grid screen, compared against the median for that same screen. Not the survey total. A battery that takes a third of its own median deserves a look.
- Reverse-worded rows behaving like their positive twins. If the negatively phrased stem scores identically to the item it contradicts, the respondent read the layout, not the words.
Straight-lining corrodes rather than merely adds noise. It inflates internal consistency, so the battery scores well on the reliability statistic you would use to defend it. It compresses variance between items, so a weak attribute and a strong one converge toward the same mean. The output looks tidy and behaves like a dead instrument.
Flat Rows Are Flags, Not VerdictsSome respondents genuinely rate every attribute highly, and deleting them manufactures the distribution you expected. Write the review rule into the analysis plan in the week or so before field opens, so the threshold exists before you have seen who it removes.
What the Desktop Preview Never Renders
Three things happen to a matrix on a phone, and a resized desktop browser window will show you none of them. Test in portrait, at the handset's default browser text size.
The first is horizontal scroll: the respondent swipes right to reach the fourth and fifth scale points, and the row label slides out of view. They are now answering an unlabelled row. Second, font shrinkage, where the platform preserves the grid by shrinking everything until tap targets become unreliable. Third, silent conversion into a stacked accordion the author never opened on a real device.
That accordion is the trap worth naming. The grid becomes a sequence of separate questions anyway, which is what splitting the battery would have achieved deliberately, except now the labelling is worse and the visual break lands mid-battery without warning.
Mis-taps sit underneath all three. Narrow radio columns on a touch screen mean the recorded answer may differ from the intended one, and nothing in the export distinguishes a mis-tap from a considered response. Count drop-off on the specific screen holding the grid rather than on the survey as a whole; a battery that bleeds respondents will hide inside a healthy overall completion rate.
Handset Pass Before FieldBudget ten minutes or so per battery on a physical phone before the instrument goes live. It is the cheapest data quality intervention available to a product team.
Where a Matrix Still Earns Its Screen
A grid is defensible under four conditions together: a small row count, a scale that is genuinely identical for every row, short parallel stems, and a comparison the respondent must make side by side on one screen. Drop any one of those and you are using the layout for filing convenience.
The strongest legitimate case is comparative benchmarking and wave-on-wave tracking. There, the grid layout is part of the instrument. Changing presentation mid-programme changes the measurement, and your trend line records a formatting decision alongside whatever the market did.
From running benchmarking waves on Floq, the rule our team applies is narrow: if a battery is already in the field as a tracked series, change it at a wave boundary with a documented reason, never opportunistically. Position matters too. Place the grid early — customarily near the first five minutes — before later modules have accumulated fatigue.
One limit on that continuity argument. Wave-boundary discipline does not protect an established matrix that already returns flat rows across most respondents. That grid is locking in a dead score, and preserving it protects the appearance of a trend rather than the trend.
Rewriting a Twelve-Row Battery, Step by Step
- Audit the rows. List every stem and mark which ones have appeared in an actual results readout. Stems that have never reached a report get deleted before any layout work begins. Most twelve-row batteries contain rows that exist because adding them cost one click.
- Re-expand the fragments. Turn each surviving label back into a sentence that names the product, the timeframe and the comparison. The fragment was a compression artefact; the sentence is the question you meant to ask.
- Split the presentation. Serve items one per screen, or in pairs, with the full scale labels repeated on every screen instead of stranded in a header row.
- Randomise item order and record the order each respondent received, so position effects stay estimable rather than baked in.
- Preserve the scale itself. Change the layout or change the scale, one at a time. Doing both leaves you unable to attribute the shift.
- Re-measure. If the series must continue, run one overlap wave carrying both formats inside roughly a two-week field window and quantify the gap before you retire the grid.
What Changes in Your ExportSplitting a battery typically lengthens the survey in screens while shortening it in effort per judgement. The variance you gain between items is the whole point: attributes that used to score identically start separating.
Which Battery Would Survive the Read-Aloud Test?
Open the live questionnaire and find the largest battery in it. This takes about two minutes and needs no new fieldwork.
Pull screen-level drop-off for that battery. Count the share of respondents with zero variance across its rows. Then read each row stem aloud as a standalone question, with no header scale in view, and notice which ones stop making sense halfway through.
If that battery had to justify every row to the team reading the results next quarter, which rows survive the meeting, and would you keep the grid at all?