SQUAT User ManualSteward & panels

Working with your squad

The steward & design panels

Personas answer “can this persona figure it out?” The steward asks “is this what we meant?” Design panels and concept interviews let you explore things you have not built yet.

The steward: your judgment, run systematically

The steward is an expert review of your live product, performed as your proxy — a reviewer who already knows what the product is supposed to be, and checks whether it actually is that. It's the kind of pass a thoughtful product owner would make weekly if they had the time. Now they do.

What makes the steward different from a persona:

A personaThe steward
PlaysA real user, in character, discovering the product.Nobody. No backstory, no think-aloud theater.
Asks“Can I figure this out?”“Is this what we meant?”
Judges againstTheir own goals and patience.Your recorded intent for each surface.
OutputNeutral evidence — what happened.Prescriptive — every finding comes with a recommended fix.

One firm boundary: the steward reviews the live product experience only. It never reads your source code or internal documents — its knowledge is what a very sharp reviewer could see on screen, plus the intent registry. If it can't tell what a surface is for from those two things, that's a finding in itself — it files a question for you rather than inventing an explanation.

The intent registry

The steward's ground truth is your intent registry: durable rulings about what a surface or behavior is for. A ruling is short and testable — “no vendor names anywhere a front-end user can see,” not a paragraph of context. Rulings are never inferred by the AI; each one comes from you.

  • Seeding. If the registry is thin, the steward interviews you — one ruling at a time, confirming the phrasing back before saving. Good starter categories: naming and vocabulary, who-sees-what rules, tone, layout conventions, data handling. Three good rulings beat an empty registry waiting for twenty.
  • Imports welcome. A brand brief, a positioning doc, prior review notes — SQUAT can convert them into registry entries, preserving your actual wording.
  • Rulings are superseded, never edited. When your thinking moves on, a new ruling replaces the old one, and the old one stays on record marked superseded. The registry shows how your judgment evolved — it never erases it.

What a sweep checks

Say “run a steward sweep.” With the plugin, /steward starts this review; a bare-MCP client can invoke the steward prompt or use the same words. The steward loads your active intents and project context, then works every surface in scope against a seven-point checklist:

  1. Labels & language — does every name make sense to a newcomer? Is terminology consistent — the same thing never called two names, two things never sharing one? Is the copy in your product's own voice?
  2. Legibility & visual accessibility — type size, contrast (measured against accessibility standards when in doubt, not eyeballed), hit-target size, information density.
  3. Workflow coherence — from every state a user lands in, is the next move obvious? Does completing an action prompt the natural next step, or strand the user?
  4. Help-system efficacy — does in-context help exist, match what's on screen, and answer the questions a user would actually have there?
  5. Onboarding & tours — do they fire when they should, describe what's really on screen, lead somewhere useful, and come back if dismissed?
  6. Intent alignment — does each surface serve its recorded ruling? Where behavior and intent diverge, the steward says which one should move.
  7. Product character — does the experience express your product's adopted character consistently?

Where a browser tool is connected, the steward reviews the real product — screenshotting surfaces as evidence and measuring actual computed colors when contrast is in question. Without one, it works from what you describe, and says plainly that the sweep is narrative rather than directly observed.

When the browser tool can run scripts, the steward also brings instruments: it runs axe-core — the industry-standard accessibility scanner — on the pages in scope, so accessibility grades rest on measured WCAG violations (contrast ratios, screen-reader labels, focus order) rather than impressions. Scanner results are condensed into a handful of meaningful findings, never a flood of raw violations, with the tool and count cited as evidence. Where the sweep notices performance problems, the steward may also suggest a Lighthouse audit (built into Chrome) as a deeper follow-up. Your testers never run these tools — a real user doesn't audit a page mid-task; instruments belong to the steward.

What a sweep produces

Every sweep ends with three things:

  1. Findings — in your normal ledger. Steward findings flow through the same round pipeline as persona findings, under a reserved “steward” identity, so they land in the same findings and issues system with no separate silo. Severity comes from observed impact against your intent, and every finding carries a concrete recommended fix — steward findings are never filed bare.

  2. A report card. One letter grade (A–F) for each of the seven checks, each anchored to the findings behind it — never a bare letter — plus an overall grade (set by the worst area, never averaged: an A-average product with an F on intent alignment is not a B product), trend arrows against the previous sweep, the top five items to fix, and a “what's working — keep it” section. Stewardship protects what's good, not just what's broken. The card is saved with the sweep, so grades trend over time.

  3. Questions for you. Anywhere intent was unclear or a genuine judgment call needs your ruling, the steward asks directly: “what did you intend here? what's the spirit of this?” Your answers become new registry entries — so the next sweep judges against a slightly more complete registry than this one did.

Tip

The steward is judgment-dense, low-volume work — the opposite economics of playing testers. Run it on your strongest available model, even if your actors run on something cheap.

Inviting the steward into a normal round

Once your project has an intent registry, a steward sweep can run only when the reserved steward already belongs to that project's default line-up, when one exists. If it does not, an owner/admin must version the line-up first. A separate round still resolves the same project default, and a projectless round is not an acceptable workaround because it would break attribution. SQUAT never treats “include the steward” as permission to override the saved cast. When it shares a regular round, the steward reviews the parts the testers touched and its report card appears alongside their results.

Design panels: move the squad left

A panel guide is a versioned moderator instrument, stored in the existing scenario container. It records the research purpose and ordered phases, while each turn stores its spoken question separately from moderator-only objectives, probes, listen-fors, and exercises. An imported guide is source material, not executable authority; SQUAT validates or compiles it against the pinned system-level concept-interview knowledge.

Catching “this confuses people” before it ships is strictly cheaper than after. A design panel puts an artifact — a wireframe, a mockup, a redesign proposal, a copy draft — in front of your squad before it's built. Same personas, same accuracy discipline, no live app required.

The project's default squad is still the panel's cast authority. SQUAT may use the whole line-up or a deliberate subset, but it does not assemble a supposedly “balanced” panel from outsiders. A different panel cast starts with an owner/admin versioning the project default or squad.

Compiling a concept-interview guide

Provide the product/concept, intended audience, research purpose, and hypotheses to explore. SQUAT produces a neutral guide, generally moving through:

  1. Baseline behavior — current reality and recent stories before the solution is introduced.
  2. Concept reaction — an uncontaminated first read before feature explanation.
  3. Feature trade-offs — forced choices only after the concept is understood.
  4. Price last — only when price is in scope, after earlier answers cannot be anchored by a number.
  5. Synthesis — boundaries between interesting, useful once, repeatedly useful, and worth testing further.

Questions seek concrete past/current behavior and stories before hypothetical preference. A short neutral classification question can route a conversation, but it must be followed by an open story probe; a closed preference question alone is weak evidence. The guide explores hypotheses — it never claims to prove a preferred answer.

Four panel jobs

  • Panelist/actor receives the provenance-excluding runtime persona, current spoken question, and current artifact only.
  • Moderator receives the participant reference, private guidance, and dialogue history. It chooses neutral probes without revealing hypotheses or listen-fors.
  • Neutral reviewer receives the completed exchange without persona provenance. It may approve the reaction or request a neutral follow-up or redirect.
  • Synthesizer works across the completed panel only after every panelist reaction is recorded.

SQUAT-generated actor, moderator, and reviewer bundles exclude source provenance, real/synthetic labels, and cross-role fields. The operator-facing report restores safe governance labels afterward. SQUAT records which bundle it returned, but it does not attest that an external model host kept those contexts isolated; the coordinator remains responsible for separate model calls or contexts.

Say “run a design panel” and share the artifact. Task scenarios describe work against an implemented product for scored UAT. A panel instead uses exactly one versioned moderator instrument for one research purpose. Its ordered questions organize the conversation while still allowing neutral follow-ups and redirects. SQUAT reuses an instrument only when its saved purpose fits the panel you asked for; otherwise it creates one intentionally before proceeding. It checks the project, authoritative line-up, selected panelists, moderator instrument, and finalized artifact version before starting. Panel setup stays scoped to the project, purpose, and files you selected; unrelated work from the surrounding session is not folded in.

During the panel, SQUAT names the next ready job and checks the recorded moderator question, the panelist response, the neutral review, and the result before moving on. If a client skips or repeats a required step, the panel pauses at the last valid point and directs the moderator back to the missing step. Follow-up wording remains flexible; the recorded order and evidence are what make the panel complete. A one-shot structured reaction written without dialogue is incomplete.

For a connected visual or file panel, SQUAT saves the inputs in a reusable project artifact set and pins one exact finalized version to the round. Each moderator question can name ordered artifact slots containing exact images, audio, text, or other file types. The panelist receives only the current question's slots. The coordinator places those exact files visibly in the conversation before asking the question, and SQUAT records the exposure and response times before allowing the panel to complete. Small image, audio, and file content works directly. For a larger item, SQUAT can offer an authenticated resource to clients that support safe attachment; otherwise it uses the local launcher path. If the selected model route cannot receive the exact file, the panel stops instead of inventing a reaction.

The artifact library can show the workspace, one project, or the exact artifact versions previously paired with a scenario. Rich clients use a thumbnail slide-sorter with click-to-open previews and clear metadata fallbacks; Markdown clients receive the same names, versions, types, sizes, and selectable identities. Temporary download links and storage credentials never appear in the card.

Every spoken question, probe, answer, hesitation, interjection, and disclosure is appended to the panel's canonical dialogue ledger. SQUAT reconstructs the ordered exchange from that chain; ordinary turns save directly and unusually large content uses the large-file path transparently. Hosted content is protected automatically. The panel cannot complete while dialogue or result evidence is still pending.

In a compatible rich client, open the live panel viewer to watch the moderator question, panelist reply, and any follow-up or redirect as they enter that canonical dialogue. The viewer resumes from the last recorded turn after a reconnect and checks access again while it runs. It is a view of the same evidence, not a separate transcript. It is visibly labeled SQUAT authoritative because the view is derived from your authorized ledger; direct clients such as LM Studio receive the same label before the viewer instructions. The label identifies the source of the record and does not claim that a panelist's statement is independently verified. Each live update repeats the same label and is checked before new turns appear.

A design panel requires canonical dialogue. If capture is unavailable, the panel pauses; the valid next step is to resume, recover an already-authorized turn, or invalidate the round. Eligible non-panel rounds may use the separate approved evidence-gap flow.

If the connection is interrupted

Everything SQUAT confirmed before the interruption remains safe in the hosted ledger. A durable local file may preserve an already-authorized turn whose append failed. The panel then pauses until the server can issue the next role brief. Chat history by itself is not treated as a recovery file.

When the connection returns, SQUAT checks that the same panel is still open and that its dialogue has not changed since the local checkpoint. You review how many turns will be added, then approve the load. The complete batch is accepted or none of it is; a failed check leaves both the hosted round and local file unchanged.

After a successful load, SQUAT identifies the panel by its readable project and round context and gives you the number of turns added and an import receipt. The stable round ID remains available as secondary, copyable metadata for an exact lookup. The local file is not changed. Recovered turns use the same transcript, panel-reaction, report, export, and deletion paths while remaining identified as recovered rather than performed live.

Two rules are non-negotiable:

Note

React, don't design. A persona names what confuses them; it never proposes the fix. The moment a simulated user starts designing, you've lost the signal you ran the panel for. Design decisions stay with you.

Note

Qualitative stays out of the scored ledger. A panel round is never pass/fail, never enters trend or delta math, and its observations never auto-become tracked issues — a reaction isn't a defect report. The ledger marks panel rounds distinctly, and round comparisons refuse to include them. Your trend line stays a trend line of real tests.

Preference and pricing statements from a panel are hypotheses, not customer findings, willingness-to-pay evidence, empirical research, or market validation.

From one panel to a research readout

At the end of a complete panel, SQUAT can prepare a Round Summary. It names the available evidence, then separates what the panel said from interpretation, hypotheses, next steps, and uncertainty. A disagreement stays visible; it is not averaged into a pretend consensus.

The full summary is a document, not a crowded card. You can save it, revise it, compare every point against the canonical dialogue, or leave it as a draft. Saving is always explicit, and every revision or addendum preserves the earlier version.

A Research Report combines selected saved Round Summaries across a study. You can start with the rounds available today, then add more Round Summaries in a later report version without changing the earlier report or its sources.

A panel never replaces a live-app round once the thing exists — it is an early hypothesis-generating read, not a verdict.