SQUAT User ManualRunning rounds

Working with your squad

Running rounds

A round is the instrument itself: squad members attempt scenarios against your product in character, an independent evaluator scores what actually happened, and the results land in your ledger. Say “run a UAT round” and SQUAT orchestrates the whole cycle. The plugin entry point is /run-round; clients without slash skills can invoke the run-round MCP prompt or ask for a UAT/panel round directly.

The shape of a round

  1. Select
  2. Act
  3. Evaluate
  4. Submit
  5. Summarize
  1. Select. Choose the project first. If it has a default squad, that saved line-up is authoritative: run the full line-up or choose a subset of it. SQUAT never silently adds someone from the wider roster or substitutes another squad — including for a design panel. Each persona×scenario pair is a full agent run, so SQUAT confirms rather than assuming “everyone, everything.”

  2. Start. You can give the round a concise readable name, and SQUAT records it with the round metadata before anyone acts: who's running, which scenarios, what model and platform is doing the testing, and — if you know them — your product's version and what's changed since the last round. It also asks whether this round has a mission. Task scenarios describe work against an implemented product; design panels use versioned moderator instruments for their spoken questions. SQUAT checks the selected panel setup before uploading a new artifact and checks the finalized artifact version again before it creates the round.

  3. Act. For each pair, a fresh actor loads a provenance-excluding runtime persona plus the current task. The SQUAT-generated bundle contains no source provenance or real/synthetic label. It narrates what it sees, thinks, and does; abandoning is a valid ending.

  4. Evaluate. A separate evaluator reads the transcript and expectations without persona provenance until scoring is complete. It judges the outcome, classifies failures, and produces evidence/measurement. See Findings & scores.

  5. Submit and summarize. Each result is submitted to the ledger with its canonical dialogue references and available evidence. Attachments are saved in the project artifact library, where they remain available to later rounds and reports. The report then gives pass/fail counts, the change versus your previous round, regression candidates, and a worst-first list of findings.

Checking current state

When you ask “where are we?”, SQUAT reads the live project record and names the current line-up, setup gaps, open valid rounds, latest scored UAT round, and latest qualitative panel. It can also tell whether a scenario was merely selected in a valid panel or actually reached panelists through recorded dialogue or submitted reactions. Asking whether one round is open or complete, or whether a proposed panel is ready, is read-only and does not create a probe round.

Transcript origin

The canonical dialogue record carries one origin note: captured live in this SQUAT round, recovered after an interrupted append, or imported from an archive. Reports and result summaries do not repeat it, and the note never enters the actor or evaluator briefing.

The ground rules

SQUAT holds itself to a short set of operating rules during every round. You don't have to enforce these — your AI does — but knowing them explains why rounds behave the way they do:

  • A test has to be able to fail. A vacuous test is flagged, not celebrated.
  • Keep acting and evaluating in separate contexts. SQUAT supplies separate, minimized bundles and records which one it returned; it does not attest that an external model host kept them isolated.
  • Cast the actor cheap, the evaluator strong. Playing a user is high-volume work; judging impartially is judgment-dense work.
  • Metadata is written before the actor acts. No post-hoc storytelling.
  • A target change mid-round voids the run. If a deploy lands mid-session, the round is invalidated — not massaged.
  • Re-verify a “fixed” issue with the same persona that found it.
  • The project's default squad controls the cast. A round may use all of it or a subset, never an implicit override. An owner/admin versions the project or line-up first when the cast truly needs to change.
  • A round that runs short still runs. If one of your ten personas isn't available, the panel doesn't stand down — the round proceeds and the record names who was missing.
  • Invalidate a bad session; never delete it.

When a persona can't take part

Real panels run short. A persona may be unavailable, its source may have withdrawn consent, or a squad member may never have been activated. SQUAT's rule is the same one you'd apply to a room full of people: you don't cancel a panel of ten because one person is out. The round runs with everyone who's there.

What it will not do is quietly pretend the panel was full. When your cast is smaller than the squad line-up it was drawn from, the round stores the line-up it started from and every declared member who didn't take part, with a reason where one was given — unavailable, consent withdrawn, retired, not activated, or other. Asking about the round says so in plain words: “Nine of the ten declared panelists took part; Marta Nilsson was unavailable.” Your round summary should repeat it, because a reader should never have to spot a shortfall by comparing two lists.

Reasons come from that fixed list rather than free text, deliberately: an absence is a scheduling fact about your panel, not research content, so it stays plainly readable to you rather than being sealed alongside your testers' responses.

Your standards, not ours. SQUAT records what happened; how small is too small is your call, and it varies honestly between an internal design review and a regulated study. Decide up front what your minimum panel is, what you'll do below it (run it as indicative only, reschedule, or substitute another persona), and which absences actually matter to your findings — a panel that loses its only accessibility-focused tester has lost more than one voice. Tell your AI that standard when you set the project up and it will hold rounds to it, and say so when one falls below the line.

New research, or a retest?

If your workspace has open issues and you start a round without naming any of them, SQUAT says so — how many are outstanding, and a few of them by name. It does not stop the round, and it is not asking you to clear your queue first. Plenty of rounds are new research and have nothing to do with what is already open.

What it prevents is the quieter failure: a round described afterwards as a retest that never re-checked anything. If you are re-verifying, name the issues when you start — SQUAT then tracks the outcome against them when the round completes. If you are not, say so once and carry on.

Re-verify in a new round, never by finishing the round that found the bug. That round is your evidence the bug existed; completing it to test the fix spends that evidence, and you cannot get it back.

Missions: your question for this round

A scenario is the reusable instrument; a mission is your specific question for this round — “shake out the new capture flow before Tuesday's launch,” “focus on mobile, ignore desktop this time.” Missions are optional. A project can carry a standing default mission that's offered whenever you don't give a round-specific one. Mission text is protected like other sensitive content; round reports note only that a mission was present, and the review step closes the loop on it — did we learn what we sent them for?

Governed-runtime preflight

Before SQUAT creates a round, it checks the exact version of every selected squad member. Every persona in the cast needs an active communication capsule for that same version — not only the governed ones. A newly recruited or imported governed persona additionally needs an approved communication range, a synthetic calibration decision, and an active provenance-excluding actor briefing. SQUAT pins those exact artifacts into the round.

The round stops before anyone performs if a persona is still pending, is pinned to another version, or has been withdrawn (a source was pulled) or rejected (it did not pass its audition). Those two are separate outcomes and SQUAT names which one it hit — they need different remedies: a withdrawal is a provenance problem, a rejection is a quality one.

Older personas remain readable for compatibility, but they run as legacy-unverified until recompiled. Every scored round also pins one framework version. The evaluator receives that exact framework from the round brief; it never silently grades against a newer rubric.

Inviting the steward

If your project has an intent registry and the reserved steward already belongs to its default line-up, SQUAT can ask one extra question while casting: “include a steward sweep this round?” Otherwise an owner/admin must version the line-up first. A separate steward round resolves the same project default; it never bypasses the saved cast. Details in Steward & panels.

Transcripts and evidence

New rounds record every spoken question, probe, answer, hesitation, interjection, and disclosure in one canonical event ledger. SQUAT reconstructs the ordered dialogue from those inline turns, including who spoke and which earlier turn a response answered. Binary inputs and evidence use the separate project artifact library and remain attachable to later rounds and reports.

For newly created rounds, the ledger commits to the exact meaning of each spoken turn while storing its protected envelope separately. This lets an authorized protection-key maintenance operation replace only the envelope without changing the immutable event chain. Reading the transcript and its public shape stays the same. Before a round closes, SQUAT verifies both the structural chain and the revealed committed content, then refuses completion if the stream or its protected-content revision changed during that check. Existing legacy rounds keep their original event bytes and remain readable; they are not silently rewritten into the newer format.

Long transcripts arrive in pages. When you ask for a transcript, SQUAT returns up to a page of turns at a time (250 by default) and keeps fetching until the exchange is complete — on clients that cap how much a single reply can carry, it simply uses smaller pages. You never need to manage this; if a transcript looks truncated, ask SQUAT to continue and it picks up exactly where the page ended.

You can give a study, panel, or scenario an optional readable two-word code for conversation and lookup. The code is an alias; the saved record and exact version remain the source of truth.

Event-ledger rounds require the active dialogue coordinator credential at completion so another runner cannot close the stream. A contract panel requires canonical dialogue: after an interruption it must resume, recover an already-authorized failed append, or be invalidated. Only eligible non-panel rounds may use the separate operator-approved evidence-gap path.

Browser-driven rounds and screenshots

Full detail — how a tester perceives a page, live vs. narrative rounds, and the one-time browser-control setup per client (including the Chrome debugger route) — lives on How testers work.

When a scenario targets a web app and a browser automation tool is connected to your client, the actor drives your real application — clicking, typing, navigating — rather than narrating hypothetical clicks. Along the way it captures screenshots at the moments that matter:

  • First impression — what the persona sees on arrival.
  • Any friction point — wherever the persona hesitates or the UI surprises them.
  • Anything the evaluator must judge visually — layout, spacing, visual bugs. A transcript can describe an action; it can't describe a misaligned button.

Screenshots are uploaded to the protected project artifact library, where the evaluator — and you, later — can review them as first-class evidence, reuse them in a future round, or attach them to a report. On visually-focused scenarios, SQUAT prefers a vision-capable evaluator model where your setup allows one; if the evaluator genuinely can't see images, it says so plainly rather than guessing.

Your artifact library

You can browse all workspace artifact sets, narrow the view to one project, or ask which exact versions appeared beside a scenario in valid rounds. A rich client shows a thumbnail slide-sorter with a metadata fallback for files it cannot preview; click a tile to open its preview and details in the card. A Markdown client lists the same named sets, versions, and file metadata.

When a larger file cannot travel through the ordinary inline path, a compatible remote client can request an authenticated SQUAT resource for the exact item. If that client cannot safely attach the resource, SQUAT uses the local transfer path or stops; it does not create an ad hoc upload script.

Owners and admins can rename a set, change its retention policy, rename an item for display, add a new immutable version, or delete an unreferenced set. Versions already pinned to rounds and reports remain exact, and referenced sets refuse deletion.

ClientBrowser control
Claude Code / Claude appRichest option: the Claude in Chrome extension, or a Chrome DevTools MCP server for a standalone browser.
LM StudioNone built in — rounds are narrative unless you wire up an external browser MCP (e.g. Playwright) yourself. If you do, the actor uses it automatically.
Codex CLIUses whatever browsing Codex itself is configured with; SQUAT neither grants nor restricts it.

Whatever the client, the rule is the same: use what's actually connected; never invent UI interactions that didn't happen. If a scenario needs a browser and none is connected, the transcript says so plainly.

The credential rule

Warning

No persona reads, types, prints, or pastes a password. Ever. This applies to actors and evaluators, in every client, at every tier, with no exceptions.

If a scenario needs an authenticated session, it records a non-secret authentication policy: allowed setup methods, exact login and SSO origins, a safe post-login check, session reuse, account isolation, and how often you must be present. It never stores an account, credential reference, cookie, browser state, local path, selector, or script.

Establish the session before the actor starts in one of three ways: use an existing authenticated UAT session; use an approved local login helper when your operator has configured one; or hand the headed browser to you.

Typing directly into the browser

For operator entry, SQUAT navigates to the exact approved login URL and then pauses browser automation, DOM and screenshot capture, recording, transcription, HAR, console, and network-body capture. You type or password-manager-autofill directly in the browser — never in chat. SQUAT resumes only after you confirm and the scenario's non-secret post-login check succeeds.

Warning

Do not put login secrets in a repo .env, ~/.squat/agentcredentials-*.txt, command argument, environment variable, MCP field, or generic local-file tool. File permissions and .gitignore do not stop an approved agent process running as your OS user from reading or emitting the file. Actual secrets belong in macOS Keychain/Passwords, Windows Credential Locker, Linux Secret Service, or an approved password-manager adapter.

If no allowed prepared/local method exists, round_start fails closed as local-auth-required before it creates a round, so there is no result to classify. If a session was validly attested but later disappears or a login wall reappears during the actor pass, that in-round attempt may be recorded as blocked and classified as a harness or test-design problem — never as an app failure. The actor never asks for a password.

Tip

Use dedicated least-privilege UAT accounts. Parallel state-changing personas need separate accounts; otherwise serialize them. Never use personal or production credentials. Cookies and browser storage are bearer secrets too, so keep them local and memory-only per round by default.

SQUAT does not store UAT login secrets. Only the non-secret scenario policy and readiness attestation enter the workspace.

Returning users

Two things make a persona's second visit different from their first:

  • Learnings — the behavioral observations they've accumulated ride along in their briefing: they remember where things were and notice what moved. First-visit behavior is only correct for a persona with no history.
  • Acknowledgments — before a returning persona is briefed, SQUAT checks whether issues they found have since been addressed, and folds that in, in character: “the team fixed the save bug you found.” Real testers remember what they reported; so do squad members.

Model policy

A project can carry an advisory model policy for actor, evaluator, and interviewer roles. The client honors it through whatever routing capability it provides; separate calls/contexts are required even when one model serves several roles. The round records what actually ran.

Harness notes

Before casting a round, SQUAT skims your workspace's harness notes — operating lessons about running the tests themselves: a scenario with inconsistent results, a staging server that needs a minute to warm up, or a login quirk. Notes are about running SQUAT (distinct from findings, which are about your product); new lessons get saved after the round, and stale ones get retired during maintenance.

Starter scenarios and frameworks

No scenarios yet? The starter library ships eight generic UAT passes ready to import — first-visit orientation, core task completion, returning-user pickup, error recovery, import/migration, find-help-and-use-it, mobile/small-screen pass, and destructive-action safety — plus two scoring frameworks: the full five-dimension rubric, and a minimal pass/fail alternative. Say “show me the starter scenarios” and import what fits; imported copies are yours to edit.

When a session goes wrong

A deploy that lands mid-round, a harness error, wrong test credentials — sessions get compromised. When that happens, the result (or the whole round) is invalidated with a plain-language reason: it stays visible in the ledger, marked, but is excluded from every trend and comparison. Nothing is deleted, and there's deliberately no un-invalidate — the record of what went wrong with the harness is real information too. See Findings & scores.