SQUAT User ManualCore concepts

Getting started

Core concepts

SQUAT has a small vocabulary that everything else builds on. Ten minutes here and the rest of the product — and this manual — will read naturally.

The big picture

The full arc of using SQUAT looks like this:

  1. Set up
  2. Recruit a squad
  3. Write scenarios
  4. Run rounds
  5. Measure findings
  6. Track issues
  7. Review trends
  8. Maintain & export

Every stage is optional in the sense that SQUAT still works if you skip it — but each exists because skipping it costs you something real later: the ability to attribute a change, to trust a trend, or to believe a “fixed” claim.

Your squad

Squad member (persona)

An AI-performed tester persona. It may represent a human, be deliberately synthetic, or combine permitted sources. Its portrait keeps claims, inferred motivations, conditional behavior patterns, tensions, and knowledge limits distinct; an approved communication range defines how they talk without demographic caricature. A provenance-excluding actor briefing is compiled for rounds without raw sources or real/synthetic labels.

Source basis and origin

Origin records acquisition workflow: interview, template, or import. Source basis records what contributed: first-person human, documentary, synthetic, or unclassified; combined is a derived operator-facing label. The two are never substituted for each other.

Learnings

Short behavioral observations from rounds. Learnings create returning-user continuity. When one conflicts with a claim or prior pattern, SQUAT preserves the tension for operator review rather than treating either side as ground truth.

Squad (named line-up)

A saved selection of squad members you reuse round after round — “Mobile squad,” “Checkout crew.” A project can name one squad as its default cast. Once it does, rounds, panels, and steward work use that line-up or a deliberate subset; SQUAT never silently pulls in someone from the wider roster.

What gets tested, and why

Project

The product under test. A workspace can hold several products; each round belongs to a project so history stays attributable and filterable per product. A project can also carry a default squad, a standing default mission, an advisory model policy (which AI model should play testers vs. evaluate them), and a list of core features (problems found there are weighted as more urgent). Login secrets are deliberately not project data: scenarios store only non-secret local-auth policy, while credentials and browser state remain on the operator's machine.

Context

Everything the squad should know about the product itself — what it is, who it's for, how to reach it, the feature map, known limitations. Context is versioned (every save adds a version, nothing is overwritten) and folded into every tester's and evaluator's briefing.

Scenario

The tester's reusable task: a goal stated in the user's own terms — “you have three ideas you don't want to lose” — never UI instructions like “click New Idea.” That distinction is the whole point: a goal lets the persona's own judgment (and confusion) show up; a scripted walkthrough only tests whether an AI can follow instructions. Scenarios are the stable measuring instrument — the same scenario runs across dozens of rounds over months, which is what makes the trend comparable. A scenario may also carry strict non-secret authentication policy (allowed local setup methods and exact origins), never an account or credential.

Moderator instrument

A scenario used for concept interviews or design panels rather than a scored task. It stores purpose and ordered phases, with each spoken question protected separately from moderator-only objectives, probes, and listen-fors. It cannot run as a scored UAT task.

Mission

Your specific question for one round: “shake out the new capture flow before Tuesday's launch,” “focus on mobile this time.” Missions are optional and kept separate from scenarios on purpose — the scenario keeps the long-term trend comparable, while the mission lets a single round answer a timely question without polluting the instrument. Missions are protected like other sensitive content; reports record only that a mission existed, never what it said.

The test cycle

Round

One testing session: one or more squad members each attempting one or more scenarios, scored and recorded. A round captures who ran (clientInfo: the platform and model that played the testers) and can optionally record what was tested (targetBuild: your product's version or build). Together they help answer the later question “did the app change, or did the tester change?” You may skip the target build, especially for a concept panel.

Actor and evaluator

Every persona×scenario attempt runs as two strictly separated passes. The actor plays the persona attempting the scenario in character, narrating what it sees, thinks, and does. The evaluator — a fresh context that sees only the transcript, never the actor's private reasoning — judges whether the goal was actually accomplished and scores impartially. The same context never both plays the tester and grades it: that separation is what makes a pass mean something.

In concept interviews, a third moderator receives the required versioned instrument and dialogue history while the actor receives only the current spoken turn. SQUAT-generated performance bundles exclude source provenance and real/synthetic labels, but SQUAT does not attest that an external model host isolated the contexts.

Result

One persona×scenario attempt's outcome: a status (Pass, Fail, Blocked, or Inconsistency), a failure class on any fail (was it the app, the test design, or the testing agent itself?), evidence excerpts, canonical dialogue references, and optional project-artifact references.

Structured measurement

Findings

Specific problems observed during an attempt, each with a severity from S0 (worst) to S3, set from observed impact — never from how loudly the persona complained. The same underlying problem reported across different rounds groups together automatically.

Scores

Five dimensions per attempt — discoverability, comprehension, task success, effort, and confidence/trust — each rated separately and never blended into one number, because a single blended score hides exactly the information you need.

Coverage

A per-feature record of what each attempt actually touched and how it went, so you can tell “this feature is fine” from “this feature was never exercised.”

Issue

The persistent, cross-round record of a problem. Findings are per-attempt; issues live across rounds. Every finding automatically opens or updates a matching issue. Issues move through a lifecycle — open, fixed (your claim), verified-fixed (proven by a re-test), or regressed (it came back — the loudest alarm in SQUAT). See Findings & scores.

The record

Ledger

Your longitudinal history: every round, every result, every finding, in chronological order — including history you backfilled from before SQUAT. The ledger is the point of the product: a single round tells you what's broken today; the ledger tells you whether you're getting better.

SQUAT holds the record

The research record lives here, not in a folder you have to remember to keep. That means all of it: the transcripts of what your personas actually said, the artifacts and screenshots the round was run against, your harness notes, the drafts you are still working on, and the summaries and reports you finish — Round Summaries, Research Reports, and the UAT Reports you hand to whoever is doing the fixing. You do not assemble the record afterwards from pieces scattered across tools — it accumulates as you work, and it is the thing the ledger is built from.

Anything you or your testers wrote is encrypted in storage under your workspace's own key. An unfinished draft is treated exactly like a finished report, because being unfinished does not make it less yours.

Copies elsewhere — a local folder, a Drive mirror — are yours to choose, not something SQUAT does behind you. Where your client supports it, exports give you the readable current content or the full protected history. A copy is a copy; the record you can cite, compare, and build a trend line from is the one in your workspace.

Workspace

The container for all of it — squad, projects, scenarios, rounds, and issues. Your license key connects your AI client to your workspace. Teammates join a workspace with their own roles and seats (see Roles & security). If your account owns more than one active workspace, the account page lets you choose where new pack capacity belongs and move one whole existing pack between workspaces. SQUAT never guesses a destination, and checks the source's currently measured use against its remaining allowance before a move.

Cards and Markdown

Focused results — workspace status, one squad member, a completed round, an issue, a steward report, an artifact-library view, or next actions — use the same card contract in every client. A client with standard MCP Apps support may draw a rich card. A client without it receives the same facts and action order as complete Markdown. Buttons start an ordinary conversational turn; they never perform a hidden mutation or skip SQUAT's confirmation rules.

A normal workspace card stays operational: workspace-wide counts of projects, personas, squads, scenarios, and rounds, the roster, and the latest round. Account-plan, billing, license, internal entitlement, and human-seat details appear only when you explicitly ask about account or license status.

Cards lead with readable names and titles. Stable record IDs appear only as secondary metadata when they help with disambiguation or a follow-up action. Rich cards place a copy control beside every displayed ID; Markdown fallbacks show displayed IDs as selectable inline code.

An artifact-library card appears as a thumbnail slide-sorter. Click an item to open its preview and details inside the card. Files that cannot be previewed still show their name, type, size, and identity; the Markdown fallback lists the same metadata. Cards never carry temporary download links or storage credentials.

Rich rendering does not weaken protection. Hosted cards accept only strict operational fields. Protected profiles, transcripts, and evidence are not placed in card payloads.

Finding your way

Whenever you feel lost, ask “where am I?” SQUAT answers with a journey card: the whole map of the process you're in, with your position marked. The workspace-setup journey runs from first connection through your first reviewed results; when you decide to run a research round, a round-preparation journey shows exactly what that round still needs before it can start; and a round in progress has its own smaller journey from start through synthesis. Steps you can take by more than one route (recruiting personas versus importing them) show the alternative marked optional.

Your position is always read from your workspace as it actually is right now — never from a script of what you were “supposed” to do next, so the map can't drift out of date. In a client without rich cards you get the same map as one line of text: (1) Start ✓ — ((2)) Create a project ← you are here — (3) Create personas… The journey shows where you are in the whole; when you want a recommendation for the very next move, the next-steps menu still does that job, and may carry a compact copy of the journey strip above its options.

Two more ways to use the squad

Steward

An expert reviewer that acts as your proxy — not a simulated user. Where a persona asks “can I figure this out?”, the steward asks “is this what we meant?” It judges your live product against your recorded intent registry (durable rulings about what each surface is for) and, unlike a persona, prescribes fixes. See Steward & panels.

Design panel

The squad, moved earlier: personas react to a design artifact through a moderator instrument. Panel reactions are qualitative, generate hypotheses, and never enter the scored trend. See Steward & panels.

Principles worth knowing by name

PrincipleWhat it means for you
A test has to be able to failA round where nothing could have gone wrong tells you nothing. SQUAT's discipline is built around accurate, falsifiable results.
Invalidate, don't deleteA compromised session (a deploy landed mid-round, wrong credentials) is flagged with a reason and excluded from trends — but never erased. The record stays accurate.
“Fixed” is a claim, not a resultOnly a re-test by the same persona that found an issue proves a fix. SQUAT tracks the difference.
Provenance stays attachedPersona records retain their acquisition/source context, and dialogue records identify whether their transcript came from the SQUAT round or an archive.
Protected automaticallyThe authenticated service protects customer content before storage. Clients send ordinary semantic values and do not manage content-encryption keys.
Everything is portableA complete export of your workspace is free at every tier, always. Your data is yours.