Concept design, entering validation

HCI 271 Capstone 1 - Activision + GUII Lab

Activision Playtest Triage Dashboard

UX researchers at Activision have an AI that detects anomalies in playtest sessions, but no interface to interpret, trust, or act on what it finds. We designed one.

Watch the presentation
RoleProduct Designer
TeamEesha Gupta, Tanvi Reddy Kamanuri
MentorsEmily Chen, Ahmad Azadvar, Vicky Ni
Activision playtest triage dashboard hero

The problem

An AI that detects, but no interface to interpret, trust, or act on what it finds.

The users affected are UX researchers at game studios who run playtest sessions, analyze player behavior, and present findings to game designers and producers, often under extreme deadline pressure.

The AI already flags moments of potential frustration during a session. The problem is what happens next:

  1. Read study plan and set scope
  2. Watch playtest session video
  3. Manually flag key moments
  4. AI flags the same session, with no explanation attached
  5. Re-watch the video to verify the flag
  6. Synthesize into themes
  7. Challenge or accept the flag, with no record of why
  8. Present to stakeholders

The AI generates hundreds of signals. Researchers still do the verification work manually. The result is high cognitive load, reduced trust, and time spent validating AI output instead of generating insight.

100s
Flags generated per playtest session
0
Tools that expose the AI's reasoning
15
Days lost per study to manual verification

Why now

Studios are rapidly deploying AI to detect player frustration and reduce cognitive load, but without researcher-facing interpretation tools, those insights are unlikely to be trusted or adopted.

If nothing changes, researchers keep spending two weeks per study verifying AI claims they can't trust, and the AI investment goes unused. One bad experience is all it takes to abandon the tool entirely.

Competitors and gap

Every tool solves part of the problem. None of them show their reasoning.

CompetitorKey strengthCritical gap
Playtest CloudFast setup, game-specific automated analysisBlack box, no reasoning shown
Lysto.ggMost sophisticated AI in the space for playtest analysisNo mechanism to trust or interact with the AI
HotjarStrong behavioral visualization: heatmaps, session recording, funnel analysisNot game-specific, no frustration detection
Our tool bridges the gap between AI detection and human decision-making by showing why the AI made a prediction, how certain it is, and how researchers can improve it.

The research

Five interviews. Thirty-four codes. Zero assumptions.

We talked to 2 UC Santa Cruz researchers and 3 Activision practitioners, spanning game telemetry, data science, UX research, and telemetry infrastructure. Sessions ran 30 to 45 minutes over Zoom, with 2 researchers present and standardized notes.

We independently coded every transcript into 34 codes, then affinity-mapped them into 10 themes.

01

One bad AI experience eliminates tool usage

The user will refrain from using the tool after one bad experience or fallout.
02

Every session starts with a specific question

A dashboard that surfaces all flags equally without context filtering feels overwhelming and irrelevant.
03

Override should be a conversation

Researchers do not want to delete a wrong flag. They want to understand it well enough to challenge it.

The researcher journey

Where trust breaks down across one playtest study.

1

Presession setup

Read the study plan, set research questions, set up tools.

"What's my question for this session."
2

Data quality check

Review data schema, confirm sample, run outlier detection.

"Is this data clean enough to use."
3

Flag review and analysis

Watch the video to understand the AI flag, check whether flags recur across the session.

"AI says 17 people mentioned this, when no one did."
4

Synthesis and insights

Cluster themes manually, cross-check every claim, rewrite AI language entirely.

"It was faster to read the comments by myself."
5

Stakeholder presentation

Trace everything back to raw data, present findings, defend methodology.

"Did this come from me or AI."

From what we heard to what we built

Three pain points. Three concepts.

Researchers are overwhelmed by flags that have nothing to do with their question.

Research Question Upfront

No reasoning shown, so nothing to trust or challenge in what the AI flagged.

Conversational Challenge Panel

No fast way to find past flags or navigate the dashboard's features.

Smart Search Bar and Chat History

Concept 01

Research Question Upfront

A collapsible prompt gates the flag list behind a research question. Only relevant flags surface; out-of-scope ones dim.

Provides the user a filtered overview of relevant flags, dismissing the ones which are not useful, reducing the cognitive load of skimming through hundreds of flags.

A single playtest session can carry around 1,000 raw data points. The research question is the first filter, cutting that down to what's actually relevant before the researcher sees anything. Once inside the dashboard, a second filter icon lets them narrow further by category and other facets. The goal at every step is the same: fewer results, more relevant ones, less to sift through.

1User adds a research statement and selects a playtest theme.
User adds a research statement and selects a playtest theme
2Opened graph view, with collapsed fields to edit the research question and playtest theme.
Opened graph view, with collapsed fields to edit the research question and playtest theme
3Collapsed graph view, with the same fields still available to edit.
Collapsed graph view, with the same fields still available to edit

Concept 02

Conversational Challenge Panel

Override opens a dialogue space. When a researcher disagrees with a flag, they have a conversation with the AI, which produces a suggestion and logs the outcome.

Gives the user more control by letting them understand the flag better through a conversation with the AI, prompting it to correct or override the log.
1User lands on the dashboard: all flags, confidence scores, and the action menu.
User lands on the dashboard: all flags, confidence scores, and the action menu
2Flag detail view: telemetry data, anomalies, and more.
Flag detail view: telemetry data, anomalies, and more
3Conversation with the AI, then override or keep the flag.
Conversation with the AI, then override or keep the flag

Concept 03

Smart Search Bar and Chat History

An always-visible search bar helps researchers navigate the dashboard's many features and filters, the dashboard is large enough that finding the right screen isn't always obvious. Chat history lives in its own icon beside the search bar, letting researchers retrace past conversations without starting over.

Acts as a directory for the dashboard, letting users find any information quickly without digging through complex pages.
1The search icon is always available at the top of the dashboard.
The search icon is always available at the top of the dashboard
2A query redirects the user straight to the right information.
A query redirects the user straight to the right information
3Chat history opens from its own icon beside the search bar.
Chat history opens from its own icon beside the search bar

The problem

UX researchers at Activision have an AI that detects, but no interface that helps them interpret, trust, or act.

What we learned

Explanation panels are not a feature. They are the mechanism through which the dashboard earns the right to be used at all.

What we built

Three connected concepts, context filtering, conversational override, and smart navigation, grounded in five interviews and zero assumptions.

Next: a 10-week sprint

Validation and refinement, before this becomes a working prototype.

Wk 1-2

User testing

Test paper prototypes with Activision researchers. Document findings.

Wk 3-4

Lo-fi digital

Build wireframes in Figma with corrected designs.

Wk 5-6

Heuristic review

Peer evaluation using Norman's 10 usability principles.

Wk 7-8

Mid-fi prototype

Interactive Figma prototype with real synthetic data.

Wk 9-10

Testing, round 2

Second round with Activision researchers. Iterate.

Biggest risks

  • Limited Activision access: schedule sessions early through sponsor contacts.
  • AI capability gap: some concepts may not be technically feasible; still exploring with Activision stakeholders.

Success metrics

  • Challenge a flag via the conversational panel in under 1 minute, without assistance.
  • Find a dashboard feature via smart search in under 30 seconds.