Skip to main content
Reply quality grades every reply that goes out on an app — written by MAX through an automation, or typed by a teammate — against a rubric you define. The monitor grades only replies that actually published to the store. AI drafts written by an automation are also checked against the same scorecard before they publish; see Checks before publishing. Open Reply quality in the sidebar to reach its three tabs: Overview (this page), Flagged, and Scorecards.
The Reply quality overview: a verdict sentence, three stat tiles for pass rate, voice match and flagged replies, and the coverage footnote
Reply quality is available on the Pro and Ultimate plans (an active trial counts). Accounts below Pro see a blurred preview of this page with an upgrade prompt.

Before anything grades

Reply quality needs two things to produce a verdict: a scorecard for the app, and scoring turned on for it. Until a scorecard exists, the Overview shows the same page you’d see once you set one up, sample-filled and blurred, with a button to create one. Once a scorecard exists but scoring is paused, the page tells you scoring is off and links to the switch. A reply doesn’t grade instantly, either. Each one clears a short quiesce window after publishing first, so a reply you edit seconds after sending isn’t graded twice against two different drafts. Only replies from the last 90 days are picked up, and only ones published to a real app store integration are eligible — nothing in a draft or preview state is graded.

Checks before publishing

When an automation writes a reply with AI and the app has a scorecard, the draft is checked against that scorecard before it goes to the store. You do not need to turn anything on for this.
  • A draft that passes publishes as normal.
  • A draft that fails goes back to the writer once, with the check that failed and why, the exact sentence that caused it when there is one, and the text of the rule it broke for a custom check. The writer keeps the review, its first draft and everything it looked up, so the rewrite is a correction rather than a fresh guess.
  • The rewrite publishes. If it still fails a check, it publishes with a flag naming the problem and appears in Flagged.
The language check and the scorecard checks share that one rewrite, so a reply never goes through a third draft. Output that must never reach the store, such as leaked model reasoning or placeholder text, is always blocked, however many attempts it takes. Replies a person writes, and drafts waiting in approval mode, are not checked before publishing. The monitor grades them once they publish.

The verdict and the three tiles

The top of the page reads in one pass: a sentence stating whether reply quality is holding, mixed, or not holding, followed by a link naming the single rule doing the most damage when one exists. Below it:
  • Pass rate — the share of graded replies at or above the scorecard’s pass mark
  • Voice match — whether replies in this range sound like this app’s own history (see below)
  • Flagged this range — how many graded replies broke a rule and still need a look, linking straight to the Flagged queue
A muted line under the tiles discloses what the numbers leave out: replies still being graded, replies outside the grading window, and replies you’ve already marked handled. Reply quality never rounds a thin sample into a verdict — with too few graded replies in range, the sentence says so instead of guessing.
A reply’s score is a weighted average over the criteria that reached a verdict; a criterion that abstained (missing setup, or nothing to judge) is left out of that average rather than counted as a fail. If any critical criterion fails, the reply is marked failed outright regardless of its numeric score.

Quality trend, by rule or by language

The trend chart plots the daily pass rate over the day each reply was published, not the day it was graded. A day built from very few graded replies renders as a hollow dot rather than a solid one. A rose ring marks a day that contains a reply that broke a rule. The segmented control above the chart splits the same trend three ways: All, By reply rule, and By language. Splitting by rule draws one line per automation that sent replies in range, plus a line for replies a person typed manually.
Quality trend split by reply rule, one line per automation plus a line for manually-written replies
Splitting by language draws one line per language the reviews were written in, using the same flag-plus-name labeling used throughout the feature.
Quality trend split by language, one line per language with its flag
If the AI grader itself changed underneath the numbers — a different model or a different grading prompt version graded part of the range — a banner appears under the charts naming what changed and when, with a link to the grading breakdown on the scorecard. That’s measurement drift, not a real change in your replies, and the banner exists so you don’t mistake one for the other.

Voice match

Voice match is the one signal on this page that isn’t a rule: it measures how closely a reply’s language matches this app’s own historical reply voice, per language, using embeddings rather than any written rubric. A language needs at least 30 of its own replies embedded before it has a baseline to compare against; until then that language shows as a progress ladder (“English 24 of 30”) instead of a chart.
Voice match: a deviation chart per language against the app's own baseline, with a Voice match tile reading Sounds like you
Once a baseline exists, the chart plots each day’s deviation from it rather than a raw similarity number, so languages with very different baseline tones stay comparable on one axis. The Voice match tile reads Sounds like you, Drifting, or Learning your voice. Once something has actually been measured, it also links to the same Flagged queue, ordered by the replies least like your own history.

The worst replies in range

Below the trend and voice charts, a small teaser shows the worst replies currently flagged and unhandled in this date range — the ones that broke a critical rule, or scored lowest — as a preview of the full Flagged queue.
A teaser of the three worst unhandled flagged replies in range, each showing the review, the score, and which criteria it broke
It only ever shows unhandled work: once every flagged reply in range has been marked handled, this section is replaced by more room for the voice chart rather than showing an empty list.

Flagged queue

Triage the replies that broke a rule and still need a look

Scorecards

Build and edit the rubric your replies are graded against