Loading. Please wait...
Reply quality grades every published response to app review against a scorecard you define, cites the exact words that broke a rule, and refuses to guess when it cannot judge. Available today on Pro and Ultimate.
Manual app store review management had one thing going for it that nobody wrote down, because nobody had to. A person read every single reply before it went out.
At ten replies a week you keep that. At eight hundred you cannot, and if you ship in a dozen languages you cannot even in principle. You are not going to read a Polish reply and catch that two of its words are Czech. Nobody on your team is either.
That is the trade nobody names when they turn on automated app review responses. Response time collapses from four days to four minutes, your rating stops sliding, and the review queue that used to eat every Monday empties itself. What quietly leaves with the manual work is the proofread. Quality stops being observable at exactly the moment volume makes it matter. Not "is AI good enough", which is an argument nobody wins in the abstract, but "how would I even find out".
It matters because a reply is not a support ticket. It is public, it is permanent, and it sits under the review for as long as the review does. Anyone deciding whether to install your app reads both. Your replies are app reputation management whether you treat them that way or not.
Every other tool in this category stops at helping you publish faster. AppReply is the first to close the loop on the brand-critical half: what actually went out, graded against the standards you set, in every language you ship in.
Reply quality is available today on Pro and Ultimate. It grades every published reply against a scorecard you define, shows you the exact words behind every verdict, and puts the failures in a queue built to be worked through.
Turn scoring on for an app and every reply that publishes to the App Store, Google Play, Samsung Galaxy Store or Huawei AppGallery gets graded. The results land on one page.
Reply quality dashboard with the pass rate, voice match, flagged count and the daily pass-rate trend
The top line is a sentence, not a number: Reply quality is holding. 21 replies stated facts we cannot back. Under it sit the three things worth knowing at a glance, your pass rate against your own pass mark, how closely replies still sound like you, and how many replies need a look.
A scorecard is a list of checks with weights, and it is yours. It is where your app review response guidelines stop being a document somebody wrote once and start being something enforced on every reply. Some checks are rules evaluated in code: does the reply reference the wrong app store, does it end with your signature, does it apologise when your policy says never admit fault. Others need a model to read the reply: does it answer the concern the reviewer actually raised, does it invent a refund window nobody promised, does it stay in the reviewer's language.
Scorecard settings showing twelve checks grouped into accuracy, language, your instructions and hygiene, each with its weight and whether it is critical, and markers on the ones still waiting on setup
Twelve checks ship by default. Seven are critical, which means failing one fails the reply outright whatever the rest scored, because "polite, well formatted, and factually invented" is not a pass. Edit the weights, turn checks off, or write your own.
Three of the twelve need something from you before they can grade anything. The signature check needs your signature. The apology check needs your list of phrases. Until those are filled in, those criteria return not scored, never fail, and the scorecard page says so out loud, including how much scoring weight is currently inert. A monitor that fails every reply the day you switch it on is a monitor you turn off by Wednesday.
An AI that grades an AI just moves the trust problem up one level. If you could not verify the reply, why would you verify the verdict?
So a verdict is not allowed to be an opinion. Every pass and every fail has to quote the text it is judging, verbatim, and for a fail the quote must contain the words that actually break the rule. A verdict the grader cannot point at in the text is discarded rather than recorded.
Reply detail with the offending sentence highlighted in the published reply, the failed criterion named critical beside its reason, and the checks that passed listed underneath
That reply reads well. It is in the right language, it is polite, it answers the reviewer. The one sentence highlighted in it is a claim about how the game works that nothing in the account backs up, and the verdict says so in three words: unsupported claim. Every other check on that reply passed, which is why this is the kind of defect that survives a human spot check. Everything about it looks fine except one sentence.
This is the part we would defend hardest. An unverifiable quality score is not a quality score, it is a second thing to trust.
Flagged replies get their own page, grouped by what broke rather than dumped in a list.
The Flagged page showing 96 failed replies filtered by criterion, each card carrying its rating, score, language, date and the rule it broke
The chips across the top are the actual failure modes in your account this month, 49 replies that stated facts you cannot back and 47 that never gave the reviewer an actionable step. Click one and you see only those. Every card carries the star rating, the score, the language, the date, and the rule it broke, so you can triage without opening anything.
Sort order is worst first, not newest first, and one-star reviews rise on their own because a weak reply to an angry reviewer costs more than a weak reply to a happy one. That is the same priority order as responding to negative app reviews by hand, applied to the replies you already sent.
Replies you have dealt with move to Handled and stop showing up in the count. The queue is meant to reach zero.
Rules catch defects. They do not catch your replies slowly stopping sounding like you.
Brand voice is the part of app store review management that is hardest to audit and easiest to lose. It is not a rule you can write down, it is a hundred small habits: how much you apologise, whether you use the customer's name, how you close, how formal you get when somebody is angry. Automation reproduces whatever it was pointed at, and it drifts.
Voice match compares each reply against your own historical voice, measured per language against that language's own average rather than against one global blend, and tells you what share of replies it could actually measure. In the screenshot it reads Drifting at 64% coverage, which is a real state and not a failure. Per language matters: a team whose English replies are perfectly on brand can be drifting badly in German, and one global average hides exactly that.
Only human-written replies enter that baseline. Feed generated replies into the standard they are measured against and the model's own output slowly becomes the definition of correct, which is drift wearing the costume of consistency.
Two decisions behind this release are worth stating plainly, because both came out of getting it wrong first.
The trend chart plots pass rate, not an average score. The first version plotted the median score per day. Take a day with four graded replies scoring 100, 100, 100 and 0. The median is 100. The day one reply went out in the wrong language rendered as flawless, indistinguishable from a perfect day, because a median does not move for rare events and rare events are the entire point. The chart now plots the share of replies that passed, so that day reads 75 percent and dips. If your quality metric cannot fall on the day something breaks, it is decoration.
A pass computed around a hole is not a pass. Scores are computed across the checks that produced a verdict. During an outage of the model behind the AI checks, the code-based checks kept answering fine, the score renormalised over the weight that had responded, and every reply graded in that window stored as a pass at 100. An outage did not make the dashboard look broken. It made it look perfect. Worse, the retry path only picks up replies marked not scored, and those inflated passes were not marked not scored, so nothing ever went back to regrade them.
If any check now fails for infrastructure reasons, the reply is stored as not scored with the reason kept, and it gets picked up again once the outage ends. A failure still stands, because the words that broke the rule are on the record regardless of what else was unreachable.
We are telling you about our own bug because it is the exact failure mode you cannot detect from the outside. Any vendor's quality number looks fine right up until the moment it is measuring nothing.
Reply quality is available on Pro and Ultimate, under Reply quality in the sidebar. It grades replies after they publish, so it never sits in the path between a review arriving and your answer going out, and it works the same whether the reply was written by an automation or typed by a person. It sits beside Performance dashboards, which answer whether your replies move your rating: that is the outcome, this is the craft. If your team owns the support queue rather than the product, AppReply for customer support is the shorter tour.
Nothing starts on its own. Adding an app creates no scorecard and starts no scoring, because grading spends money on model calls and that decision is yours. Pausing it stops new grading immediately.
The screenshots here come from a demo account with invented reviews and replies, including the language-mixing defect, because publishing a real account's quality failures is not something we would do to a customer.
If you are already sending automated replies to app store reviews, the question is not whether some of them are wrong. Some of them are. The question is whether you would find out.
New to this? Start with our guide to app store review management, which covers ratings, responses, and building a review workflow from scratch. Or see Reply Quality and what it checks.
Bring monitoring, full Analytics, MAX, and Reply Quality into one app review workflow.

Apply your existing auto-reply rules to featured reviews with one toggle on Google Play and the App Store, and get faster, higher-quality AI replies after our upgrade to GPT-5.6.

Three new things in AppReply: Performance dashboards that show whether your replies move your rating, custom AI reply disclosure, and stronger reliability for 1,000+ reviews a day.