Loading. Please wait...
Mobile app reputation management was built for a store that showed one number. Google Play and the App Store now read review text, and your developer replies sit in that text under every review. In a pilot read of 100 top Google Play apps, about 1 in 9 low-star reviews had a reply and 42% of replies repeated word for word. This guide covers what the stores read, the four ways a review response fails, what to do when your app store rating drops, how to ask for ratings without asking in replies, how iOS and Android differ, and how to measure whether any of it worked.
In September, Den Markov published a breakdown of the rubric Google's app quality audit reportedly applies to Play Store apps. One of its scores is called perceived quality from reviews. It does not grade the star average. It reads the text: which complaints repeat, which praise repeats, whether the recent mood has drifted from the historical rating, and whether the developer's replies are personal or templated.
The case study in that breakdown is an app with a 100% reply rate. Every review answered. It scored medium on reviews, because every one of those replies was a template.
The reviews score in the reported Google Play audit, and the templated case study at Medium.
Google has not published this rubric, so hold the detail loosely. But the direction matches everything else the stores have done in the past two years, and it changes what app reputation management is for.
Start with who reads the page. Above your reviews on the App Store there is now a paragraph an AI wrote about your app, drawn from what people said. Apple has shipped these summaries since iOS 18.4, and Google Play has been rolling out AI summaries of its own. Neither counts your stars. They read the words, and what repeats gets weight.
Google tells developers what it expects from a reply. Its Play Console guidelines ask for replies that address the reviewer's comment in a clear, valuable and truthful way, stay respectful, and avoid promotions and solicitations. None of that is about speed or coverage.
The breakdown of Google's audit goes further. Its reviews score reportedly weighs recurring pain points, recurring praise, the gap between the historical rating and recent sentiment, the reply rate, and whether the replies are personalised or templated.
A separate trustworthiness score looks at the tone and competence of developer responses, and at small signals a reader also notices, like a support address on a free email domain instead of your own.
The 13 scores in the reported audit. The two in violet read developer replies.
The fixes it recommends for the templated app are specific: cluster the complaints and fix the top two, and write personal replies to one-star reviews that name the bug.
Whether or not every line of that rubric is accurate, it describes what a person scrolling your reviews already does. They read a complaint, they read your answer, and they decide whether anyone on your side read it first. A store that summarises reviews with a language model is doing the same thing at scale.
That is the shift this guide is about. When the text is what gets read, a reply is not a box to tick. It is writing you publish, in your company's name, next to the words it answers.
Mobile app reputation management is the work of looking after everything publicly attached to your app on Google Play and the App Store: the star ratings, the review text, and the developer replies you publish under them. It sits next to app store optimization rather than inside it. ASO works on the listing you control before the install, the title, keywords and screenshots. Reputation management works on the record users read alongside it, and that record feeds back into ASO, because rating and review sentiment move both search ranking and install conversion.
For most of the last decade, mobile reputation work reduced to one number. The store showed a star average, the average moved slowly, and the playbook followed from it. Watch the rating. Answer quickly. Answer as much as you can. Get the cost of answering down. That playbook worked, and teams that ran it had better ratings for it. Our own product was built for that world.
Your reputation is no longer only the number customers give you. It is the public record of what they said and what you said back. Customers write the first half. You write the second, one reply at a time, and it stays under the review for as long as the review is there.
The first half is well understood. Sentiment analysis and topic extraction over ten thousand reviews is now fast and cheap, and it is still the foundation: you cannot answer a problem you have not found. The second half is the one almost nobody measures, which is odd, because it is the only part of your store listing you actually wrote.
We wanted to know what that writing looks like in practice, so this month we ran a pilot read of Google Play. We took 100 apps from the United States top free and top grossing charts, read every review dated in a 30-day window across ten review languages, and graded a sample of more than 1,800 developer replies. It is a pilot, the apps skew large, and the grading model is not yet validated against human labels, so read the figures as a well-measured first look.
About 1 in 9 low-star reviews had a reply a week later. Across 89,330 one to three star reviews with real text, 11.5% had a developer answer. Apps split hard into two camps: of the 76 apps with enough reviews to judge, 24 never answered a single low-star review and 18 answered nine in ten or more. The largest apps answered least. The median app with 500 million installs or more answered 0.9% of its low-star reviews.
Reply coverage by star rating, pilot read of 100 top Google Play apps.
When a reply did come, it came fast. The median gap between the review and its reply was 6.8 hours, 61% of replies arrived within a day and 81% within three days. Speed is not the bottleneck. Measured against every low-star review, answered or not, 9.3% got a reply within 72 hours, because most never got one at all.
Apps cluster at the two ends: no replies at all, or near-full coverage.
Answering did not mean saying something. 42% of all replies in the window repeated another reply from the same app word for word, and the real share of templates is higher, because a template that inserts the reviewer's name counts as unique. One reply was posted 1,604 times in 23 days.
Among graded replies to low-star reviews, about 3 in 4 addressed what the reviewer actually raised. The rest guessed at a fix for a vague complaint, thanked someone who had reported a bug, or answered a different problem from the one in the review.
And 13% of the replies we could assess were in a different language from the review, almost all of them English answers to reviews written in something else.
Now put a realistic mobile team behind those numbers. A team publishing 800 replies a week would need most of a working day, every week, just to read them. Across a dozen languages the reading stops being possible at all. At that volume a 2% error rate is more than 800 public mistakes a year, each one carrying your company's name and sitting under the review it got wrong.
Reply rate says nothing about any of that. You can reach 100% by publishing the same paragraph every time, and the pilot found apps doing exactly that. If you are starting from the first problem, reviews sitting unanswered, our guide to app store review management covers the workflow. The rest of this guide is about the second problem.
Every bad developer reply we have seen fails in one of four ways, and they are not equally bad. We rank them in the order support leads worry about them, because the question they ask is not "what is our average score" but "did we embarrass ourselves in public".
The worst failure is a wrong fact. The reply states something about your product that nothing supports: a refund window nobody promised, a feature that does not exist on that platform, an explanation of how the app works that sounds plausible and is not. It is rarely a lie. It is a confident guess, and in public it reads the same as a lie.
Next is the wrong language. A reply in English under a Turkish review tells the reviewer you did not read theirs. A Polish reply with two Czech words in the middle is worse in a quieter way: close enough to right that nothing flagged it, close enough to wrong that a native speaker reads it as broken.
Then come your own instructions. Every team has rules, written down or not. Never admit fault. Never promise a date. Always route billing to the help center. Sign off with the team name. A reply that breaks one is not necessarily wrong for the reviewer, but it is wrong for you, and it is the kind of thing a legal or support lead finds six months later.
Last is hygiene: the wrong store named in a Google Play reply, a missing signature, a reply four times longer than the review, leftover template markers. Cheap to catch and cheap to fix, and it still reads badly.
Reply detail with the offending sentence highlighted in the published reply, the failed criterion named critical beside its reason, and the checks that passed listed underneath
Share of graded replies that answer the review, by star rating.
The pilot read graded replies on exactly that first question, whether the reply addresses what the reviewer raised, and the answer depends on how hard the review was. Praise is easy to answer. A one-star complaint is where replies miss, and it is also the reply most people will read.
Here is what passing looks like. A two-star review of a mobile game in the pilot read:
It's a clone of that same base building game we've all played over and over. Fun for about 10 hours. The sniper action mini games are great, though.
The reply:
Great to hear the sniper action mini games hit the mark for you. We understand the rest felt like a familiar base-building formula that stopped being fun after about 10 hours. We'll keep looking at where the game falls short and work on it.
It names both halves of the review in the reviewer's own terms, owns the criticism without arguing, and could not be pasted under any other review. That last test is the one to run on your own replies.
Each of these is invisible in aggregate and obvious in the specific reply. A dashboard that says 97% of replies are fine tells you nothing about the 3%, and the 3% is the whole subject. That shapes how you check.
Some checks need no model at all. Whether a reply names the wrong store, carries the right signature, or apologises where your policy forbids it can be settled by plain rules. Whether a reply answers the concern the reviewer raised, invents a commitment, or stays in the reviewer's language needs something that reads. Whichever way you check, mark wrong facts and wrong language as automatic failures. A weighted average lets a critical defect hide behind five things that went fine.
And make every judgement quote the words it is judging. "This reply mixes languages" is a claim you have to trust. The same verdict with two Czech words highlighted inside a Polish sentence is something you can see without reading either language. A verdict that cannot point at specific text is an opinion, and it should be thrown away rather than recorded.
You can run all of this by hand before you buy anything. Write down what a good reply means for your team, five or six checks with the fatal ones marked. Read your last ten replies in a row, and if you cannot tell them apart, neither can anyone else. Grade twenty against your list and record the pass rate as a baseline.
Then sort the failures by what broke rather than by date, wrong facts first, and within each group put the one-star reviews at the top. That is the order to fix them in, because a wrong fact under a one-star review is the reply most likely to be read and least likely to be forgiven.
There are two questions, and you need both answers.
The first is review response quality: was the reply good, judged against your written standard, with the offending words quoted. The second is reply effect: did the star rating move on reviews that received a developer reply? There is good reason to expect it can. Google says in its Play Console guidance that responding to a negative review can raise that rating by 0.7 stars on average. A 2018 study by Hassan and colleagues found that 4.4% of reviews with a developer reply had their rating raised, against 0.7% of reviews without one.
Measure that movement only on reviews that could move. A review base that is mostly five stars flatters every average you compute, so look at replied reviews below five stars, and compare the movement after an AI reply against the movement after a human one. That comparison is the cleanest read on whether automation is helping or only saving time.
AppReply Performance dashboard showing reply effect for AI vs human, automation score, and review coverage
Read the two numbers together and you get a diagnosis instead of a mood. If quality is up and ratings move, the work is paying, so keep going. If quality is up and ratings are flat, your checklist is grading the wrong things, which you would never see tracking only one of them. If ratings move while quality falls, something is working that your own guidelines do not allow yet, and it is worth going to read those replies. Track only one of the two and you get a number you cannot act on.
Most reputation dashboards track coverage, speed and the star average. Those are the inputs. The table below is the set we use, with the trap each one hides.
| KPI | What it tells you | The trap |
|---|---|---|
| Reply coverage on low-star reviews | Whether complaints get an answer at all | Coverage across all reviews is inflated by praise. Measure on one to three star reviews with real text. |
| Time to first reply | Whether the answer arrives while the reviewer still cares | Median hides the tail. Report the share answered within 24 and 72 hours against every low-star review, not only the answered ones. |
| Repeated reply share | How much of your coverage is a template | Exact repeats understate it. A template with the name swapped in counts as unique. |
| Pass rate against your written standard | Whether the replies were good | An average score lets a fatal defect hide behind five things that went fine. Count passes, and fail the reply outright on wrong facts or wrong language. |
| Pass rate per language | Whether one market is failing behind a healthy blend | Small languages need a minimum sample before the number means anything. Report "not enough data" instead of a score. |
| Not-scored share | How much of your quality number is real | Anything that could not be checked has to be recorded as unmeasured, never as a pass. |
| Stale reply share | How many replies answer a review that has since been edited | Google Play shows the latest edit date. In the pilot read, one reply in 18 was older than the review it sat under. |
| Reply effect | Whether replies move ratings | Measure on replied reviews below five stars, and compare AI replies against human ones. A five-star base flatters everything. |
The first three are the operational floor and every tool reports some version of them. The next four are the quality set, and they are the ones that separate a team that reads its own replies from one that counts them. The last is the outcome, and it only means anything read together with the quality set.
The moment reputation work gets urgent is a sudden drop in your app store rating. The instinct is to reply to everything fast. The better first move is to read, because the cause decides the response, and app store rating recovery starts with the diagnosis.
A release regression shows up as a cluster of reviews naming the same crash or broken screen, starting on a specific day and often a specific app version. An outage looks similar but hits every version at once and stops when the service recovers.
A pricing or paywall change produces complaints about money and fairness rather than bugs, and those reviews are angrier and slower to change.
Review bombing, where a group reviews your app over something that has nothing to do with its quality, arrives as a spike of near-identical short reviews, often from new accounts. Games see it most, but any app with a public decision to make can be hit.
Look at a short window to find it. A spike that started three days ago disappears inside a ninety-day average, so the longer the range you check, the calmer a live incident looks. Alerts on a sudden rise in a specific complaint catch this faster than anyone refreshing a dashboard, which is why monitoring is the first thing to get right.
Once you know the cause, the replies write themselves in the right direction. For a regression, name the problem, say it is fixed or being fixed, and give the version or date. Do not paste the same sentence under all three hundred reviews. The facts are the same, but a reviewer who lost data and one who was merely annoyed need different answers.
For reviews that need account access, like billing disputes and lost purchases, route them to your helpdesk and answer publicly with where to go next. Our Zendesk and Intercom integrations create the ticket for you and bring the agent's reply back to the review.
For review bombing that breaks store policy, report the reviews through the store's flagging tools rather than arguing with them in public.
The same goes for fake reviews, whether bought for a competitor or against you. Both stores remove reviews that break their policies, and neither removes a review for being negative. The only reviews you can get taken down are the ones that were never legitimate. Google's audit rubric reportedly treats suspected review manipulation as a policy matter rather than a score, which is a reason to report it yourself before the store finds it. Everything else stays up, and the reply under it is your only say.
Then come back. The reviewers most likely to change their rating are the ones who complained about something you then fixed, and they only find out it was fixed if you tell them under their review once the fix ships. That second reply is where rating recovery actually happens.
The other lever on your app store rating is who leaves one. Happy users rarely go looking for the review screen, so the ratings you collect skew toward people with a problem unless you ask the others.
Both stores give you a sanctioned way to ask. On iOS the SKStoreReviewController prompt appears at most three times a year per user, and Apple decides whether to show it. On Android the Google Play In-App Review API works the same way under a quota Google does not disclose. Neither lets you ask only happy users first, and both reward timing.
Ask right after a user succeeds at the thing they came for, not on launch and not in the middle of a task. The breakdown of Google's audit makes the same point from the other side: an app that asks for money or attention before it delivers value loses on almost every score.
What not to do is ask in the reply. In the pilot, 12 of the 58 graded apps asked at least one reviewer to raise their rating, one of them 313 times with the same sentence. It reads as the reply's real purpose, it sits badly with Google's guidance against solicitation in replies, and it turns an answer into a request. Answer the review. If the fix is real, the rating follows, and the reviewer gets a store notification that you replied.
The same developer reply behaves differently on iOS and Android, and a few of the differences change how you should write.
On the App Store you reply through App Store Connect or its API, one response per review, and you can edit it later. The reviewer is notified and can update their review. Reviews are grouped by storefront country, so a single app can carry a very different reputation in Germany and in the United States, and your replies should follow the market they sit in.
On Google Play you reply through the Play Console or its API, with a limit of 350 characters, which forces specificity whether you want it or not. Reviews are grouped by the language they are written in rather than by country.
Google Play also shows the date of a review's latest edit, not when it was first written, and reviewers edit often. In the pilot, about one reply in 18 was dated before the review it sat under, which means it answered a version of the review that readers can no longer see. When a reviewer edits, read it again and update your reply.
Reply coverage by the language the review was written in.
That language split is where reply coverage quietly diverges. In the pilot read, German and English reviews were answered about three times as often as Korean and Indonesian ones, and when a reply was in the wrong language, 96% of the time it was English. If your team reads English, your English reviews are the ones getting answered and checked.
Samsung Galaxy Store and Huawei AppGallery follow the same basic pattern, one public reply under each review, with their own consoles and their own audiences. If your Android app is in those stores, those replies are part of the same record, which is why we added both alongside Google Play.
On every store, a reply stays up until you change it. You can edit or delete one, but you will only know which of the thousands you published need editing if something is reading them. Treat published replies as permanent by default, because in practice the ones nobody flags stay exactly as they were written.
We built our products around both halves of the record.
Review analytics covers the first half: the recurring complaints and praise, per app and per language, with alerts when something spikes.
Reply Quality covers the second. It grades every reply your team publishes on the App Store, Google Play, Samsung Galaxy Store and Huawei AppGallery against a scorecard you write, whether an automation wrote the reply or a person typed it.
Every verdict quotes the words it judged, failures land in a queue grouped by what broke with the worst first, and brand voice is tracked per language against your team's own past replies. It runs after publishing, so it never delays an answer, and nothing is graded until you turn it on. It is available on the Pro and Ultimate plans, and the launch post walks through it in detail.
MAX, our reply agent, works on the cause. Before it writes, it reads what the reviewer raised, searches the replies your team already approved, and checks the answer against your live help docs and release notes, then drafts in the reviewer's language across more than 100 of them. It is hardened on more than 100,000 published replies and included on Pro.
Performance dashboards close the loop by showing whether any of it moved your rating, split by AI and human replies.
For more depth on each part, read why canned replies are not review management, closing the loop on review analytics, and publishing replies in languages nobody on your team reads.
Whatever you use for mobile app reputation management, the rhythm matters more than the tool. Daily: answer the low-star reviews that arrived, and read the ones that spiked. Weekly: pull twenty published replies and grade them against your written standard, then look at the flagged ones worst first. Monthly: compare pass rate and reply effect per language against last month, and re-read replies under any review that was edited since. Quarterly: rewrite the standard itself, because the replies that failed will tell you which rule was missing.
Start this week with the version that costs nothing. Read ten of your replies in a row. Count the languages you publish in against the languages anyone at your company reads. Grade twenty replies against a written standard. The number you get will tell you more about your app's reputation than the star average on your listing.
Bring monitoring, full Analytics, MAX, and Reply Quality into one app review workflow.

Reply Quality grades every published app store reply against a scorecard you write, and quotes the exact words that broke a rule.

Apply your existing auto-reply rules to featured reviews with one toggle on Google Play and the App Store, and get faster, higher-quality AI replies after our upgrade to GPT-5.6.