Screenshots are the single highest-impact conversion lever in ASO. Bad screenshots — raw app screens with no context — convert far worse than well-designed ones, and it is the largest single gap in most listings. Treat the specific multiples quoted around the web with caution: they come from vendor case studies on other people’s apps, and the only figure that describes yours is the one Apple’s Product Page Optimization test returns. Good screenshots tell a story across the visible slots, lead with the strongest value proposition, use captions, and adapt per locale. See the App Audit hub for the full ASO programme.
App Store (iOS): - Up to 10 screenshots per device size - First 3 visible without scroll (most important) - 6.7" (iPhone Pro Max) is reference; other sizes auto-scaled - Portrait or landscape (consistency within app) Google Play: - Up to 8 screenshots - First 2-3 visible in listing preview - Phone, 7" tablet, 10" tablet versions
Slot 1: HEADLINE BENEFIT
Strongest value prop, brief caption
"Track every lead automatically"
Slot 2: KEY FEATURE 1
Specific capability with caption
"AI prioritises hot prospects"
Slot 3: KEY FEATURE 2
Second capability
"Drag and drop pipeline"
Slot 4-6: SECONDARY FEATURES
Other capabilities, social proof
Slot 7-10: TRUST & DEPTH
Awards, reviews, integrations, security
Raw screenshots without captions convert poorly — users don't decode app UI in 1-2 seconds. Captions explain the value:
Bad: [pure app screenshot]
Good: "TRACK EVERY LEAD" [headline]
[app screenshot showing leads]
"Never lose a prospect again" [subhead]
Caption layout:
- Headline: 4-6 words, biggest text
- Subhead: 8-12 words, supporting
- Use brand colours
- Background colour matching brand
- 60-70% screen for caption, 30-40% for screenshot
Each App Store / Play locale gets its own screenshots. Don't just translate captions — adjust visuals for cultural context. Numeric formats, currency symbols, character sets all matter:
US: "$29/month" English caption UK: "£24/month" English caption (different price possibly) DE: "29€/Monat" German caption JP: "¥3,200/月" Japanese caption + JP-localised app UI in screenshot
Slots 4-6 often include:
App Store Connect (iOS 15+) Custom Product Pages allow split tests. Google Play Store Listing experiments too. Test:
The story structure in section 2 is sound, and it quietly assumes something that is often false: that people see the story.
In search results, a listing frequently shows just one or two screenshots, at small size, beside the icon and title. The decision to tap is made right there — before anybody reaches the product page, and before slots four to ten exist for them at all.
Which means slot 1 is not the opening of a narrative. It is a standalone advertisement that has to work alone, at thumbnail scale, in roughly a second, for somebody who has never heard of you.
Design slot 1 to survive on its own: - One claim, in four to six words, legible at thumbnail size - One image that shows the outcome, not the interface - No dependency on anything in slots 2-10 - Test it by shrinking it to 20% and looking again
Then let the remaining slots do what they are good at: deepening the case for somebody who has already stopped scrolling. They are a warmer, smaller audience, and they are not the ones you have to win.
Section 3 says raw screenshots convert poorly, and the reason is worth stating plainly, because “add captions” sounds like a design preference and it is not.
A raw screenshot asks a stranger to look at an interface they have never seen, decode what it does, and infer why it would improve their life — in about a second, while scrolling. Nobody performs that task. They are not being lazy; the task is unreasonable.
A caption does the work for them. It states the outcome in the words a person would use, and the screenshot beneath it stops being a puzzle and becomes evidence for the claim. That is the entire mechanism, and it is why the caption is doing more of the persuading than the image.
Screenshot decisions generate more internal debate than almost anything else in a product team, because everybody can see them and everybody has an opinion. Both stores have made that debate unnecessary.
Apple: Product Page Optimization - Split test up to three variants against your current page - Real store traffic, real installs, a real result - Results in days to weeks depending on volume Google Play: store listing experiments - Same principle, run from Play Console - Test screenshots, icon, description independently
The rule that makes this work: change one thing at a time. A redesign that alters the captions, the ordering, the colours and the device frames simultaneously will tell you that something got better and nothing about what. That is a wasted test cycle, and cycles are the scarce resource here.
And be prepared for the result to be unwelcome. The variant the team likes least wins often enough to be worth planning for, which is the whole reason the test exists — taste is not the same as conversion, and the store is the only judge whose opinion is being counted.
Section 4 is right that localisation matters, and the failure mode is specific: teams translate the captions and ship the same images.
A translated caption on an unchanged screenshot is a listing built for a market you have not looked at. The claim that lands in one country is frequently not the claim that lands in another — a value proposition built on speed can matter far less than one built on price, or on privacy, or on working offline, depending on where somebody is standing when they read it. That is a market fact, not a language one, and no translator will surface it.
The screenshots themselves also carry assumptions that do not travel: currency symbols, date formats, name conventions, the sample data in the UI, and the app content itself. A demo screen full of English place names is quietly telling a German user that this app is not really for them.
So the honest version of localisation is: pick the two or three markets that actually matter to you, work out what people there care about, and build the assets for that — then use a straightforward translation everywhere else. Doing three markets properly beats doing twelve of them in a way that fools nobody.
The hybrid answer in the FAQ below is the right one, and it is worth spelling out where the boundary actually sits, because both extremes fail for different reasons.
Pure marketing graphics
- Apple requires screenshots to show the actual app
- A listing that shows an app you do not have is
a review problem, not an optimisation
- And it sets an expectation the product then breaks
Pure raw UI
- Asks a stranger to decode an interface in one second
- Nobody does this
- Converts worst of all, despite being the most honest
The workable middle is a real screen, cropped to the part that matters, with a caption above it doing the explaining. The screen is genuine — it is your app, it is what they will get — but it has been framed, in the way a photographer frames rather than the way a novelist invents.
The line worth holding: never show a capability the app does not have. Cropping, zooming, highlighting and using well-chosen sample data are all fair. Mocking up a feature that does not exist, or filling the screen with data no real user would have, buys an install from somebody who will uninstall in a day and leave a one-star review explaining why — and the rating costs more than the install was worth.
Minimum 3 (the visible ones). Optimal 6-8 (story plus depth). Don't skip the visible 1-3 — that's where conversion happens; later slots are for diligent browsers.
Portrait for most apps (matches phone use). Landscape only for genuinely landscape-first apps (racing games, video editors). Don't mix in same product page.
Hybrid is best. Caption + cropped-but-real UI screenshot. Pure marketing graphics violate App Store guidelines (must show actual app). Pure raw UI doesn't convert.
Every 8-12 weeks during optimisation. Test one variable at a time. Quarterly minimum once stable to refresh with new features.
By testing them, not by preferring them. Apple's Product Page Optimization runs a real split test on real store traffic and returns a number for your listing; Google Play offers store listing experiments. Both settle in a fortnight the arguments a team can otherwise have for a quarter. The specific conversion multiples quoted in ASO articles describe other people's apps in other categories, and the cases where a redesign made things worse do not get written up — so use them for direction and your own test for magnitude.
The first one, and often it is the only one anybody sees. In search results a listing frequently shows one or two screenshots at small size beside the icon and title, and the decision to tap is made there — before anyone reaches the product page and the story you carefully built across slots four to ten. So the first slot is not the opening of a sequence; it is a standalone advertisement that must work alone, at thumbnail size, in about a second. Design it that way and let the rest support it.
Yes, because nobody decodes an unfamiliar interface in one second. A raw screenshot asks a stranger to look at a grid of controls they have never seen and infer the value of your product from it — which is a task nobody performs while scrolling. A caption does that work for them: it states the outcome in four to six words, and the screenshot beneath becomes evidence for the claim rather than a puzzle. And make the text large: screenshots are viewed at a fraction of their exported size, and text sized for a desktop mockup is unreadable on the device that matters.
Test caption variants, slot order, social proof placement.
Run App Audit →aiwebpageseo.com is a data-driven SEO and AEO (Answer Engine Optimisation) platform providing a free suite of technical website tools. Rather than relying on AI-theorised assumptions, the platform analyses live URL performance, delivering objective diagnostics, page speed metrics, CLS debugging, and site crawl data alongside actionable technical tutorials.