Starlit Reviews
A Shopify reviews app that collects photo and video reviews and proves which ones came from real buyers.
00
problem
Shoppers have learned to distrust star ratings, because most review apps let a store quietly bury the reviews that hurt and reward the ones that help. The thing that is supposed to make a product safe to buy has become the least reliable element on the page.
solution
Starlit treats authenticity as a feature rather than a disclaimer. Reviews tied to a confirmed order are labelled, moderation is neutral by construction so a merchant cannot filter display by rating, and incentives are conditioned on adding a photo or video and never on being positive, so the reward path is identical for a one star and a five star review. Around that sits what merchants actually shop for: photo first widgets that inherit the store's own type, colour and corner radius from a single token layer, plus home sections and a shoppable social reel. Published reviews go out as structured data for Google and as feeds that AI shopping assistants can read and cite.
Context
Product reviews are the last thing a shopper reads before deciding, and the part of a storefront they trust least. The apps that supply them have quietly taught people why: the category norm is to let a merchant hold reviews for approval, filter what appears by rating, and offer a discount in exchange for one. Every individual setting is defensible. Together they produce a five star wall that means nothing.
Starlit Reviews is a Shopify app built on the opposite bet: that a review system which cannot be gamed by the store that owns it is worth more, to the shopper and eventually to the merchant, than one that can. It collects photo and video reviews after purchase, labels the ones tied to a confirmed order, refuses to gate display on rating, and publishes the result as structured data that both search engines and AI shopping assistants can read.
It is a personal product, designed and built between June and September 2026, and it is in development and being prepared for Shopify App Store submission.
My Role
I did the product definition and competitive research, the visual design and the storefront design system, the merchant facing admin, and the entire build: a Python backend, an embedded React admin using Shopify Polaris, and a Shopify theme app extension written in vanilla JavaScript and Liquid.
That range is the reason the design decisions below survived. The constraint that shaped the widget most, a hard weight budget, is one a designer usually hears about after the fact.
The incentive that had to be redesigned
Most review apps offer the shopper a discount for leaving a review. The moment that reward is even loosely tied to sentiment, the reviews stop being evidence. United States FTC rules published as 16 CFR Part 465 (2024) prohibit incentives conditioned on the sentiment of a review, which turned a soft design opinion into a hard boundary.
So the reward conditions on effort, not on opinion. A shopper earns the incentive by adding a photo or video. The path is byte for byte identical for a one star and a five star review, which means the merchant gets the thing they actually wanted, more visual reviews, without buying the thing that destroys the asset, agreeable ones.
The same logic runs through moderation. Reviews auto publish after spam screening. A merchant can reply, and can hide a review with the action logged, but there is no filter that removes low ratings from display, because a control like that gets used. On top of that sits an earned transparency badge in four tiers, computed from a shop's real behaviour: verified buyer share, volume, and how the merchant responds. A store that suppresses by rating earns no badge at all, which is the entire point of having one.
One token layer, every store
A review widget lives inside somebody else's design. Most competitors solve this with a settings panel of colour pickers, and the widget still reads as a bolt on because its type, spacing and corner language never match the theme around it.
Starlit ships one token layer instead. Colour, spacing, type scale and radius are CSS custom properties. Every derived surface, hairlines, soft fills, shadows, is mixed from the merchant's own background and text colours rather than hardcoded grey, so the widget stays correct in any palette. The whole radius language, from small chips to large cards to pills, derives from one merchant set number. Type is set with clamp so headings scale down on phones without a breakpoint. Font family inherits, always, so the widget takes the store's typography rather than imposing its own.
The proof is in the gallery: three of those images are the same component code with different token values, and they read as three different products.
The second constraint was weight. A widget is third party code on someone else's storefront, so it is an ethical load, not just a technical one. The build fails if any asset exceeds 10 KB gzipped, which is Shopify's Built for Shopify budget. Staying under it forced the write a review wizard and the media lightbox to load only when a shopper opens them, with dependencies passed in explicitly rather than shared through a global.
Users and Constraints
Two users with opposed instincts. The merchant wants the wall to look good and is one checkbox away from curating it. The shopper wants to know whether the thing is worth buying and can smell curation instantly. Designing for both means occasionally refusing the merchant a control they expect, and being clear about why.
Constraints that shaped it: every asset under 10 KB gzipped; no front end dependencies at all on the storefront, so no framework and no runtime library; the admin had to be Shopify Polaris, so the design work there was restraint and information hierarchy rather than surface; touch targets at least 44 by 44; reduced motion respected; and six languages in the widget from the start.
Key Insight
The category sells social proof, but its pricing and its controls both push merchants toward reviews that are prettier than they are true. Pricing that meters by order volume punishes the stores that grow, and moderation that filters by rating punishes the shopper. Those are the same mistake in two places: optimising the number rather than the signal behind it.
Once I stopped treating authenticity as trust-badge decoration and started treating it as the product constraint, most of the hard calls answered themselves. It told me what the incentive could condition on, what moderation was allowed to do, what the dashboard should lead with, and even what belongs on the free tier: a reviews app whose free plan cannot collect reviews fails at the job it exists for.
Design Principles
Authenticity is structural, not a badge. If a control can be used to fake the signal, it does not ship.
Inherit, never impose. The widget takes the store's type, colour and radius from one token layer.
Weight is a design decision. Every asset stays under 10 KB gzipped, enforced by the build.
Reward the effort, never the opinion. One star and five star follow identical paths.
Show the merchant the truth first. The dashboard leads with verified share, not review count.
Never sell what is not built. A paid feature flag that nothing enforces fails the test suite.
Impact
This is a personal product in development, so there is no traction data and I am not going to invent any. What I can report is measured, and measured on 2026-09-24 in the codebase itself.
Measured in the repository: 345 automated backend tests, 18 storefront blocks, 16 database tables, 6 widget languages, built solo across 192 commits over roughly three and a half months.
Measured at build time: every storefront asset ships under the 10 KB gzipped Built for Shopify budget, the largest sitting at 9.6 KB.
Internal analysis during design suggested the flat pricing model is defensible without order caps, which is what made an uncapped free tier possible rather than a marketing line.
From my own competitive research during design in 2026: the two category leaders differentiate on beauty and on price, and neither publishes reviews in a form aimed at AI shopping assistants. That gap became the forward looking part of the positioning rather than another price argument.
The honest outcome is a craft one. The product stayed coherent because one constraint, provable authenticity, was allowed to overrule feature requests, including some I wanted.
Reflection / Takeaway
The most useful thing I did was let a legal constraint act as a design brief instead of a compliance checkbox. Reading the FTC rule closely did not restrict the product, it resolved it: once an incentive cannot condition on sentiment, the reward has to condition on something else, and effort turned out to be the better thing to ask for anyway.
The second was building it end to end. Several of the decisions I am proudest of, the lazy loaded wizard, the derived colour surfaces, the plan flag that fails the build if nothing enforces it, only exist because the same person hit the constraint and owned the design response to it. Designing at that distance from the material is a habit worth keeping.
The storefront sections are what a merchant actually picks the app for. A quote, its rating, its author and the photo rail beside it advance as one unit, so the wall reads as a run of real people rather than a grid of text. Every surface in it, the hairlines, the soft fills, the card backgrounds, is mixed from the store's own colours rather than hardcoded grey, which is why it sits inside this theme instead of on top of it.
On a product page the question narrows to whether this specific thing is worth buying and what goes wrong with it. The summary answers that in a sentence and then in chips, and the critical one stays in view, because a block that only lists what people liked is exactly what shoppers learned to discount. Underneath, the list can be filtered to reviews with photos and reordered by recency, but not by rating, which is the one control the category offers and this product does not have. What the shopper sees is what the merchant sees.
The badge is the claim and opening it is the proof. It states in three lines what the tier actually guarantees, reviews from confirmed buyers, no hiding or sorting by rating, spam removal only, then shows the verified share and the published count it was computed from, so the claim can be checked rather than taken. The tier is earned from the shop's own behaviour and cannot be bought, which is why a store that suppresses low ratings carries no badge at all. A trust mark every merchant can buy is decoration, and this is the one element of the product worth faking, so it is the one that had to be computed.
Around those two decisions sits what merchants actually shop for: a rating summary that answers the buying question, the widget family that covers every placement, and the moderation queue they will spend real time in. The last two are drawings rather than screens, because the two claims this product rests on are not things a screenshot can prove.
The first sets what a merchant can do, reply, hide and delete, against what the product refuses to offer: hiding the one and two star reviews, ranking the happy ones to the top, approving each review before it shows. Those are not settings someone turned off. They were never built, so there is nothing to turn on. The second draws the incentive as a grid. Read it down the columns, from no media to a photo, and the reward changes. Read it across the rows, from one star to five, and nothing changes at all. That is the whole compliance posture, and it is a shape in the product rather than a sentence in a help page.









