A casual games portal has one job that almost no one outside the industry thinks about: deciding what does and does not go on the catalog. At Suntongames, that decision is the single biggest determinant of whether the site gets better or worse over time. We currently have dozens of prototypes and submissions waiting in our review queue at any given moment. The fraction that actually ships is small — under one in five, in most months. This article is an honest account of how that selection works, what we look for, and why we say no far more often than we say yes.

The Queue: More Prototypes Than We Can Ship

The starting point is volume. Puzzle prototypes are cheap to produce compared to most genres, which means submissions arrive faster than any responsible review process can absorb them. Some come from internal prototyping; many come from external partners and independent developers; a few arrive through what we politely call serendipity. At any given time, the queue contains everything from a near-finished, well-tuned sort puzzle to a five-minute experiment with one good idea and no polish.

The temptation, with a queue this large, is to ship the safe middle — the competent, unobjectionable entries that will not embarrass anyone. We have made that mistake in the past, and the result is exactly what you would expect: a catalog that grows without getting better, and players who visit less. So we treat the queue the opposite way. The default answer to a submission is no. The work of the review process is to find the rare submission that earns a yes.

Shipping a mediocre game is not a neutral act. It displaces attention from a good one and teaches players that the catalog is uneven. The cost of a bad ship is paid by every good game on the site.
Only 5–10% of prototypes ship
Industry-wide, the fraction of game prototypes that make it from concept to a live release is estimated in the single digits. Our own ship rate sits in the same band — under one in five submissions in a typical month — which is a feature of the process, not a failure of it.
Source: Industry prototype-to-ship ratios discussed in GameDeveloper.com coverage and Game Developers Conference postmortems

The Playtesting Rubric

Every prototype that survives an initial screen enters a structured playtesting process. We do not rely on a single reviewer’s gut feel. Each candidate is played by multiple people across multiple sessions, and scored against a four-axis rubric. The rubric is deliberately short. A longer rubric produces noisier, not better, decisions, because reviewers start optimizing for the form rather than the game. This mirrors the consensus that has come out of years of Game Developers Conference (GDC) talks on prototyping: test early, test with strangers, and trust observed behavior over stated opinions.

1. Fun — the only axis that can veto alone

This is the only axis where a failing score rejects the game by itself, regardless of how well it scores elsewhere. The question is not “is this game fun for someone, somewhere” — almost any game is fun for someone. The question is whether the core loop produces the “one more attempt” pull during playtesting, in people who had no prior attachment to the prototype. We are looking for the moment where a tester, unprompted, restarts a level they just lost. If that moment does not occur across multiple testers, the game is not fun enough, and no amount of polish will fix it. Polish makes a fun game better; it does not make a boring game fun.

"The goal of a successful game designer is to create experiences that players find genuinely enjoyable... A game is an experience; the rules and the technology are merely the medium through which the experience is delivered." Fun is therefore not a feature you bolt on at the end — it is the experience the prototype either produces or does not, long before any art or polish is applied.— Jesse Schell, The Art of Game Design: A Book of Lenses (a standard reference text in game-design programs and GameDeveloper.com pedagogy)

2. Session length — short, but not too short

A casual puzzle is not supposed to be a 40-minute commitment, but it is also not supposed to be over in 20 seconds. We measure natural session length during playtesting — how long testers choose to play before voluntarily stopping — and we look for a sweet spot of roughly three to eight minutes per session, with a clear sense that more is available. A game that ends too fast reads as shallow. A game that demands too long a session reads as a commitment, which is not what casual players are looking for. The exact target varies by mechanic, but the curve matters more than the absolute number: sessions should be long enough to feel like something happened, and short enough that finishing one feels like a clean stopping point.

3. Learning curve — flat entry, real gradient

This is the axis we care most about after fun, and it is the one most submissions fail. The ideal casual puzzle teaches itself in under fifteen seconds and then keeps revealing depth for the next hundred levels. Most prototypes do one or the other: they are easy to learn but have no depth, or they have depth but require a tutorial nobody will sit through. We specifically test whether a tester can play the first level with no instructions at all. If they cannot, the onboarding is broken and the game goes back for rework before it gets a fair score on anything else.

4. Accessibility — the non-negotiable floor

Accessibility is treated as a floor, not a bonus. A game can be the most fun prototype in the queue and still be rejected if it fails basic accessibility checks. The floor includes: targets large enough for a finger on a mid-tier phone, color palettes that do not rely on hue alone (we test with simulated color vision deficiencies), no mechanics that depend on audio alone, no time pressure without a relaxed mode, and keyboard-equivalent input for any mouse-only interaction. A game that ships on Suntongames is playable by as wide an audience as we can reasonably engineer for. We are not perfect on this axis, but we refuse to ship games we know are excluding people unnecessarily.

The Rubric in Practice

The four axes are scored, but the scoring is less important than the conversation around it. A typical review produces a table like the one below, and the table is the start of the decision, not the end.

AxisWhat a “pass” looks likeCommon failure
FunTesters restart unprompted“It is fine, I guess”
Session length3–8 min natural sessionsEnds in 20s, or demands 20min
Learning curveFirst level no-instructionsNeeds a tutorial to start
AccessibilityMeets floor checksHue-only colors, tiny targets

The reason the conversation matters more than the score is that fun, in particular, is not a number. It is a pattern observed across testers, and the role of the reviewer is to describe that pattern accurately, not to average it into a misleading digit.

Why We Reject Most Submissions

It is worth being specific about why the rejection rate is as high as it is. Most submissions do not fail because they are bad games. They fail because they are not meaningfully better than games already on the catalog. The bar is not “is this fun in isolation” — it is “does this earn a place on a shelf that already has good games on it.” A perfectly competent color-sort clone, for example, will be rejected not because it is broken but because it does not add anything the catalog does not already have. We say no to competence all the time, and that is the point.

Other common rejection reasons, in roughly descending frequency:

  • Cloned without improvement. A copy of an existing game that does not tune, refine, or extend the mechanic in any meaningful way. The market does not need another identical version.
  • No depth past level ten. The first few levels are charming and then the game runs out of ideas. The difficulty engine is missing or shallow.
  • Unsolvable or unfair states. Boards generated by random shuffle that produce configurations the player cannot actually solve. This is a design bug, not a difficulty feature.
  • Onboarding failure. Requires reading instructions to play. Casual players will not read instructions, and the game should not ask them to.
  • Accessibility regressions. Hue-only color coding, tiny touch targets, audio-dependent cues. Fixable, but the submission is sent back until fixed.
  • Friction-by-design. Energy systems, forced interstitials between every level, artificial wait timers. These are not game mechanics; they are retention theatre, and they contradict the design philosophy described in our piece on why casual puzzles still win in 2026.

The Quality Control Process, End to End

For the small fraction of submissions that pass the rubric, the work is not over. Shipping is a separate process from selection, and it has its own gates.

  1. Initial screen. A single reviewer spends five minutes with the prototype. Most are rejected here for obvious reasons — clone, broken, or off-category.
  2. Rubric playtesting. Multiple testers play across multiple sessions and score against the four axes. The conversation around the scores decides whether the prototype advances.
  3. Tuning pass. The prototype is tuned — difficulty curve, board generation, accessibility fixes. This is where a “yes, but” submission becomes a “yes.”
  4. Engineering integration. Code-split chunk, performance budget check (see how we hit sub-3-second loads), asset pipeline, lazy-loaded banner. The game has to meet the same performance bar as the rest of the site.
  5. Final cross-device check. iPhone, mid-tier Android, tablet, laptop with mouse and trackpad. The game must feel right on all of them, per our mobile vs desktop guidance.
  6. Ship. The game goes live. Then the real test begins: do players, given the choice, actually play it? A game that earns its place in the queue but not in player sessions is retired, not defended.

That last step is the one that keeps the catalog honest. Selection does not end at ship. Every game is, in effect, re-reviewed every month by the players who do or do not open it. A game we were sure would succeed that nobody plays is removed to make room. A game we were unsure about that quietly becomes a favorite stays. The rubric picks the candidates. The players pick the winners.

What This Means for Players and Submissions

For players, the practical implication is that the Suntongames catalog is intentionally small relative to the size of the queue behind it. We would rather ship twenty games we believe in than two hundred we are unsure about. For developers submitting prototypes, the implication is that rejection is the default outcome and not a judgment on the work — most submissions are competent, and competence is not enough to earn a slot. We try to be specific in rejection feedback, because the same prototype, retuned, may well pass on a second submission.

The bar is high on purpose. The only way a casual puzzle catalog stays worth visiting is if the next game added is at least as good as the one above it. That is the standard we hold ourselves to, and the one we will keep holding the queue to.