Skip to content
All projects

Case study17 min read

Romish

A CS2 10-player matchmaking platform where captains draft teams and ban maps, then play on a server the app sets up itself.

Role
Solo, design to deployment
Timeline
2024 to present
Status
Pre-launch
Stack
Next.js 16, TypeScript, MongoDB, Redis, Pusher
  • Next.js 16 (App Router)
  • React 19
  • TypeScript
  • Tailwind CSS v4
  • MongoDB Atlas + Mongoose
  • Upstash Redis
  • Upstash QStash
  • Pusher
  • NextAuth v5
  • Steam OpenID 2.0
  • DatHost API, FTP, RCON
  • MatchZy (CS2 plugin)
  • Stripe
  • Zod
  • Pino
  • Sentry
  • Jest
  • Playwright
  • GitHub Actions
  • Vercel
The landing page. Sign-in is Steam only.

At a glance

  • 10

    players synced in real time per match

  • 7

    match phases, from Find Match to Results

  • 8

    server-side deadline types, none depending on an open browser tab

  • 40

    automated failure-path checks run by bot scripts

01 / TL;DR

What I built, and what it shows

Five lines for a quick read. The rest of the page is the detail behind each one.

  • A full product, end to end.

    I designed and built the platform that takes ten players from Steam sign-in to updated Elo, on Next.js 16, React 19 and TypeScript.

  • Real-time systems.

    Ten people act in sync over Pusher channels, and every deadline runs on the server, so no outcome depends on a browser being open.

  • Distributed state.

    Redis locks and Lua scripts keep the queue, sessions and timers correct across serverless instances. MongoDB holds the durable record.

  • Third-party integrations.

    Steam OpenID, DatHost, FTP, RCON, the MatchZy game plugin, Stripe and QStash, with a signature check on every one that calls back in.

  • Security and testing.

    CSRF, rate limits, CSP and a role-gated admin panel with an audit log. Jest and Playwright in CI, plus bot scripts that play whole matches.

02 / The problem

Ten people, one match, no admin

Ten people have to act in sync, and any one of them can stall the match.

Community CS2 10-mans are usually run by hand. Someone collects ten players, checks who is actually there, balances two teams, runs a map veto in chat, sets up a server, and then updates everyone's rating in a spreadsheet. It's harder to automate than it looks.

  • 01

    Everyone waits on someone

    Ten accepts, eight picks, six bans. Each step blocks on one person, so the flow has to move on by itself when they don't act.

  • 02

    Anyone can vanish

    Players go AFK, close the tab or lose their connection at any point, and the match still needs a clean outcome.

  • 03

    The server has to match the lobby

    A real CS2 server has to be started and configured for exactly those ten players on the chosen map, and then report the result back.

03 / The player journey

From Find Match to results

Seven phases on one continuous screen, and the server owns every clock.

For the player it is one flow. The same map backdrop and frame carry through from the ready check to the results, and every screen is built phone-first.

  1. 01Solo or party

    Find Match

    • Queue solo or as a party of up to five. The party leader queues everyone.
    • Every member is checked before joining: bans, queue access, cooldowns and any other active session.
    • The Play button follows the player's state: Find Match, Ready Up, a cooldown countdown, or back to their match.
    The Play page: party slots, the Find Match button, rank and Elo history, and live matches.
  2. 0225s

    Ready check

    • When ten players are found, everyone gets Accept or Decline and a 25-second countdown (admins can change it).
    • A decline or a no-show sends everyone who accepted back to the front of the queue, keeping their original wait time.
    • The timeout is resolved on the server, even if nobody has the page open.
    Ten players found: everyone has 25 seconds to accept.
  3. 0330s per pick

    Draft

    • Each party captains its own team through its highest-Elo member. With no parties, the two highest-Elo players captain.
    • Captains pick in the order A A B B A B A B. Party members are already placed on their team.
    • If a captain runs out of time, the server picks a random available player.
    The captain draft. The server picks at random if the clock runs out.
  4. 0420s per ban

    Veto

    • Captains take turns banning maps until one is left, and that map is played.
    • A missed turn bans a random map, so a captain who leaves can't stall the match.
    • The game server has been booting in the background since the match was created.
    The map veto: alternating bans until one map is left.
  5. 053 min cap

    Server setup

    • The platform starts the server, uploads a match config for these ten players and loads it, with no human involved.
    • Only the ten players see the connect details: the address, copy buttons and a steam:// link.
    • If setup fails twice or takes over 3 minutes, the match is cancelled with no penalties and all ten go back to the front of the queue.
    Server setup, shown while the platform starts and configures the game server.
  6. 06MatchZy

    Live

    • The MatchZy plugin runs the knife round, warmup, match and overtime on the game server.
    • The score updates round by round from a signed webhook, and anyone signed in can spectate.
    • A player who disconnects has 5 minutes to reconnect before it counts as an abandon.
    A live match. The score updates round by round from the game server.
  7. 07Elo applied

    Results

    • The results screen shows the winner, score, MVP, K/D/A and every player's Elo change.
    • Every player is released from the match at once, so Queue again works straight away.
    • The game server is stopped a minute later if no other match needs it.
    Results: every player's K/D/A and Elo change, and the match MVP.
The same match flow on a phone: draft, veto, server setup and results.

Back to top

04 / Engineering deep dives

What made it harder than a CRUD app

Six problems, each with the problem, what I built, and why it holds up.

a.

Server-side timers

Timeouts are decided on the server, exactly once, whether or not anyone has the page open.

The problem

A match has eight kinds of deadline: ready checks, turns, setup and result timeouts, abandons, a live score poller and the server release. If browsers decided them, one closed tab could freeze a match and two open tabs could fire the same timeout twice.

Show what I built and why it holds up

What I built

  • Every deadline lives in one Redis sorted set: the member is the task, the score is when it is due. Scheduling the same task again just moves its deadline.
  • A Lua script reads the due tasks and removes them in one step, so each task goes to exactly one caller, however many instances are sweeping.
  • On a long-running server, a sweeper runs every 2 seconds, holding a short Redis lease so only one instance sweeps at a time.
  • On serverless hosting, scheduling a deadline also publishes a delayed QStash message that wakes a signed endpoint one second after it is due.

Why it holds up

  • An early, duplicate or stale wake-up finds nothing to do, because the Redis set is the source of truth.
  • A task that throws is rescheduled 10 seconds later instead of being lost.
  • Handlers re-check state before acting. An auto-pick is a conditional write on the pick index, so it lands once even if the sweeper and several open pages report the same timeout.
Show the scheduler pseudocode
Pseudocode
schedule(task, dueAt):
  ZADD deadlines dueAt task        # same task again = new deadline
  if serverless:
    qstash.publish("/api/cron/match-tick", delay = dueAt - now + 1s)

sweep():                           # every 2s, or on a QStash wake-up
  tasks = EVAL takeDue(now)        # Lua: read due members + remove them, atomically
  for task in tasks:
    try:   run(task)               # re-checks the match before acting
    catch: schedule(task, now + 10s)

b.

One active session per player

A player can only be in one flow at a time, and a stale lock can never trap them.

The problem

A double click, two tabs, or a party leader queueing while a member joins somewhere else could put one player in two matches. Nine people waiting on someone who is busy elsewhere is a dead match.

Show what I built and why it holds up

What I built

  • Each player has a session record in Redis: what they are in (queue, ready check or match), the phase, and its deadline.
  • Joining the queue claims the lock for the whole unit, solo or party, in one Lua call. Either every member gets it or nobody does.
  • Every way into a flow refuses a player who is already in one, and names who is blocking ("Viper is currently in a match").
  • One browser tab drives the session. Other tabs can watch, and show "Session active in another window" with a Use this tab button.

Why it holds up

  • Reads heal the lock. Every read checks it against the queue, the ready check and the match: a lock whose match is over is released on the spot, and a player who is really in a match gets their lock back.
  • Every way out (result, forfeit, any cancel) goes through one release function, so no path forgets a player.
  • Refreshing, bookmarking or typing a URL lands on the real phase page through server-side redirects.
Show the session lock pseudocode
Pseudocode
claimSessions(members, session):   # one Lua script: all or nothing
  for m in members:
    if EXISTS session:{m}: return CONFLICT(m)
  for m in members:
    SET session:{m} session EX safetyTtl
  return OK

readSession(user):                  # every read reconciles
  lock  = GET session:{user}
  truth = queued(user) or readyCheck(user) or liveMatch(user)
  if lock and not truth: DEL session:{user}      # match over: free them
  if truth and not lock: SET session:{user} ...   # really playing: restore

c.

Automating the game server

From the last ban to a configured CS2 server with no human involved, and a clean exit if it fails.

The problem

Before anyone can play, a real CS2 server has to be running, locked to exactly these ten players, on the map the veto chose, and able to report the result back. Every one of those steps can fail.

Show what I built and why it holds up

What I built

  • Start early. The DatHost server is started when the match is created, so it has usually booted by the last ban. Setup checks again and waits for the boot if needed.
  • Configure. The app builds a MatchZy match config (teams, Steam IDs, map, a knife round for sides) and a whitelist of the ten Steam IDs, and uploads them over FTP.
  • Load. It runs the config over RCON, reads the server address, and flips the match to live only if it is still in setup, with a conditional update.
  • Report back. MatchZy posts going live, every round, disconnects and the final result to a webhook that checks a shared secret in constant time.

Why it holds up

  • One automatic retry after 5 seconds, and a 3-minute cap on the whole setup.
  • If it still fails, the match is cancelled with no penalty for anyone, and all ten players go back to the front of the queue.
  • A setup that finishes after its match was cancelled never goes live. The server is released instead.
  • Finishing a match is idempotent. MatchZy can send both of its end-of-match events and Elo is still applied once.
Show the MatchZy match config
The match config MatchZy loads (trimmed)
{
  "matchid": 1042,
  "team1": { "name": "<alpha captain>", "players": { "<steam64>": "<name>" } },
  "team2": { "name": "<beta captain>",  "players": { "<steam64>": "<name>" } },
  "num_maps": 1,
  "maplist": ["de_mirage"],
  "map_sides": ["knife"],
  "players_per_team": 5,
  "cvars": {
    "matchzy_remote_log_url": "<app>/api/webhooks/matchzy",
    "matchzy_remote_log_header_key": "X-MatchZy-Secret"
  }
}
Server health in the admin panel. Only the owner can power the real server.

d.

Fair play

Rules that punish bad behaviour, not bad luck.

The problem

Teams have to be balanced and captains chosen fairly, and players who dodge or abandon need a consequence. None of that should punish anyone for the platform's own failures.

Show what I built and why it holds up

What I built

  • Matchmaking takes the oldest waiting group first. A party is taken whole or not at all and never larger than one team, with an optional cap on the Elo spread. It is a pure, unit-tested function shared with the admin Match Lab.
  • Captains: each party goes on one team and its highest-Elo member captains it. Otherwise the two highest-Elo solo players captain.
  • Elo: each side's rating is its team average, K = 32, and every player on a side gets that side's change, so the result is zero-sum.
  • Cooldowns and Trust Score escalate within a rolling 24 hours. Admins can lift a cooldown, and the offense still counts toward the next one.

Why it holds up

  • Server setup failures, admin cancels and missing results never penalise anyone.
  • Each player is penalised at most once per match.
  • Every timer, cooldown and penalty number lives in one rules file.

Penalties within a rolling 24 hours. Trust Score starts at 100 and never drops below 0.

  • Decline a ready check

    Queue cooldown (1st, 2nd, 3rd and later)
    5 min, 15 min, 1 h
    Trust Score
    No change
  • Miss a ready check

    Queue cooldown (1st, 2nd, 3rd and later)
    5 min, 15 min, 1 h
    Trust Score
    No change
  • Abandon after accepting

    Queue cooldown (1st, 2nd, 3rd and later)
    5 min, 15 min, 1 h
    Trust Score
    -5, -10, -20
  • Captain turn timeout

    Queue cooldown (1st, 2nd, 3rd and later)
    None
    Trust Score
    -5 each, from the 3rd
Show the Elo calculation
Elo after a match
ratingA   = average(team A Elo)
ratingB   = average(team B Elo)
expectedA = 1 / (1 + 10 ^ ((ratingB - ratingA) / 400))
changeA   = round(32 * (resultA - expectedA))   # resultA: 1 win, 0 loss
changeB   = -changeA                           # zero-sum: team B mirrors it
Every player starts with a Trust Score of 100.

e.

Realtime and privacy

Everyone sees the same state instantly, and nobody sees what they shouldn't.

The problem

Ten players, spectators and admins all watch the same match. Picks and bans have to appear instantly, but connect details, party chat and direct messages must only reach the right people.

Show what I built and why it holds up

What I built

  • Pusher channels are split by audience: public channels carry only public state, private channels carry per-user, party and team data, and a presence channel runs global chat.
  • The server authorises private channels by checking membership (your own user channel, your party, your team in this match) and the CSRF token.
  • The browser keeps one subscription per channel, however many components listen to it.
  • Spectators get an allow-list of match fields. Any field added to the model later stays hidden until it is listed.

Why it holds up

  • No public event carries connect details, IPs or passwords. When the server is ready, players refetch the match, and only the ten players get the address back.
  • Draft, veto and server setup can't be spectated at all.
  • Pages resync on focus, on reconnect and on a slow poll, so a dropped connection never leaves a stale screen. Clocks use the server's time, so a refresh shows the same number.

Pusher channels by audience.

  • queue, match-{id}, live-matches

    Type
    Public
    Carries
    Queue changes, picks, bans and scores. Never private data.
  • private-user-{id}

    Type
    Private
    Carries
    Session updates, party, friends, notifications and DMs
  • private-party-{id}, private-team-{match}-{side}

    Type
    Private
    Carries
    Party chat and team chat
  • presence-global-chat

    Type
    Presence
    Carries
    Global chat, with who's online

f.

Security

Every way into the app is checked: browsers, game servers, payments and the scheduler.

The problem

The app takes requests from browsers and from four machine callers (the game server, DatHost, Stripe and QStash), and it has an admin panel that can ban players and power a real game server.

Show what I built and why it holds up

What I built

  • A proxy runs before every request. It applies rate limits (stricter on login and queue join, plus per-action limits on invites, friend requests and chat) and checks a CSRF double-submit token on every mutating API call.
  • Webhooks skip CSRF and prove who they are instead: MatchZy with a shared secret compared in constant time, DatHost with an HMAC, Stripe with its signature, and QStash with a signed JWT that supports key rotation.
  • Steam sign-in is verified server-side with Steam. Sessions are stored in the database behind a host-only, HttpOnly cookie.
  • Security headers: a Content Security Policy, X-Frame-Options DENY, nosniff, and strict referrer and permissions policies.
  • Admin routes are gated by role (moderator, admin, owner), and every admin change is written to an audit log. Config changes are stored as a before and after diff.

Why it holds up

  • Request bodies are validated with Zod.
  • User input in database search queries is escaped.
  • Blocks are never revealed: a blocked player's card just returns "not found".

Back to top

05 / Architecture

How it fits together

One Next.js app, two stores, and managed services for realtime, timers, game servers and payments.

Romish is a single Next.js app. Route handlers call service modules that own the rules, and the rule modules themselves (matchmaking, draft, veto, Elo) never touch a database, so the API, the admin Match Lab and the tests all run the same code.

Romish system overviewThe browser talks to the Next.js app over HTTPS and receives live updates from Pusher. The app stores durable data in MongoDB and live state in Redis, schedules wake-ups through QStash, pushes events through Pusher, drives the CS2 server through DatHost, and uses Steam for sign-in and Stripe for billing. The CS2 server reports results back to the app through the MatchZy webhook, and QStash wakes the app when a deadline is due.HTTPSeventspushstart, configresults webhookNext.js appApp Router, TypeScriptproxy: rate limits, CSRF, page guardroute handlers → lib/ servicestimer sweeper (long-running hosts)BrowserReact pagesPusherrealtime channelsMongoDB Atlasdurable recordsUpstash Redisqueue, locks, timersQStashdeadline wake-upsSteamOpenID, Web APIStripecheckout, webhooksDatHostAPI, FTP, RCONCS2 serverMatchZy plugin

Where state lives

If it must survive a restart or feed history, it goes in MongoDB. If it changes every few seconds or expires on its own, it goes in Redis.

  • MongoDB

    Holds
    Users, matches, friendships, messages, notifications, reports, roles, audit and system logs, runtime config
    Why
    Durable and queried by history
  • Redis

    Holds
    Queue, ready checks, live parties, presence, session locks, cooldowns, timers, live scores
    Why
    Changes every few seconds or expires on its own

Back to top

06 / Handling failure

What happens when things go wrong

Every failure I could find has a defined outcome, and the platform's own faults never cost a player anything.

  • Nobody answers the ready check

    What happens
    Resolved on the server at the deadline. No-shows get a cooldown, and everyone who accepted goes back to the front of the queue.
  • A captain goes AFK or disconnects

    What happens
    The server picks or bans at random every 30 or 20 seconds, so the match never deadlocks. Trust Score drops from the 3rd timeout in 24 hours.
  • A player leaves the site during draft or veto

    What happens
    Presence is checked every 15 seconds. Five minutes away counts as an abandon: a cooldown and a Trust Score penalty.
  • The game server won't start, or FTP or RCON fails

    What happens
    One retry, then the match is cancelled with no penalty and all ten players are requeued at the front.
  • Setup finishes after the match was cancelled

    What happens
    The conditional update refuses to go live. The match stays cancelled and the server is released.
  • MatchZy never reports a result

    What happens
    After 3 hours the match is cancelled and flagged for an admin to review. Nobody is penalised.
  • MatchZy sends the final result twice

    What happens
    Finishing a match is idempotent: the second call returns early, so Elo is applied once.
  • Two admins cancel the same match at once

    What happens
    A conditional status update lets one of them win. The other is told the match has already ended.

Back to top

07 / Beyond the match

The platform around it

Social features, a full admin panel and billing sit around the match flow.

  • Social

    • Friends with live presence (online, in queue, in party, in match). One record per pair, so one-sided states are impossible.
    • Parties of up to five, with invites by name search, from friends, or through a 15-minute invite link.
    • Chat in four channel types (global, party, team and DMs) with unread badges, edits, reports, slow mode and a profanity filter.
    • Notifications with toasts and inline actions that clear in every tab once acted on.
  • Admin panel

    • Moderation: reports with evidence, bans, chat mutes and player notes.
    • Live config without a deploy: queue rules, party size, ready-check time, maintenance mode and the site banner, pushed to open tabs.
    • Analytics: player, matchmaking and server charts, retention cohorts and the Elo distribution.
    • Every admin action recorded in an audit log.
  • Billing

    • Stripe Checkout for three monthly tiers.
    • A signed Stripe webhook activates, renews and cancels subscriptions, and players can cancel themselves.
    • Admins can grant access by hand, and an open beta switch lets everyone queue.
The friends panel slides out from any page.
The admin overview: live counts and the health of every dependency.

Back to top

08 / Brand and design

From placeholder to the Ready ring

A mark that still reads at 16px, and a quiet system with one warm accent.

Romish ran for most of its life on a working name and a Spartan-helmet placeholder. Before launch I rebuilt the brand around the product itself.

The name is Rom-ish: Romanian plus Irish. I explored eight logo concepts over three rounds, from geometric monograms to heritage ideas like the Dacian draco and a wolf-teeth shield, and tested each one as an app icon, a 16px favicon and inside the navbar.

The winner is the Ready ring: a heavy ring with a notch and an amber dot. It reads as a crosshair and as the ready state, the moment a match pops, and it still works at 16px where the detailed concepts fell apart.

The system is deliberately quiet: one warm accent against a dark UI, Unbounded for the display voice, Geist for the interface, and Geist Mono for anything that changes or lines up, like scores, Elo and timers. The colours live as tokens in Tailwind v4 and are being mapped onto shadcn's variables, so existing components pick up the brand without rewrites.

Screenshot coming soon

Logo concept board

Eight concepts, each tested as an icon, a favicon and in the navbar.

The mark

32px
16px

Palette

  • Background#0B0C0E
  • Surface#121316
  • Ink#EDEBE6
  • Muted#9398A1
  • Amber#FFAA1F
  • Team beta#4C7DFF

Type

  • ROMISH

    Unbounded 600, display

  • Map veto · 20s per ban

    Geist, interface

  • 13 : 11 · +18 Elo

    Geist Mono, scores, Elo and timers

Before and after

Before

After

Screenshot coming soon

Play page, after

Before and after the rebrand.

Back to top

09 / Testing and tooling

How I know it works

Bots play whole matches, so the failure paths get tested without ten real people.

Jest unit and API tests
The pure rules (matchmaker, Elo, draft and veto), the session lock, scheduler, penalties, server setup and spectator privacy, plus route handlers for the queue, matches, sessions and admin role gates.
A fake Redis that runs the real Lua
An in-memory stand-in for Upstash with strings, sets, sorted sets, hashes, pipelines and the app's own Lua scripts, so the atomic claims are tested as written.
Bot scripts
Scripts that act as seeded test users and play through the flow. The failure-path script runs 40 checks: ready-check timeouts, declines, setup failure and requeue, captain timeouts, two admins cancelling at once, and spectator privacy. It fast-forwards deadlines through the production code, so it never waits on real clocks.
CI on every push
GitHub Actions runs lint, the type-check, npm audit, Jest with coverage, a Playwright smoke set and a production build.
Match Lab and UI Studio
Admin tools that run each match stage in a sandbox against fake players using the production rule functions, and render every real screen with simulated data, each state reachable by URL.
UI Studio opens every real screen with simulated data. Match Lab, in the next tab, runs the match rules.

Back to top

10 / Lessons learned

What broke, and what I changed

The bugs that taught me the most, and the rule each one left behind.

  1. 01

    One definition for the session cookie

    Cookie settings drifted apart between NextAuth, the Steam login route and the user endpoint, and logins failed silently: players signed in, then looked signed out on the next page. Now all three read one shared definition.

  2. 02

    Don't log players out everywhere

    Deleting other sessions on login looked like good hygiene, but it signed players out on every other device. I removed it and left a comment in the code so it does not come back.

  3. 03

    Load images in batches

    Loading a dozen large map images at once caused connection resets in development. Map grids now load in small, staggered batches.

  4. 04

    Let the server own the clock

    When timers ran in the browser, what happened at a deadline depended on who had the page open. Moving every deadline to the server made outcomes the same whether or not anyone is watching.

11 / Status and what's next

Where it is now

Pre-launch. The match flow works end to end, from queue to results.

The full match flow is built and covered by tests and bot scripts. Before opening it up, I am working on:

  • Map pools set from the admin panel and used in every veto
  • A priority queue for paid tiers
  • Complete GDPR data export and account deletion
  • Tournaments

Want a walkthrough of the code?

The repo is private, but I'm happy to walk through it on a call: the timers, the session lock, or anything else on this page.