Skip to content

MagicTask · Project management SaaS

From latency spikes at 5,000 users to 100,000 concurrent

A gamified project management platform re-architected from a single Node server into an event-driven system that holds 100,000+ concurrent connections without the reward engine touching the critical path.

100K+concurrent users

At a glance

Duration
78 weeks
Team size
8 people
Engagement
Replatform
Project type
Web application
Industry
B2B SaaS

The situation

The challenge

One Node server handled everything — WebSocket connections, task mutations, reward calculations and database writes. It was already spiking at 5,000 concurrent users, and the target was 100,000 within twelve months. The reward algorithm was the worst of it: every completion, comment, mention, reaction and login triggered a multi-step scoring pass across a user's entire history, synchronously, at peak.

What we did

We took the reward engine off the critical path first, because it was the thing making every other interaction slow, then scaled the connection layer behind it.

The calls that mattered

  • The reward engine became a queue, not a function call

    BullMQ over Redis, with a dedicated worker consuming events and writing XP to PostgreSQL in batches. The UI updates instantly and the arithmetic settles behind it within 200ms p99 — users were never waiting on the scoring, they were only ever waiting on the architecture.

  • WebSockets across instances instead of inside one

    In-memory Socket.IO rooms replaced with a Redis Pub/Sub adapter, sticky sessions at the ALB and heartbeat pruning. That is what turned one box into twelve EC2 instances sharing 100,000+ connections.

  • Leaderboards stopped being computed live

    Materialized views refreshed every five minutes, served from read replicas. A leaderboard that is five minutes stale is indistinguishable from a live one to a user, and enormously cheaper.

What changed

concurrent users
100K+concurrent users
API p95 latency
85msAPI p95 latency
user retention
+340%user retention
uptime
99.99%uptime
  • 100K+ — from 5,000 before the rebuild
  • 85ms — down from 1.2 seconds

The platform carries 100,000+ concurrent users and 2M+ events a day at 99.99% uptime, with p95 API latency down from 1.2 seconds to 85ms and retention up 340%.

Services used

  • Event-driven reward engine on BullMQ with batched writes
  • Horizontally scalable WebSocket layer across twelve instances
  • Read-replica and materialized-view strategy for leaderboards
  • Load-test harness and capacity plan to 100K concurrent

What we would do differently

Every project has one of these. Publishing it is the point — a case study with no regrets in it is marketing, not evidence.

Knowing when to evolve an architecture without stopping product velocity is the actual skill. The original monolith was not a mistake — it was correct for its stage. The judgement is in avoiding premature optimisation and premature architectural pessimism at the same time, while continuing to ship features throughout.