Content par jayen

MagicTask · Project management SaaS

From latency spikes at 5,000 users to 100,000 concurrent

A gamified project management platform re-architected from a single Node server into an event-driven system that holds 100,000+ concurrent connections without the reward engine touching the critical path.

100K+concurrent users

Aik nazar mein

Muddat
78 hafte
Team ka hajm
8 afraad
Muahide ki naueeyat
Platform tabdeeli
Project ki qism
Web application
Shoba
B2B SaaS

Soorat-e-haal

Challenge kya tha

One Node server handled everything — WebSocket connections, task mutations, reward calculations and database writes. It was already spiking at 5,000 concurrent users, and the target was 100,000 within twelve months. The reward algorithm was the worst of it: every completion, comment, mention, reaction and login triggered a multi-step scoring pass across a user's entire history, synchronously, at peak.

Hum ne kya kiya

We took the reward engine off the critical path first, because it was the thing making every other interaction slow, then scaled the connection layer behind it.

Wo faisle jo aham the

  • The reward engine became a queue, not a function call

    BullMQ over Redis, with a dedicated worker consuming events and writing XP to PostgreSQL in batches. The UI updates instantly and the arithmetic settles behind it within 200ms p99 — users were never waiting on the scoring, they were only ever waiting on the architecture.

  • WebSockets across instances instead of inside one

    In-memory Socket.IO rooms replaced with a Redis Pub/Sub adapter, sticky sessions at the ALB and heartbeat pruning. That is what turned one box into twelve EC2 instances sharing 100,000+ connections.

  • Leaderboards stopped being computed live

    Materialized views refreshed every five minutes, served from read replicas. A leaderboard that is five minutes stale is indistinguishable from a live one to a user, and enormously cheaper.

Kya badla

concurrent users
100K+concurrent users
API p95 latency
85msAPI p95 latency
user retention
+340%user retention
uptime
99.99%uptime
  • 100K+ — from 5,000 before the rebuild
  • 85ms — down from 1.2 seconds

The platform carries 100,000+ concurrent users and 2M+ events a day at 99.99% uptime, with p95 API latency down from 1.2 seconds to 85ms and retention up 340%.

Istemal shuda khidmaat

  • Event-driven reward engine on BullMQ with batched writes
  • Horizontally scalable WebSocket layer across twelve instances
  • Read-replica and materialized-view strategy for leaderboards
  • Load-test harness and capacity plan to 100K concurrent

Hum kya mukhtalif karte

Har mansoobe mein aisi aik baat hoti hai. Ise shaya karna hi asal nukta hai — jis case study mein koi pachhtawa na ho wo saboot nahi, tashheer hai.

Knowing when to evolve an architecture without stopping product velocity is the actual skill. The original monolith was not a mistake — it was correct for its stage. The judgement is in avoiding premature optimisation and premature architectural pessimism at the same time, while continuing to ship features throughout.

Shuru karne ke liye tayyar hain?