
MagicTask · Project management SaaS
From latency spikes at 5,000 users to 100,000 concurrent
A gamified project management platform re-architected from a single Node server into an event-driven system that holds 100,000+ concurrent connections without the reward engine touching the critical path.
لمحة سريعة
- المدة
- 78 أسبوعاً
- حجم الفريق
- 8 أشخاص
- نوع التعاقد
- إعادة بناء المنصّة
- نوع المشروع
- تطبيق ويب
- القطاع
- B2B SaaS
الوضع
التحدي
One Node server handled everything — WebSocket connections, task mutations, reward calculations and database writes. It was already spiking at 5,000 concurrent users, and the target was 100,000 within twelve months. The reward algorithm was the worst of it: every completion, comment, mention, reaction and login triggered a multi-step scoring pass across a user's entire history, synchronously, at peak.
ما الذي فعلناه
We took the reward engine off the critical path first, because it was the thing making every other interaction slow, then scaled the connection layer behind it.
القرارات التي صنعت الفارق
The reward engine became a queue, not a function call
BullMQ over Redis, with a dedicated worker consuming events and writing XP to PostgreSQL in batches. The UI updates instantly and the arithmetic settles behind it within 200ms p99 — users were never waiting on the scoring, they were only ever waiting on the architecture.
WebSockets across instances instead of inside one
In-memory Socket.IO rooms replaced with a Redis Pub/Sub adapter, sticky sessions at the ALB and heartbeat pruning. That is what turned one box into twelve EC2 instances sharing 100,000+ connections.
Leaderboards stopped being computed live
Materialized views refreshed every five minutes, served from read replicas. A leaderboard that is five minutes stale is indistinguishable from a live one to a user, and enormously cheaper.
ما الذي تغيّر
- concurrent users
- 100K+concurrent users
- API p95 latency
- 85msAPI p95 latency
- user retention
- +340%user retention
- uptime
- 99.99%uptime
- 100K+ — from 5,000 before the rebuild
- 85ms — down from 1.2 seconds
The platform carries 100,000+ concurrent users and 2M+ events a day at 99.99% uptime, with p95 API latency down from 1.2 seconds to 85ms and retention up 340%.
الخدمات المستخدمة
- Event-driven reward engine on BullMQ with batched writes
- Horizontally scalable WebSocket layer across twelve instances
- Read-replica and materialized-view strategy for leaderboards
- Load-test harness and capacity plan to 100K concurrent
ما الذي كنا سنفعله بشكل مختلف
لكل مشروع واحدة من هذه. ونشرها هو المقصد — فدراسة حالة بلا ندم فيها تسويق لا دليل.
Knowing when to evolve an architecture without stopping product velocity is the actual skill. The original monolith was not a mistake — it was correct for its stage. The judgement is in avoiding premature optimisation and premature architectural pessimism at the same time, while continuing to ship features throughout.
الخدمات المستخدمة
تطوير الويب
مواقع تركّز على التحويل، سريعة التحميل ومتصدّرة في نتائج البحث.
استكشف الخدمةالسحابة وDevOps
أطلق كل يوم، على بنية تحتية تكلّف ما ينبغي لها.
استكشف الخدمةالبيانات والتحليلات
أرقام يثق بها الجميع، في مكان واحد، تُحدَّث ليلاً.
استكشف الخدمةالدعم التقني المُدار
رقم واحد تتصل به، وفريق يعرف منظومتك مسبقاً.
استكشف الخدمة
أعمال أخرى
- Internet Money · Crypto financial services$10M+monthly transaction volume
One wallet interface across five blockchains
A multi-chain crypto platform where every chain hides behind one interface, with a tamper-evident audit trail underneath it. $10M+ moves through it monthly.
اقرأ دراسة الحالة - Bitnob · Consumer fintech<3scross-border settlement
Cross-border settlement in under three seconds
A regulated savings and payments platform across 10+ African markets, settling over Bitcoin Lightning in seconds rather than the days a SWIFT transfer takes — on a double-entry ledger that has never disagreed with itself.
اقرأ دراسة الحالة