Content par jayen

Education technology

Reference build

Fifty thousand courses, two hundred providers, one schema

A course discovery engine that normalises 200+ provider APIs into one canonical schema and ranks them with a scoring model tuned against real click behaviour.

50K+courses indexed

Aik nazar mein

Muddat
52 hafte
Team ka hajm
5 afraad
Muahide ki naueeyat
Nayi tameer
Project ki qism
Data platform
Shoba
Education

Soorat-e-haal

Challenge kya tha

A Coursera specialization and a Udemy course are structurally different products, and 200+ providers meant 200+ inconsistent schemas updating on their own cadences. On top of normalising all of that, search had to feel like Google.

Hum ne kya kiya

A plugin ingestion pipeline so each provider's weirdness stays contained, then a canonical schema and a scoring model that learns from behaviour.

Wo faisle jo aham the

  • One adapter per provider, one schema behind them

    Provider-specific adapters own auth, rate limiting and normalisation; everything downstream sees a canonical course with provider quirks in JSONB. Deduplication happens by fingerprint before indexing.

  • BM25 as a starting point, not an answer

    Custom Elasticsearch scoring layering rating, freshness, provider reputation and enrolment velocity on top of BM25 — with the weights set by A/B test rather than intuition.

Kya badla

courses indexed
50K+courses indexed
monthly active learners
500K+monthly active learners
search p95 latency
94mssearch p95 latency
recommendation click-through
3.2xrecommendation click-through
  • 50K+ — with under 0.1% normalisation errors
  • 3.2x — after A/B tested scoring weights

50,000+ courses from 200+ providers at under 0.1% normalisation error, 94ms p95 search, and a 3.2x lift in recommendation click-through.

Istemal shuda khidmaat

  • Plugin ingestion pipeline with 200+ provider adapters
  • Canonical course schema with fingerprint deduplication
  • Custom Elasticsearch scoring model with A/B tested weights

Hum kya mukhtalif karte

Har mansoobe mein aisi aik baat hoti hai. Ise shaya karna hi asal nukta hai — jis case study mein koi pachhtawa na ho wo saboot nahi, tashheer hai.

Relevance is a product problem before it is an infrastructure one. Instrumenting what people actually clicked turned out to be worth more than any further tuning of the retrieval layer.

Shuru karne ke liye tayyar hain?