How to Design Music Streaming Platform like Spotify
Building audio streaming, personalized playlists, and social features at 600M+ user scale — A Senior+ System Design Guide
Table of Contents
- Introduction — The Scale of Spotify
- Requirements Gathering
- Capacity Estimation
- Data Model
- API Design
- High-Level Architecture
- Audio Streaming Pipeline
- Music Storage & CDN
- Search & Discovery
- Recommendation Engine
- Playlist Management
- Social Features
- Podcast Platform
- Offline Mode & Caching
- Audio Quality & Codec Selection
- Real-Time Listening Analytics
- Artist Dashboard & Royalties
- Rights Management & Licensing
- Ad-Supported Tier
- Database Design & Sharding
- Caching Strategy
- Multi-Region Design
- Cost Estimation
- Interview Q&A
- Full C# Implementation
- Conclusion
1. Introduction — The Scale of Spotify
Spotify has fundamentally transformed how the world consumes music. With over 600 million users, 226 million premium subscribers, and a catalog of 100+ million tracks and 80+ million podcast episodes, Spotify is the largest audio streaming platform on the planet. The platform serves 5 billion playlists and processes over 4 billion hours of audio streamed per quarter.
What makes Spotify's engineering fascinating isn't just the scale — it's the combination of real-time audio streaming, machine learning-powered recommendations, social features, and a complex royalty payment system that must work flawlessly across 184 markets and 40+ languages.
- 600M+ total users (226M+ premium subscribers)
- 100M+ music tracks in the catalog
- 80M+ podcast episodes
- 5B+ user-generated playlists
- 11M+ artist and creator accounts
- 4B+ hours of audio streamed per quarter
- 184 markets, 40+ languages
- ~$14B annual revenue (2025)
- 99.99%+ uptime SLA for audio delivery
- Sub-100ms time-to-first-byte for audio chunks
Spotify runs on a massive microservices architecture on Google Cloud Platform, processing terabytes of event data daily. Their internal back-end framework, Loki, handles tens of thousands of requests per second across hundreds of services. The recommendation system processes listening behavior from hundreds of millions of users to deliver personalized experiences that keep users engaged for an average of 30+ minutes per session.
In this comprehensive system design guide, we'll break down every major component — from adaptive bitrate audio streaming and content delivery to collaborative filtering recommendation engines and real-time royalty calculations. This article targets Senior+ engineers preparing for system design interviews or architects evaluating music streaming platforms.
2. Requirements Gathering
Functional Requirements
Core Streaming
- F1: Users can search for songs, artists, albums, playlists, and podcasts
- F2: Users can play audio tracks with adaptive bitrate streaming
- F3: Users can create, edit, and share playlists
- F4: Users can follow artists, other users, and curated playlists
- F5: System provides personalized recommendations (Discover Weekly, Release Radar, Daily Mixes)
- F6: Users can download tracks for offline listening (Premium only)
- F7: Users can view friend activity and sharing
- F8: System supports podcast playback with episode management
- F9: Users can like/save tracks to their library
- F10: Cross-device playback (Spotify Connect)
Creator & Business
- F11: Artists can upload and manage their catalog
- F12: Artists get analytics dashboard (streams, listeners, demographics)
- F13: Royalty calculation and payment system
- F14: Ad-serving for free-tier users
- F15: Rights management across territories
Non-Functional Requirements
| Property | Target | Rationale |
|---|---|---|
| Availability | 99.99% (52 min downtime/year) | Audio playback must be uninterrupted globally |
| Latency (API) | <100ms P99 | Search, metadata, and UI responsiveness |
| Latency (Audio) | <200ms time-to-first-byte | Smooth playback start experience |
| Throughput | 10M+ concurrent streams | Peak evening hours across all markets |
| Durability | 99.999999999% (11 nines) | Audio files and user data must never be lost |
| Consistency | Eventual (metadata); Strong (payments) | Playlists can lag; royalty math must be exact |
| Scalability | Linear with users | Add resources proportionally to growth |
| Offline | Full catalog access for downloaded content | Premium users expect offline playback |
3. Capacity Estimation
Bandwidth & Storage Calculations
Audio Storage
Assume an average track duration of 3.5 minutes:
- Ogg Vorbis quality 5 (160 kbps): 160,000 bits/s x 210s = 33.6 Mbit = ~4.2 MB per track
- FLAC lossless (1000 kbps): 1,000,000 bits/s x 210s = 210 Mbit = ~26.25 MB per track
- 100M tracks x 3 qualities x 4.2 MB average: ~1.26 PB raw audio storage
- With metadata, artwork, and podcasts: ~2-3 PB total storage
Bandwidth Estimation
- 600M users x 30 min/day average: 18B minutes/day
- Concurrent (peak): ~10M streams simultaneously
- Per stream at 160 kbps: 160 kbps = 20 KB/s
- Peak bandwidth: 10M x 20 KB/s = 200 GB/s = 1.6 Tbps
- Daily CDN transfer: 18B min x 4.2 MB / 210s = ~360 TB/day through CDN
Request Rate Estimation
- API requests: 600M users x 100 requests/day = 60B requests/day = ~700K RPS
- Search queries: ~20% of API traffic = ~140K QPS
- Audio chunk requests: Each chunk is 10-30 seconds of audio. At 10M concurrent streams, refreshing every 10s = 1M chunk requests/s
- Event ingestion (listening events): 10M concurrent x ~1 event/30s = ~333K events/s
Data Size Summary
| Data Type | Size | Notes |
|---|---|---|
| Audio files (multi-quality) | ~2-3 PB | OGG, AAC, FLAC variants |
| Album artwork | ~50-100 TB | Multiple resolutions per album |
| Podcast episodes | ~500 TB | Growing ~1M episodes/year |
| User data (profiles, playlists) | ~10-20 TB | Relational + document stores |
| Listening history | ~500 TB/year | Append-heavy, time-series |
| Search index | ~5-10 TB | Inverted index for catalog |
| Recommendation models | ~100 GB | ML model weights + embeddings |
| Real-time event stream | ~1-2 PB/year | Kafka topics, retained 7-30 days |