Ensiklopedia VibeKoding: System Design Methodology.Ensiklopedia VibeKoding: System Design Methodology.
System design is not about sketching architecture diagrams on a whim โ it's a structured methodology. Whether it's a system design interview question or real-world architecture design, both follow a similar thinking framework: first understand the problem, then estimate the scale, then design the solution, and finally dive deep into optimization.System design is not about sketching architecture diagrams on a whim โ it's a structured methodology. Whether it's a system design interview question or real-world architecture design, both follow a similar thinking framework: first understand the problem, then estimate the scale, then design the solution, and finally dive deep into optimization.
What will you learn from this article?What will you learn from this article?
After reading this chapter, you will gain:After reading this chapter, you will gain:
| Chapter | Content | Core Concepts |
|---|---|---|
| Chapter 1 | Four-Step Design Method | Requirements clarification, capacity estimation, architecture design, deep optimization |
| Chapter 2 | Capacity Estimation | QPS, storage, bandwidth, back-of-envelope estimation |
| Chapter 3 | Core Design Patterns | Caching, database sharding, message queues, CDN |
| Chapter 4 | Trade-off Thinking | Consistency vs. availability, performance vs. cost |
| Chapter 5 | Classic Case Studies | URL shortener, feed system, flash sale system |
------
System design is not about drawing architecture diagrams right away. Whether in an interview or in practice, you should follow a structured process.System design is not about drawing architecture diagrams right away. Whether in an interview or in practice, you should follow a structured process.
Many people start drawing diagrams as soon as they get the prompt, only to design a system that is "correct but not what the interviewer wanted." Spending 5 minutes clarifying requirements can prevent 30 minutes of rework later. Common clarification questions: - What are the core features of the system? (Don't design every feature) - What is the user scale? (Determines whether distribution is needed) - What is the read/write ratio? (Determines caching strategy) - How long does data need to be retained? (Determines the storage solution)Many people start drawing diagrams as soon as they get the prompt, only to design a system that is "correct but not what the interviewer wanted." Spending 5 minutes clarifying requirements can prevent 30 minutes of rework later. Common clarification questions: - What are the core features of the system? (Don't design every feature) - What is the user scale? (Determines whether distribution is needed) - What is the read/write ratio? (Determines caching strategy) - How long does data need to be retained? (Determines the storage solution)
------
"Back-of-envelope estimation" is a core skill in system design. You don't need precise calculations โ just knowing the order of magnitude is enough."Back-of-envelope estimation" is a core skill in system design. You don't need precise calculations โ just knowing the order of magnitude is enough.
| Magnitude | Conversion | Memory Trick |
|---|---|---|
| 1 day | 86,400 seconds | โ 100K seconds |
| 100M requests/day | โ 1,200 QPS | Divide by 100K |
| 1 KB ร 100M | โ 100 GB | 100M small records |
| 1 MB ร 1M | โ 1 TB | 1M images |
Most systems follow the 80/20 rule: 20% of the data handles 80% of the requests. This means:Most systems follow the 80/20 rule: 20% of the data handles 80% of the requests. This means:
------
Patterns that appear repeatedly in system design โ mastering these will prepare you for most scenarios.Patterns that appear repeatedly in system design โ mastering these will prepare you for most scenarios.
| Pattern | Read Path | Write Path | Use Cases |
|---|---|---|---|
| Cache-Aside | Check cache first; on miss, query DB and backfill | Write DB first, then invalidate cache | General purpose, most commonly used |
| Read-Through | Cache layer automatically loads from DB | Same as Cache-Aside | Requires caching framework support |
| Write-Behind | Same as Cache-Aside | Write to cache first, async write to DB | Write-heavy, can tolerate data loss |
Updating the cache is prone to data inconsistency in concurrent scenarios: threads A and B update simultaneously, A writes to DB first but B updates the cache first, resulting in B's stale value in the cache. Invalidating the cache causes the next read request to reload from DB, naturally avoiding this problem.Updating the cache is prone to data inconsistency in concurrent scenarios: threads A and B update simultaneously, A writes to DB first but B updates the cache first, resulting in B's stale value in the cache. Invalidating the cache causes the next read request to reload from DB, naturally avoiding this problem.
When a single table exceeds tens of millions of rows, or when a single database's QPS hits a bottleneck, it's time to consider database sharding.When a single table exceeds tens of millions of rows, or when a single database's QPS hits a bottleneck, it's time to consider database sharding.
| Strategy | Approach | Advantages | Disadvantages |
|---|---|---|---|
| Vertical sharding | Split databases by business domain | Business decoupling, independent scaling | Cross-database JOINs are difficult |
| Horizontal sharding | Split one table into multiple tables by rule | Controllable data volume per table | Shard key selection is critical |
| Vertical table splitting | Split large columns into a separate table | Reduces I/O, improves query efficiency | Requires additional JOINs |
Shard Key Selection Principles:Shard Key Selection Principles:
Message queues are the "shock absorbers" of distributed systems. Their core roles are decoupling, async processing, and peak shaving.Message queues are the "shock absorbers" of distributed systems. Their core roles are decoupling, async processing, and peak shaving.
| Scenario | Without Queue | With Queue |
|---|---|---|
| Send notification after order | Order API calls notification service synchronously; notification failure causes order failure | Send message after order success; notification service consumes asynchronously |
| Flash sale | Burst traffic overwhelms the database | Requests enter queue first; backend processes at its own pace |
| Data synchronization | Service A calls Service B's API directly | Service A publishes event; Service B subscribes and handles it |
------
The essence of architecture design is trade-offs. Every decision has a cost โ the key is understanding the cost and making choices appropriate for the current stage.The essence of architecture design is trade-offs. Every decision has a cost โ the key is understanding the cost and making choices appropriate for the current stage.
| Trade-off Dimension | Option A | Option B | Decision Basis |
|---|---|---|---|
| Consistency vs. Availability | Strong consistency (CP) | High availability (AP) | Can the business tolerate brief inconsistency? |
| Performance vs. Cost | Full caching | On-demand caching | Data volume and budget |
| Simplicity vs. Flexibility | Monolithic architecture | Microservices | Team size and business complexity |
| Real-time vs. Batch | Stream processing | Batch processing | Data timeliness requirements |
| Self-managed vs. Hosted | Build your own MySQL | Use cloud database RDS | Operations capability and cost |
Every important architecture decision should be documented: what was the context, what options were considered, why this one was chosen, and what the trade-offs are. This isn't about assigning blame โ it's about helping future teams understand "why it was designed this way." The format is simple: - Title: Using X instead of Y - Context: What problem we encountered - Decision: What solution we chose - Rationale: Why we chose this - Consequences: The drawbacks and risks of this decisionEvery important architecture decision should be documented: what was the context, what options were considered, why this one was chosen, and what the trade-offs are. This isn't about assigning blame โ it's about helping future teams understand "why it was designed this way." The format is simple: - Title: Using X instead of Y - Context: What problem we encountered - Decision: What solution we chose - Rationale: Why we chose this - Consequences: The drawbacks and risks of this decision
| Mistake | Manifestation | Correct Approach |
|---|---|---|
| Premature optimization | Sharding at 1,000 daily active users | Start with a single database; shard when you hit bottlenecks |
| Technology-driven | "I want to use Kafka" instead of "I need async processing" | Start from the problem, not the technology |
| Ignoring operations cost | Choosing the optimal solution that the team can't maintain | Solutions must match team capability |
| Pursuing perfect consistency | Using distributed transactions for every scenario | Eventual consistency is sufficient for most scenarios |
------
Let's connect the methodology we've learned through three classic examples.Let's connect the methodology we've learned through three classic examples.
The URL shortener is a classic system design interview question โ small but comprehensive.The URL shortener is a classic system design interview question โ small but comprehensive.
Requirements Clarification:Requirements Clarification:
Capacity Estimation:Capacity Estimation:
| Metric | Calculation | Result |
|---|---|---|
| Write QPS | 100M / 100 / 86,400 | โ 12 QPS |
| Read QPS | 100M / 86,400 | โ 1,200 QPS |
| Peak read QPS | 1,200 ร 3 | โ 3,600 QPS |
| 5-year storage | 1M/day ร 365 ร 5 ร 100B | โ 18 GB |
| Cache (20%) | 18 GB ร 20% | โ 3.6 GB |
Architecture Design:Architecture Design:
CODE Write path: Client โ API Server โ ID Generator โ Base62 Encode โ Write to MySQL + Redis Read path: Client โ CDN โ API Server โ Redis lookup โ 302 redirect โ (cache miss) MySQL query โ backfill Redis
Key Design Decisions:Key Design Decisions:
Social platform feeds (WeChat Moments, Twitter home timeline) are another classic question.Social platform feeds (WeChat Moments, Twitter home timeline) are another classic question.
Core Challenge: When a user publishes a post, how do all their followers see it?Core Challenge: When a user publishes a post, how do all their followers see it?
| Approach | How It Works | Advantages | Disadvantages |
|---|---|---|---|
| Pull model | Aggregate followees' posts in real time at read time | Simple writes, less storage | Slow reads; high latency with many followees |
| Push model | Write to all followers' inboxes at publish time | Extremely fast reads | Severe write amplification for accounts with many followers |
| Hybrid (Push-Pull) | Push for regular users, pull for celebrities | Balanced read/write performance | Complex implementation |
Hybrid Push-Pull Approach:Hybrid Push-Pull Approach:
The core challenge of a flash sale: instant ultra-high concurrency + inventory must not be oversold.The core challenge of a flash sale: instant ultra-high concurrency + inventory must not be oversold.
Traffic Characteristics:Traffic Characteristics:
Layered Peak Shaving Strategy:Layered Peak Shaving Strategy:
CODE User request โ CDN (static pages) โ Gateway (rate limiting) โ Message queue (peak shaving) โ Inventory service (deduction)
| Layer | Strategy | Effect |
|---|---|---|
| Frontend | Button gray-out + random delay + CAPTCHA | Filters bots, disperses requests |
| CDN | Static resource caching | Reduces 90% of page requests |
| Gateway | Token bucket rate limiting | Only allows traffic the system can handle |
| Message queue | Requests queued, processed asynchronously | Peak shaving, protects the database |
| Inventory service | Redis pre-deduction + Lua atomic operations | Prevents overselling, millisecond response |
1. Intercept upstream whenever possible: If you can block it at the CDN, don't let it reach the application layer 2. Separate reads and writes: Product detail pages use cache; only orders go to the database 3. Async processing: After the user clicks "buy," immediately return "queuing" and process in the background 4. Fallback plans: Rate limiting, circuit breaking, degradation โ every layer needs a Plan B1. Intercept upstream whenever possible: If you can block it at the CDN, don't let it reach the application layer 2. Separate reads and writes: Product detail pages use cache; only orders go to the database 3. Async processing: After the user clicks "buy," immediately return "queuing" and process in the background 4. Fallback plans: Rate limiting, circuit breaking, degradation โ every layer needs a Plan B
------
System design is a highly practical skill. The core lies in structured thinking and making trade-offs.System design is a highly practical skill. The core lies in structured thinking and making trade-offs.
Key takeaways from this chapter:Key takeaways from this chapter: