Ensiklopedia VibeKoding: Principles of Load Balancing and Gateways.Ensiklopedia VibeKoding: Principles of Load Balancing and Gateways.
When a single server can't handle the load, how do you "smartly" distribute traffic across multiple server instances? Load balancing is the "dispatcher" of modern distributed systems. This article uses real-world analogies (bubble tea shop checkout, package sorting, traffic control) to deeply explore the design philosophy and engineering practices of load balancing.When a single server can't handle the load, how do you "smartly" distribute traffic across multiple server instances? Load balancing is the "dispatcher" of modern distributed systems. This article uses real-world analogies (bubble tea shop checkout, package sorting, traffic control) to deeply explore the design philosophy and engineering practices of load balancing.
------
A startup encountered severe performance issues as its user base grew rapidly:A startup encountered severe performance issues as its user base grew rapidly:
Scenario:Scenario:
CODE Phase 1: Single Server Users โ Server (1 vCPU, 2 GB RAM) โ 1,000 DAU โ Peak: 1,000 concurrent users โ Problem: CPU at 100%, slow responses, frequent crashes
- Performance bottleneck: CPU at 100%, response time > 5 seconds - Single point of failure: If the server goes down, the entire site is unavailable - Scaling difficulty: Only vertical scaling is possible (adding CPU, RAM), which is expensive and limited- Performance bottleneck: CPU at 100%, response time > 5 seconds - Single point of failure: If the server goes down, the entire site is unavailable - Scaling difficulty: Only vertical scaling is possible (adding CPU, RAM), which is expensive and limited
Improved Architecture (with Load Balancing):Improved Architecture (with Load Balancing):
CODE Phase 2: Multiple Servers + Load Balancer Users โ Load Balancer (Nginx) โ โโ Server 1 (1 vCPU, 2 GB RAM) โโ Server 2 (1 vCPU, 2 GB RAM) โโ Server 3 (1 vCPU, 2 GB RAM)
- Better performance: 3 servers processing in parallel, response time < 1 second - High availability: If one server fails, others continue serving - Horizontal scaling: Need more capacity? Just add more servers- Better performance: 3 servers processing in parallel, response time < 1 second - High availability: If one server fails, others continue serving - Horizontal scaling: Need more capacity? Just add more servers
The Bubble Tea Shop CounterThe Bubble Tea Shop Counter
Imagine you run a popular bubble tea shop:Imagine you run a popular bubble tea shop:
The load balancer is the "counter assignment person":The load balancer is the "counter assignment person":
------
Operates at the transport layer (TCP/UDP) โ like a delivery driver who only looks at your address (IP + port) without caring about who you are or what you do.Operates at the transport layer (TCP/UDP) โ like a delivery driver who only looks at your address (IP + port) without caring about who you are or what you do.
Characteristics:Characteristics:
`` Client request โ L4 Load Balancer โ Backend server โ Only looks at IP + Port โ Fast forwarding (no content inspection) ```` Client request โ L4 Load Balancer โ Backend server โ Only looks at IP + Port โ Fast forwarding (no content inspection) ``
Operates at the application layer (HTTP/HTTPS) โ like a delivery driver who not only checks the address but also opens the package to inspect its contents before deciding how to deliver.Operates at the application layer (HTTP/HTTPS) โ like a delivery driver who not only checks the address but also opens the package to inspect its contents before deciding how to deliver.
Characteristics:Characteristics:
`` Client request โ L7 Load Balancer โ Parses HTTP content โ Inspects URL, Header, Cookie โ Intelligent routing to specific server ```` Client request โ L7 Load Balancer โ Parses HTTP content โ Inspects URL, Header, Cookie โ Intelligent routing to specific server ``
| Dimension | L4 Load Balancing | L7 Load Balancing |
|---|---|---|
| OSI Layer | Transport Layer (TCP/UDP) | Application Layer (HTTP/HTTPS) |
| Routing Basis | IP address + port | URL, Header, Cookie, Body |
| Processing Speed | Extremely fast (kernel-space) | Fast (user-space parsing) |
| Feature Richness | Basic forwarding | SSL termination, caching, compression, WAF |
| Typical Scenarios | Databases, gaming, long connections | Web apps, API gateways, microservices |
| Representative Products | LVS, AWS NLB | Nginx, HAProxy, AWS ALB |
------
Imagine one of your checkout counters breaks, but the assignment person doesn't know and keeps sending customers there. The queue grows longer, and customers grow angrier.Imagine one of your checkout counters breaks, but the assignment person doesn't know and keeps sending customers there. The queue grows longer, and customers grow angrier.
Health checks are the "sentinels" that prevent this scenario. They periodically "examine" each server, immediately removing any "sick" ones from the pool and bringing them back once they "recover."Health checks are the "sentinels" that prevent this scenario. They periodically "examine" each server, immediately removing any "sick" ones from the pool and bringing them back once they "recover."
Active Health Check: The load balancer actively "knocks on the door" asking the server, "Are you still there?"Active Health Check: The load balancer actively "knocks on the door" asking the server, "Are you still there?"
Passive Health Check: The load balancer "observes" the response patterns of real business trafficPassive Health Check: The load balancer "observes" the response patterns of real business traffic
| Metric | Healthy Threshold | Unhealthy Threshold | Notes | |:---|:---|:---|:---| | HTTP Status Code | 200-399 | 400+ or timeout | 4xx/5xx are all considered failures | | TCP Connection | Successfully established | Connection timeout | Checks whether the port is reachable | | Response Time | < 500 ms | > 2000 ms | Timeout typically set to 2-5 seconds | | Consecutive Failures | - | 3 times | Avoids false positives from transient blips | | Check Interval | - | 5 s | Too frequent increases load || Metric | Healthy Threshold | Unhealthy Threshold | Notes | |:---|:---|:---|:---| | HTTP Status Code | 200-399 | 400+ or timeout | 4xx/5xx are all considered failures | | TCP Connection | Successfully established | Connection timeout | Checks whether the port is reachable | | Response Time | < 500 ms | > 2000 ms | Timeout typically set to 2-5 seconds | | Consecutive Failures | - | 3 times | Avoids false positives from transient blips | | Check Interval | - | 5 s | Too frequent increases load |
A team set the health check response time threshold to 100ms, but their application's average response time fluctuated between 80-120ms. As a result, servers were frequently marked as "unhealthy," causing traffic to oscillate between healthy and unhealthy states, and overall system availability actually dropped.A team set the health check response time threshold to 100ms, but their application's average response time fluctuated between 80-120ms. As a result, servers were frequently marked as "unhealthy," causing traffic to oscillate between healthy and unhealthy states, and overall system availability actually dropped.
The right approach: Thresholds should be set to 2-3x the P99 response time, leaving enough buffer for normal fluctuations.The right approach: Thresholds should be set to 2-3x the P99 response time, leaving enough buffer for normal fluctuations.
--- ## 4. Core Problem 2: Approach to Ensuring Returning Users Always Get "Backend Instance" ### 4.1 Session Persistence: Let "Returning User" Always Find "Backend Instance" Imagine you're a regular at a bubble tea shop, and the same staff member serves you every time. She knows your preferences (half sugar, no ice) and serves you quickly and thoughtfully. But if you get a new person every time, you have to repeat the same requests over and over โ a huge efficiency loss. Session persistence (sticky sessions) solves this problem: ensuring that requests from the same user are always routed to the same backend server.
--- ## 5. Core Problem 3: Approach to achieving Zero-Downtime Deployment ### 5.1 Blue-Green Deployment: "One-Click Switch" for Zero-Downtime Releases Core idea: Maintain two identical production environments (blue and green) simultaneously, but only one serves live traffic at any given time.
| Pros | Cons |
|---|---|
| โ Zero downtime; switch completes in milliseconds | โ High resource cost; requires maintaining two environments |
| โ Fast rollback; switch back immediately if issues are found | โ Database schema changes require special compatibility handling |
| โ New environment can be fully tested before taking over traffic | โ Not suitable for stateful services (e.g., WebSocket long connections) |
### 5.2 Canary Release: "Small Steps, Fast Iteration" Canary Strategy The canary release is named after the historical "coal mine canary" โ miners brought canaries into the mines; if the canary showed signs of distress, it indicated toxic gas leakage, and miners would evacuate immediately. In software releases, a canary release means exposing a small subset of users to the new version first, observing for issues, and then gradually expanding the rollout.
| Advantage | Description |
|---|---|
| ๐ฏ Controlled risk | Even if the new version has severe bugs, only a small number of users are affected |
| ๐ Real-world validation | Validated in the real production environment, more reliable than staging |
| Fast iteration | Teams can release new features more confidently and frequently |
| ๐ฐ Resource-friendly | Doesn't require two complete environments like blue-green deployment |
--- ## 6. Core Problem 4: Approach to making the System "Auto-Scale" on Its Own ### 6.1 Auto Scaling: Let the System Flexibly Schedule Workloads Imagine you run a restaurant: - Lunch peak: You need 10 servers, but at 3 PM during the afternoon lull, you only need 2 - If you always keep 10: labor costs explode - If you always keep only 2: customers during peak hours can't wait and all leave Auto Scaling lets the system "flexibly schedule" like a restaurant โ automatically adding servers when busy and removing them when idle.
Pitfall 1: Scaling responds too slowly; the traffic surge already crashes the systemPitfall 1: Scaling responds too slowly; the traffic surge already crashes the system
During a major e-commerce promotion, the team set CPU > 80% as the scale-up trigger, but metric collection had a 1-minute delay, and new instance startup took 3 minutes. Traffic arrived too fast โ before scaling completed, the servers were already overwhelmed.During a major e-commerce promotion, the team set CPU > 80% as the scale-up trigger, but metric collection had a 1-minute delay, and new instance startup took 3 minutes. Traffic arrived too fast โ before scaling completed, the servers were already overwhelmed.
Solutions:Solutions:
Pitfall 2: Scaling is too aggressive; costs explodePitfall 2: Scaling is too aggressive; costs explode
A startup set an aggressive auto-scaling policy: scale up if CPU > 50%. As a result, a normal business fluctuation triggered scaling, and the server count ballooned from 5 to 30. The end-of-month cloud bill terrified the CTO.A startup set an aggressive auto-scaling policy: scale up if CPU > 50%. As a result, a normal business fluctuation triggered scaling, and the server count ballooned from 5 to 30. The end-of-month cloud bill terrified the CTO.
Solutions:Solutions:
Pitfall 3: Scaling down too fast; newly added machines are removed immediatelyPitfall 3: Scaling down too fast; newly added machines are removed immediately
A team set CPU < 30% as the scale-down trigger. After scaling up, traffic was still settling, and CPU briefly dropped to 25%, triggering a scale-down. Right after scaling down, CPU spiked back to 80%, triggering another scale-up โ the system oscillated wildly in a "scale-up, scale-down, scale-up" loop.A team set CPU < 30% as the scale-down trigger. After scaling up, traffic was still settling, and CPU briefly dropped to 25%, triggering a scale-down. Right after scaling down, CPU spiked back to 80%, triggering another scale-up โ the system oscillated wildly in a "scale-up, scale-down, scale-up" loop.
Solutions:Solutions:
--- ## 7. Practical Guide: Approach to choosing a Load Balancer ### 7.1 Comparison of Mainstream Load Balancers | Feature | Nginx | HAProxy | Envoy | Cloud Provider LB | | --------------------- | ------------------------------------ | -------------------------- | ------------------- | ------------------ | | Positioning | High-performance reverse proxy / LB | Open-source load balancer | Cloud-native proxy | Managed load balancer | | Performance | Extremely high (C, event-driven) | High (event-driven) | High (C++/Rust) | Extremely high | | Feature Richness | Basic LB, static files, caching | Rich LB algorithms | Advanced routing, observability | Full-featured | | Configuration | Config file (nginx.conf) | Config file (haproxy.cfg) | API / config file | UI console | | Extensibility | C modules / Lua scripts | Lua scripts | WASM / Filters | Plugins | | Use Cases | Static assets, L7 LB, SSL termination | L7 LB, high availability | Service mesh, multi-cloud | Quick start |--- ## 7. Practical Guide: Approach to choosing a Load Balancer ### 7.1 Comparison of Mainstream Load Balancers | Feature | Nginx | HAProxy | Envoy | Cloud Provider LB | | --------------------- | ------------------------------------ | -------------------------- | ------------------- | ------------------ | | Positioning | High-performance reverse proxy / LB | Open-source load balancer | Cloud-native proxy | Managed load balancer | | Performance | Extremely high (C, event-driven) | High (event-driven) | High (C++/Rust) | Extremely high | | Feature Richness | Basic LB, static files, caching | Rich LB algorithms | Advanced routing, observability | Full-featured | | Configuration | Config file (nginx.conf) | Config file (haproxy.cfg) | API / config file | UI console | | Extensibility | C modules / Lua scripts | Lua scripts | WASM / Filters | Plugins | | Use Cases | Static assets, L7 LB, SSL termination | L7 LB, high availability | Service mesh, multi-cloud | Quick start |
Decision Tree:Decision Tree:
CODE Choose a load balancer: โ โโ Only need basic L4 load balancing? โ โโ Yes โ LVS (open-source, free) or cloud provider NLB โ โโ No โ Continue โ โโ Need service mesh or multi-cloud deployment? โ โโ Yes โ Envoy โ โโ No โ Continue โ โโ Need extremely complex configuration and plugins? โ โโ Yes โ HAProxy โ โโ No โ Continue โ โโ Need high performance + simple configuration? โ โโ Yes โ Nginx (first choice) โ โโ Continue โ โโ Want managed operations? โ โโ Yes โ Cloud provider LB (AWS ALB, Alibaba Cloud SLB) โ โโ Self-host Nginx
--- ## 8. Summary: Core Mindset of Load Balancing ### 8.1 Core Principles Recap | Principle | Meaning | Key Practice Points | | --------------- | ---------------------------------------------- | ------------------------------------------------------------ | | Layering | L4 handles "package sorting" (fast but simple) | L4 for databases, gaming; L7 for web, APIs | | Redundancy | Single points of failure are the enemy | Improve availability through multi-instance, multi-region deployment | | Gradualism | Don't release new versions with "one big cut" | Blue-green deployment for zero downtime; canary for controlled risk | | Elasticity | The system should "breathe" like a living organism | Automatically scale up when busy, scale down when idle | ### 8.2 Design Checklist Before introducing load balancing, ask yourself the following questions: - [ ] Is load balancing really needed? (Is single-machine performance truly insufficient?) - [ ] Choose L4 or L7? (Based on business scenario) - [ ] How to handle session persistence? (Cookie, IP hash, session table) - [ ] How to implement health checks? (Active, passive, threshold settings) - [ ] How to achieve zero downtime? (Blue-green deployment, canary) - [ ] How to implement elasticity? (Scaling metrics, cooldown periods, max instance count) --- ## 9. Glossary | Term | Chinese Translation | Explanation | | ----------------------------- | ------------------- | ------------------------------------------------------------------------------------- | | Load Balancer | ่ด่ฝฝๅ่กกๅจ | A device or software that distributes traffic across multiple backend servers | | L4 Load Balancing | ๅๅฑ่ด่ฝฝๅ่กก | Load balancing based on the transport layer (TCP/UDP) | | L7 Load Balancing | ไธๅฑ่ด่ฝฝๅ่กก | Load balancing based on the application layer (HTTP/HTTPS) | | Health Check | ๅฅๅบทๆฃๆฅ | A mechanism that periodically checks the health status of backend servers | | Session Persistence | ไผ่ฏไฟๆ | Ensures requests from the same user are always routed to the same server | | Sticky Session | ็ฒๆงไผ่ฏ | Another term for Session Persistence | | Blue-Green Deployment | ่็ปฟ้จ็ฝฒ | A zero-downtime release strategy using two environments that switch over | | Canary Release | ้ไธ้ๅๅธ | A canary release strategy that validates with a small traffic portion first | | Auto Scaling | ่ชๅจๆฉ็ผฉๅฎน | Automatically increasing or decreasing the number of servers based on load | | Horizontal Scaling | ๆฐดๅนณๆฉๅฑ | Increasing server count to improve processing capacity | | Vertical Scaling | ๅ็ดๆฉๅฑ | Upgrading individual machine specs (CPU, RAM) to improve processing capacity | | Multi-Region | ๅคๅบๅ | Deploying services across multiple geographic regions | | Active-Active | ๅคๆดป | Multiple regions simultaneously serving traffic | | Active-Standby | ไธปๅค | Only one region serves traffic; others are on standby | | Data Replication | ๆฐๆฎๅๆญฅ | Cross-region data replication mechanism | | RTO | ๆขๅคๆถ้ด็ฎๆ | Recovery Time Objective โ the target time within which a system must recover after failure | | RPO | ๆขๅค็น็ฎๆ | Recovery Point Objective โ the acceptable amount of data loss after a system failure |--- ## 8. Summary: Core Mindset of Load Balancing ### 8.1 Core Principles Recap | Principle | Meaning | Key Practice Points | | --------------- | ---------------------------------------------- | ------------------------------------------------------------ | | Layering | L4 handles "package sorting" (fast but simple) | L4 for databases, gaming; L7 for web, APIs | | Redundancy | Single points of failure are the enemy | Improve availability through multi-instance, multi-region deployment | | Gradualism | Don't release new versions with "one big cut" | Blue-green deployment for zero downtime; canary for controlled risk | | Elasticity | The system should "breathe" like a living organism | Automatically scale up when busy, scale down when idle | ### 8.2 Design Checklist Before introducing load balancing, ask yourself the following questions: - [ ] Is load balancing really needed? (Is single-machine performance truly insufficient?) - [ ] Choose L4 or L7? (Based on business scenario) - [ ] How to handle session persistence? (Cookie, IP hash, session table) - [ ] How to implement health checks? (Active, passive, threshold settings) - [ ] How to achieve zero downtime? (Blue-green deployment, canary) - [ ] How to implement elasticity? (Scaling metrics, cooldown periods, max instance count) --- ## 9. Glossary | Term | Chinese Translation | Explanation | | ----------------------------- | ------------------- | ------------------------------------------------------------------------------------- | | Load Balancer | ่ด่ฝฝๅ่กกๅจ | A device or software that distributes traffic across multiple backend servers | | L4 Load Balancing | ๅๅฑ่ด่ฝฝๅ่กก | Load balancing based on the transport layer (TCP/UDP) | | L7 Load Balancing | ไธๅฑ่ด่ฝฝๅ่กก | Load balancing based on the application layer (HTTP/HTTPS) | | Health Check | ๅฅๅบทๆฃๆฅ | A mechanism that periodically checks the health status of backend servers | | Session Persistence | ไผ่ฏไฟๆ | Ensures requests from the same user are always routed to the same server | | Sticky Session | ็ฒๆงไผ่ฏ | Another term for Session Persistence | | Blue-Green Deployment | ่็ปฟ้จ็ฝฒ | A zero-downtime release strategy using two environments that switch over | | Canary Release | ้ไธ้ๅๅธ | A canary release strategy that validates with a small traffic portion first | | Auto Scaling | ่ชๅจๆฉ็ผฉๅฎน | Automatically increasing or decreasing the number of servers based on load | | Horizontal Scaling | ๆฐดๅนณๆฉๅฑ | Increasing server count to improve processing capacity | | Vertical Scaling | ๅ็ดๆฉๅฑ | Upgrading individual machine specs (CPU, RAM) to improve processing capacity | | Multi-Region | ๅคๅบๅ | Deploying services across multiple geographic regions | | Active-Active | ๅคๆดป | Multiple regions simultaneously serving traffic | | Active-Standby | ไธปๅค | Only one region serves traffic; others are on standby | | Data Replication | ๆฐๆฎๅๆญฅ | Cross-region data replication mechanism | | RTO | ๆขๅคๆถ้ด็ฎๆ | Recovery Time Objective โ the target time within which a system must recover after failure | | RPO | ๆขๅค็น็ฎๆ | Recovery Point Objective โ the acceptable amount of data loss after a system failure |