VibeKoding / Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat / Principles of Load Balancing and GatewaysPrinciples of Load Balancing and Gateways
VK

Principles of Load Balancing and GatewaysPrinciples of Load Balancing and Gateways

๐Ÿ“š Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat ๐ŸŒ Dual Bahasa (ID / EN) โšก VibeKoding Native

Ensiklopedia VibeKoding: Principles of Load Balancing and Gateways.Ensiklopedia VibeKoding: Principles of Load Balancing and Gateways.

๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

When a single server can't handle the load, how do you "smartly" distribute traffic across multiple server instances? Load balancing is the "dispatcher" of modern distributed systems. This article uses real-world analogies (bubble tea shop checkout, package sorting, traffic control) to deeply explore the design philosophy and engineering practices of load balancing.When a single server can't handle the load, how do you "smartly" distribute traffic across multiple server instances? Load balancing is the "dispatcher" of modern distributed systems. This article uses real-world analogies (bubble tea shop checkout, package sorting, traffic control) to deeply explore the design philosophy and engineering practices of load balancing.

------

1. Motivation for the "Load Balancing"1. Motivation for the "Load Balancing"

1.1 A Real-World Case: The Architecture Evolution of a Website1.1 A Real-World Case: The Architecture Evolution of a Website

A startup encountered severe performance issues as its user base grew rapidly:A startup encountered severe performance issues as its user base grew rapidly:

Scenario:Scenario:

CODE
Phase 1: Single Server Users โ†’ Server (1 vCPU, 2 GB RAM) โ†“ 1,000 DAU โ†’ Peak: 1,000 concurrent users โ†“ Problem: CPU at 100%, slow responses, frequent crashes
โš ๏ธ Catatan Keamanan / Peringatanโš ๏ธ Warning / Security Note

- Performance bottleneck: CPU at 100%, response time > 5 seconds - Single point of failure: If the server goes down, the entire site is unavailable - Scaling difficulty: Only vertical scaling is possible (adding CPU, RAM), which is expensive and limited- Performance bottleneck: CPU at 100%, response time > 5 seconds - Single point of failure: If the server goes down, the entire site is unavailable - Scaling difficulty: Only vertical scaling is possible (adding CPU, RAM), which is expensive and limited

Improved Architecture (with Load Balancing):Improved Architecture (with Load Balancing):

CODE
Phase 2: Multiple Servers + Load Balancer Users โ†’ Load Balancer (Nginx) โ†“ โ”œโ†’ Server 1 (1 vCPU, 2 GB RAM) โ”œโ†’ Server 2 (1 vCPU, 2 GB RAM) โ””โ†’ Server 3 (1 vCPU, 2 GB RAM)
๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

- Better performance: 3 servers processing in parallel, response time < 1 second - High availability: If one server fails, others continue serving - Horizontal scaling: Need more capacity? Just add more servers- Better performance: 3 servers processing in parallel, response time < 1 second - High availability: If one server fails, others continue serving - Horizontal scaling: Need more capacity? Just add more servers

1.2 Load Balancing in Everyday Terms1.2 Load Balancing in Everyday Terms

The Bubble Tea Shop CounterThe Bubble Tea Shop Counter

Imagine you run a popular bubble tea shop:Imagine you run a popular bubble tea shop:

The load balancer is the "counter assignment person":The load balancer is the "counter assignment person":

------

2. Overview of Load Balancing2. Overview of Load Balancing

2.1 Layer 4 Load Balancing (L4): Only Looking at the Address2.1 Layer 4 Load Balancing (L4): Only Looking at the Address

Operates at the transport layer (TCP/UDP) โ€” like a delivery driver who only looks at your address (IP + port) without caring about who you are or what you do.Operates at the transport layer (TCP/UDP) โ€” like a delivery driver who only looks at your address (IP + port) without caring about who you are or what you do.

Characteristics:Characteristics:

๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

`` Client request โ†’ L4 Load Balancer โ†’ Backend server โ†“ Only looks at IP + Port โ†“ Fast forwarding (no content inspection) ```` Client request โ†’ L4 Load Balancer โ†’ Backend server โ†“ Only looks at IP + Port โ†“ Fast forwarding (no content inspection) ``

2.2 Layer 7 Load Balancing (L7): Inspecting the Build Artifact Contents2.2 Layer 7 Load Balancing (L7): Inspecting the Build Artifact Contents

Operates at the application layer (HTTP/HTTPS) โ€” like a delivery driver who not only checks the address but also opens the package to inspect its contents before deciding how to deliver.Operates at the application layer (HTTP/HTTPS) โ€” like a delivery driver who not only checks the address but also opens the package to inspect its contents before deciding how to deliver.

Characteristics:Characteristics:

๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

`` Client request โ†’ L7 Load Balancer โ†’ Parses HTTP content โ†“ Inspects URL, Header, Cookie โ†“ Intelligent routing to specific server ```` Client request โ†’ L7 Load Balancer โ†’ Parses HTTP content โ†“ Inspects URL, Header, Cookie โ†“ Intelligent routing to specific server ``

2.3 L4 vs L7 Comparison2.3 L4 vs L7 Comparison

DimensionL4 Load BalancingL7 Load Balancing
OSI LayerTransport Layer (TCP/UDP)Application Layer (HTTP/HTTPS)
Routing BasisIP address + portURL, Header, Cookie, Body
Processing SpeedExtremely fast (kernel-space)Fast (user-space parsing)
Feature RichnessBasic forwardingSSL termination, caching, compression, WAF
Typical ScenariosDatabases, gaming, long connectionsWeb apps, API gateways, microservices
Representative ProductsLVS, AWS NLBNginx, HAProxy, AWS ALB

------

3. Core Problem 1: Approach to preventing "Broken" Servers from Continuing to Serve3. Core Problem 1: Approach to preventing "Broken" Servers from Continuing to Serve

3.1 Health Checks: Don't Let Unhealthy Instances Drag Down the System3.1 Health Checks: Don't Let Unhealthy Instances Drag Down the System

Imagine one of your checkout counters breaks, but the assignment person doesn't know and keeps sending customers there. The queue grows longer, and customers grow angrier.Imagine one of your checkout counters breaks, but the assignment person doesn't know and keeps sending customers there. The queue grows longer, and customers grow angrier.

Health checks are the "sentinels" that prevent this scenario. They periodically "examine" each server, immediately removing any "sick" ones from the pool and bringing them back once they "recover."Health checks are the "sentinels" that prevent this scenario. They periodically "examine" each server, immediately removing any "sick" ones from the pool and bringing them back once they "recover."

3.2 Active Health Checks vs Passive Health Checks3.2 Active Health Checks vs Passive Health Checks

Active Health Check: The load balancer actively "knocks on the door" asking the server, "Are you still there?"Active Health Check: The load balancer actively "knocks on the door" asking the server, "Are you still there?"

Passive Health Check: The load balancer "observes" the response patterns of real business trafficPassive Health Check: The load balancer "observes" the response patterns of real business traffic

๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

| Metric | Healthy Threshold | Unhealthy Threshold | Notes | |:---|:---|:---|:---| | HTTP Status Code | 200-399 | 400+ or timeout | 4xx/5xx are all considered failures | | TCP Connection | Successfully established | Connection timeout | Checks whether the port is reachable | | Response Time | < 500 ms | > 2000 ms | Timeout typically set to 2-5 seconds | | Consecutive Failures | - | 3 times | Avoids false positives from transient blips | | Check Interval | - | 5 s | Too frequent increases load || Metric | Healthy Threshold | Unhealthy Threshold | Notes | |:---|:---|:---|:---| | HTTP Status Code | 200-399 | 400+ or timeout | 4xx/5xx are all considered failures | | TCP Connection | Successfully established | Connection timeout | Checks whether the port is reachable | | Response Time | < 500 ms | > 2000 ms | Timeout typically set to 2-5 seconds | | Consecutive Failures | - | 3 times | Avoids false positives from transient blips | | Check Interval | - | 5 s | Too frequent increases load |

A team set the health check response time threshold to 100ms, but their application's average response time fluctuated between 80-120ms. As a result, servers were frequently marked as "unhealthy," causing traffic to oscillate between healthy and unhealthy states, and overall system availability actually dropped.A team set the health check response time threshold to 100ms, but their application's average response time fluctuated between 80-120ms. As a result, servers were frequently marked as "unhealthy," causing traffic to oscillate between healthy and unhealthy states, and overall system availability actually dropped.

The right approach: Thresholds should be set to 2-3x the P99 response time, leaving enough buffer for normal fluctuations.The right approach: Thresholds should be set to 2-3x the P99 response time, leaving enough buffer for normal fluctuations.

๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

--- ## 4. Core Problem 2: Approach to Ensuring Returning Users Always Get "Backend Instance" ### 4.1 Session Persistence: Let "Returning User" Always Find "Backend Instance" Imagine you're a regular at a bubble tea shop, and the same staff member serves you every time. She knows your preferences (half sugar, no ice) and serves you quickly and thoughtfully. But if you get a new person every time, you have to repeat the same requests over and over โ€” a huge efficiency loss. Session persistence (sticky sessions) solves this problem: ensuring that requests from the same user are always routed to the same backend server. ### 4.2 Comparison of Three Session Persistence Mechanisms | Mechanism | How It Works | Advantages | Disadvantages | Use Cases | | :--------------- | :--------------------------------------------------- | :--------------------------------------- | :------------------------------------------- | :----------------------------- | | Cookie Insert | LB inserts a cookie in the response; subsequent requests carry this cookie | Unaffected by IP changes; persists from the first request | Client must support cookies; may be disabled | Shopping carts, login sessions | | IP Hash | Hashes the client IP and maps it to a specific server | No client-side support needed; stateless | Session lost if IP changes; hard to distribute evenly | Cookie-free environments, WebSocket | | Sticky Session Table | LB maintains a mapping table of sessions to servers | Supports session replication and failover | Consumes LB memory; requires additional synchronization | Scenarios with strict high-availability requirements |--- ## 4. Core Problem 2: Approach to Ensuring Returning Users Always Get "Backend Instance" ### 4.1 Session Persistence: Let "Returning User" Always Find "Backend Instance" Imagine you're a regular at a bubble tea shop, and the same staff member serves you every time. She knows your preferences (half sugar, no ice) and serves you quickly and thoughtfully. But if you get a new person every time, you have to repeat the same requests over and over โ€” a huge efficiency loss. Session persistence (sticky sessions) solves this problem: ensuring that requests from the same user are always routed to the same backend server. ### 4.2 Comparison of Three Session Persistence Mechanisms | Mechanism | How It Works | Advantages | Disadvantages | Use Cases | | :--------------- | :--------------------------------------------------- | :--------------------------------------- | :------------------------------------------- | :----------------------------- | | Cookie Insert | LB inserts a cookie in the response; subsequent requests carry this cookie | Unaffected by IP changes; persists from the first request | Client must support cookies; may be disabled | Shopping carts, login sessions | | IP Hash | Hashes the client IP and maps it to a specific server | No client-side support needed; stateless | Session lost if IP changes; hard to distribute evenly | Cookie-free environments, WebSocket | | Sticky Session Table | LB maintains a mapping table of sessions to servers | Supports session replication and failover | Consumes LB memory; requires additional synchronization | Scenarios with strict high-availability requirements |

๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

--- ## 5. Core Problem 3: Approach to achieving Zero-Downtime Deployment ### 5.1 Blue-Green Deployment: "One-Click Switch" for Zero-Downtime Releases Core idea: Maintain two identical production environments (blue and green) simultaneously, but only one serves live traffic at any given time. Workflow: 1. Initial state: Blue environment runs v1.0 (production), green environment stands by. 2. Deploy new version: Deploy v1.1 to the green environment and run internal smoke tests. 3. Switch traffic: Point the load balancer to the green environment; traffic instantly switches to v1.1. 4. Monitor: Observe the green environment's behavior and confirm no anomalies. 5. Keep old version: Keep the blue environment on v1.0 for a period (e.g., 24 hours) as a safety net for rapid rollback.--- ## 5. Core Problem 3: Approach to achieving Zero-Downtime Deployment ### 5.1 Blue-Green Deployment: "One-Click Switch" for Zero-Downtime Releases Core idea: Maintain two identical production environments (blue and green) simultaneously, but only one serves live traffic at any given time. Workflow: 1. Initial state: Blue environment runs v1.0 (production), green environment stands by. 2. Deploy new version: Deploy v1.1 to the green environment and run internal smoke tests. 3. Switch traffic: Point the load balancer to the green environment; traffic instantly switches to v1.1. 4. Monitor: Observe the green environment's behavior and confirm no anomalies. 5. Keep old version: Keep the blue environment on v1.0 for a period (e.g., 24 hours) as a safety net for rapid rollback.

ProsCons
โœ… Zero downtime; switch completes in millisecondsโŒ High resource cost; requires maintaining two environments
โœ… Fast rollback; switch back immediately if issues are foundโŒ Database schema changes require special compatibility handling
โœ… New environment can be fully tested before taking over trafficโŒ Not suitable for stateful services (e.g., WebSocket long connections)
๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

### 5.2 Canary Release: "Small Steps, Fast Iteration" Canary Strategy The canary release is named after the historical "coal mine canary" โ€” miners brought canaries into the mines; if the canary showed signs of distress, it indicated toxic gas leakage, and miners would evacuate immediately. In software releases, a canary release means exposing a small subset of users to the new version first, observing for issues, and then gradually expanding the rollout. Core idea: 1. Small traffic first: Route 1% of traffic to the new version servers initially. 2. Observe metrics: Continuously monitor error rates, latency, and key business metrics. 3. Gradual rollout: If everything is normal, gradually increase the proportion to 5%, 10%, 25%, 50%, and 100%. 4. Rapid rollback: If any anomaly is detected, immediately switch all traffic back to the old version.### 5.2 Canary Release: "Small Steps, Fast Iteration" Canary Strategy The canary release is named after the historical "coal mine canary" โ€” miners brought canaries into the mines; if the canary showed signs of distress, it indicated toxic gas leakage, and miners would evacuate immediately. In software releases, a canary release means exposing a small subset of users to the new version first, observing for issues, and then gradually expanding the rollout. Core idea: 1. Small traffic first: Route 1% of traffic to the new version servers initially. 2. Observe metrics: Continuously monitor error rates, latency, and key business metrics. 3. Gradual rollout: If everything is normal, gradually increase the proportion to 5%, 10%, 25%, 50%, and 100%. 4. Rapid rollback: If any anomaly is detected, immediately switch all traffic back to the old version.

AdvantageDescription
๐ŸŽฏ Controlled riskEven if the new version has severe bugs, only a small number of users are affected
๐Ÿ“Š Real-world validationValidated in the real production environment, more reliable than staging
Fast iterationTeams can release new features more confidently and frequently
๐Ÿ’ฐ Resource-friendlyDoesn't require two complete environments like blue-green deployment
๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

--- ## 6. Core Problem 4: Approach to making the System "Auto-Scale" on Its Own ### 6.1 Auto Scaling: Let the System Flexibly Schedule Workloads Imagine you run a restaurant: - Lunch peak: You need 10 servers, but at 3 PM during the afternoon lull, you only need 2 - If you always keep 10: labor costs explode - If you always keep only 2: customers during peak hours can't wait and all leave Auto Scaling lets the system "flexibly schedule" like a restaurant โ€” automatically adding servers when busy and removing them when idle. ### 6.2 Choosing Scaling Metrics The core question of auto scaling is: When should you add machines? When should you remove them? Common decision metrics: | Metric | Scale-Up Threshold | Scale-Down Threshold | Use Case | | :---------------------- | :----------------- | :------------------- | :------------------------------ | | CPU Utilization | > 70% | < 30% | Compute-intensive applications | | Memory Utilization | > 75% | < 40% | Memory-intensive applications | | QPS (Queries/sec) | > 1000/s | < 400/s | API gateways, web services | | Connection Count | > 5000 | < 1000 | Databases, message queues | | Custom Business Metrics | Depends on business | Depends on business | Specific business scenarios |--- ## 6. Core Problem 4: Approach to making the System "Auto-Scale" on Its Own ### 6.1 Auto Scaling: Let the System Flexibly Schedule Workloads Imagine you run a restaurant: - Lunch peak: You need 10 servers, but at 3 PM during the afternoon lull, you only need 2 - If you always keep 10: labor costs explode - If you always keep only 2: customers during peak hours can't wait and all leave Auto Scaling lets the system "flexibly schedule" like a restaurant โ€” automatically adding servers when busy and removing them when idle. ### 6.2 Choosing Scaling Metrics The core question of auto scaling is: When should you add machines? When should you remove them? Common decision metrics: | Metric | Scale-Up Threshold | Scale-Down Threshold | Use Case | | :---------------------- | :----------------- | :------------------- | :------------------------------ | | CPU Utilization | > 70% | < 30% | Compute-intensive applications | | Memory Utilization | > 75% | < 40% | Memory-intensive applications | | QPS (Queries/sec) | > 1000/s | < 400/s | API gateways, web services | | Connection Count | > 5000 | < 1000 | Databases, message queues | | Custom Business Metrics | Depends on business | Depends on business | Specific business scenarios |

Pitfall 1: Scaling responds too slowly; the traffic surge already crashes the systemPitfall 1: Scaling responds too slowly; the traffic surge already crashes the system

During a major e-commerce promotion, the team set CPU > 80% as the scale-up trigger, but metric collection had a 1-minute delay, and new instance startup took 3 minutes. Traffic arrived too fast โ€” before scaling completed, the servers were already overwhelmed.During a major e-commerce promotion, the team set CPU > 80% as the scale-up trigger, but metric collection had a 1-minute delay, and new instance startup took 3 minutes. Traffic arrived too fast โ€” before scaling completed, the servers were already overwhelmed.

Solutions:Solutions:

Pitfall 2: Scaling is too aggressive; costs explodePitfall 2: Scaling is too aggressive; costs explode

A startup set an aggressive auto-scaling policy: scale up if CPU > 50%. As a result, a normal business fluctuation triggered scaling, and the server count ballooned from 5 to 30. The end-of-month cloud bill terrified the CTO.A startup set an aggressive auto-scaling policy: scale up if CPU > 50%. As a result, a normal business fluctuation triggered scaling, and the server count ballooned from 5 to 30. The end-of-month cloud bill terrified the CTO.

Solutions:Solutions:

Pitfall 3: Scaling down too fast; newly added machines are removed immediatelyPitfall 3: Scaling down too fast; newly added machines are removed immediately

A team set CPU < 30% as the scale-down trigger. After scaling up, traffic was still settling, and CPU briefly dropped to 25%, triggering a scale-down. Right after scaling down, CPU spiked back to 80%, triggering another scale-up โ€” the system oscillated wildly in a "scale-up, scale-down, scale-up" loop.A team set CPU < 30% as the scale-down trigger. After scaling up, traffic was still settling, and CPU briefly dropped to 25%, triggering a scale-down. Right after scaling down, CPU spiked back to 80%, triggering another scale-up โ€” the system oscillated wildly in a "scale-up, scale-down, scale-up" loop.

Solutions:Solutions:

๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

--- ## 7. Practical Guide: Approach to choosing a Load Balancer ### 7.1 Comparison of Mainstream Load Balancers | Feature | Nginx | HAProxy | Envoy | Cloud Provider LB | | --------------------- | ------------------------------------ | -------------------------- | ------------------- | ------------------ | | Positioning | High-performance reverse proxy / LB | Open-source load balancer | Cloud-native proxy | Managed load balancer | | Performance | Extremely high (C, event-driven) | High (event-driven) | High (C++/Rust) | Extremely high | | Feature Richness | Basic LB, static files, caching | Rich LB algorithms | Advanced routing, observability | Full-featured | | Configuration | Config file (nginx.conf) | Config file (haproxy.cfg) | API / config file | UI console | | Extensibility | C modules / Lua scripts | Lua scripts | WASM / Filters | Plugins | | Use Cases | Static assets, L7 LB, SSL termination | L7 LB, high availability | Service mesh, multi-cloud | Quick start |--- ## 7. Practical Guide: Approach to choosing a Load Balancer ### 7.1 Comparison of Mainstream Load Balancers | Feature | Nginx | HAProxy | Envoy | Cloud Provider LB | | --------------------- | ------------------------------------ | -------------------------- | ------------------- | ------------------ | | Positioning | High-performance reverse proxy / LB | Open-source load balancer | Cloud-native proxy | Managed load balancer | | Performance | Extremely high (C, event-driven) | High (event-driven) | High (C++/Rust) | Extremely high | | Feature Richness | Basic LB, static files, caching | Rich LB algorithms | Advanced routing, observability | Full-featured | | Configuration | Config file (nginx.conf) | Config file (haproxy.cfg) | API / config file | UI console | | Extensibility | C modules / Lua scripts | Lua scripts | WASM / Filters | Plugins | | Use Cases | Static assets, L7 LB, SSL termination | L7 LB, high availability | Service mesh, multi-cloud | Quick start |

Decision Tree:Decision Tree:

CODE
Choose a load balancer: โ”‚ โ”œโ”€ Only need basic L4 load balancing? โ”‚ โ”œโ”€ Yes โ†’ LVS (open-source, free) or cloud provider NLB โ”‚ โ””โ”€ No โ†’ Continue โ”‚ โ”œโ”€ Need service mesh or multi-cloud deployment? โ”‚ โ”œโ”€ Yes โ†’ Envoy โ”‚ โ””โ”€ No โ†’ Continue โ”‚ โ”œโ”€ Need extremely complex configuration and plugins? โ”‚ โ”œโ”€ Yes โ†’ HAProxy โ”‚ โ””โ”€ No โ†’ Continue โ”‚ โ”œโ”€ Need high performance + simple configuration? โ”‚ โ”œโ”€ Yes โ†’ Nginx (first choice) โ”‚ โ””โ”€ Continue โ”‚ โ”œโ”€ Want managed operations? โ”‚ โ”œโ”€ Yes โ†’ Cloud provider LB (AWS ALB, Alibaba Cloud SLB) โ”‚ โ””โ”€ Self-host Nginx
๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

--- ## 8. Summary: Core Mindset of Load Balancing ### 8.1 Core Principles Recap | Principle | Meaning | Key Practice Points | | --------------- | ---------------------------------------------- | ------------------------------------------------------------ | | Layering | L4 handles "package sorting" (fast but simple) | L4 for databases, gaming; L7 for web, APIs | | Redundancy | Single points of failure are the enemy | Improve availability through multi-instance, multi-region deployment | | Gradualism | Don't release new versions with "one big cut" | Blue-green deployment for zero downtime; canary for controlled risk | | Elasticity | The system should "breathe" like a living organism | Automatically scale up when busy, scale down when idle | ### 8.2 Design Checklist Before introducing load balancing, ask yourself the following questions: - [ ] Is load balancing really needed? (Is single-machine performance truly insufficient?) - [ ] Choose L4 or L7? (Based on business scenario) - [ ] How to handle session persistence? (Cookie, IP hash, session table) - [ ] How to implement health checks? (Active, passive, threshold settings) - [ ] How to achieve zero downtime? (Blue-green deployment, canary) - [ ] How to implement elasticity? (Scaling metrics, cooldown periods, max instance count) --- ## 9. Glossary | Term | Chinese Translation | Explanation | | ----------------------------- | ------------------- | ------------------------------------------------------------------------------------- | | Load Balancer | ่ดŸ่ฝฝๅ‡่กกๅ™จ | A device or software that distributes traffic across multiple backend servers | | L4 Load Balancing | ๅ››ๅฑ‚่ดŸ่ฝฝๅ‡่กก | Load balancing based on the transport layer (TCP/UDP) | | L7 Load Balancing | ไธƒๅฑ‚่ดŸ่ฝฝๅ‡่กก | Load balancing based on the application layer (HTTP/HTTPS) | | Health Check | ๅฅๅบทๆฃ€ๆŸฅ | A mechanism that periodically checks the health status of backend servers | | Session Persistence | ไผš่ฏไฟๆŒ | Ensures requests from the same user are always routed to the same server | | Sticky Session | ็ฒ˜ๆ€งไผš่ฏ | Another term for Session Persistence | | Blue-Green Deployment | ่“็ปฟ้ƒจ็ฝฒ | A zero-downtime release strategy using two environments that switch over | | Canary Release | ้‡‘ไธ้›€ๅ‘ๅธƒ | A canary release strategy that validates with a small traffic portion first | | Auto Scaling | ่‡ชๅŠจๆ‰ฉ็ผฉๅฎน | Automatically increasing or decreasing the number of servers based on load | | Horizontal Scaling | ๆฐดๅนณๆ‰ฉๅฑ• | Increasing server count to improve processing capacity | | Vertical Scaling | ๅž‚็›ดๆ‰ฉๅฑ• | Upgrading individual machine specs (CPU, RAM) to improve processing capacity | | Multi-Region | ๅคšๅŒบๅŸŸ | Deploying services across multiple geographic regions | | Active-Active | ๅคšๆดป | Multiple regions simultaneously serving traffic | | Active-Standby | ไธปๅค‡ | Only one region serves traffic; others are on standby | | Data Replication | ๆ•ฐๆฎๅŒๆญฅ | Cross-region data replication mechanism | | RTO | ๆขๅคๆ—ถ้—ด็›ฎๆ ‡ | Recovery Time Objective โ€” the target time within which a system must recover after failure | | RPO | ๆขๅค็‚น็›ฎๆ ‡ | Recovery Point Objective โ€” the acceptable amount of data loss after a system failure |--- ## 8. Summary: Core Mindset of Load Balancing ### 8.1 Core Principles Recap | Principle | Meaning | Key Practice Points | | --------------- | ---------------------------------------------- | ------------------------------------------------------------ | | Layering | L4 handles "package sorting" (fast but simple) | L4 for databases, gaming; L7 for web, APIs | | Redundancy | Single points of failure are the enemy | Improve availability through multi-instance, multi-region deployment | | Gradualism | Don't release new versions with "one big cut" | Blue-green deployment for zero downtime; canary for controlled risk | | Elasticity | The system should "breathe" like a living organism | Automatically scale up when busy, scale down when idle | ### 8.2 Design Checklist Before introducing load balancing, ask yourself the following questions: - [ ] Is load balancing really needed? (Is single-machine performance truly insufficient?) - [ ] Choose L4 or L7? (Based on business scenario) - [ ] How to handle session persistence? (Cookie, IP hash, session table) - [ ] How to implement health checks? (Active, passive, threshold settings) - [ ] How to achieve zero downtime? (Blue-green deployment, canary) - [ ] How to implement elasticity? (Scaling metrics, cooldown periods, max instance count) --- ## 9. Glossary | Term | Chinese Translation | Explanation | | ----------------------------- | ------------------- | ------------------------------------------------------------------------------------- | | Load Balancer | ่ดŸ่ฝฝๅ‡่กกๅ™จ | A device or software that distributes traffic across multiple backend servers | | L4 Load Balancing | ๅ››ๅฑ‚่ดŸ่ฝฝๅ‡่กก | Load balancing based on the transport layer (TCP/UDP) | | L7 Load Balancing | ไธƒๅฑ‚่ดŸ่ฝฝๅ‡่กก | Load balancing based on the application layer (HTTP/HTTPS) | | Health Check | ๅฅๅบทๆฃ€ๆŸฅ | A mechanism that periodically checks the health status of backend servers | | Session Persistence | ไผš่ฏไฟๆŒ | Ensures requests from the same user are always routed to the same server | | Sticky Session | ็ฒ˜ๆ€งไผš่ฏ | Another term for Session Persistence | | Blue-Green Deployment | ่“็ปฟ้ƒจ็ฝฒ | A zero-downtime release strategy using two environments that switch over | | Canary Release | ้‡‘ไธ้›€ๅ‘ๅธƒ | A canary release strategy that validates with a small traffic portion first | | Auto Scaling | ่‡ชๅŠจๆ‰ฉ็ผฉๅฎน | Automatically increasing or decreasing the number of servers based on load | | Horizontal Scaling | ๆฐดๅนณๆ‰ฉๅฑ• | Increasing server count to improve processing capacity | | Vertical Scaling | ๅž‚็›ดๆ‰ฉๅฑ• | Upgrading individual machine specs (CPU, RAM) to improve processing capacity | | Multi-Region | ๅคšๅŒบๅŸŸ | Deploying services across multiple geographic regions | | Active-Active | ๅคšๆดป | Multiple regions simultaneously serving traffic | | Active-Standby | ไธปๅค‡ | Only one region serves traffic; others are on standby | | Data Replication | ๆ•ฐๆฎๅŒๆญฅ | Cross-region data replication mechanism | | RTO | ๆขๅคๆ—ถ้—ด็›ฎๆ ‡ | Recovery Time Objective โ€” the target time within which a system must recover after failure | | RPO | ๆขๅค็‚น็›ฎๆ ‡ | Recovery Point Objective โ€” the acceptable amount of data loss after a system failure |