VibeKoding / Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat / Principles of Monitoring, Logging, and AlertingPrinciples of Monitoring, Logging, and Alerting
VK

Principles of Monitoring, Logging, and AlertingPrinciples of Monitoring, Logging, and Alerting

๐Ÿ“š Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat ๐ŸŒ Dual Bahasa (ID / EN) โšก VibeKoding Native

Ensiklopedia VibeKoding: Principles of Monitoring, Logging, and Alerting.Ensiklopedia VibeKoding: Principles of Monitoring, Logging, and Alerting.

> ๐Ÿ’ก Learning Guide: This chapter requires no programming background. Through interactive demos, you'll gain a comprehensive understanding of operations โ€” from monitoring and alerting to troubleshooting, from capacity planning to automated operations, mastering all the skills needed to run production systems.> ๐Ÿ’ก Learning Guide: This chapter requires no programming background. Through interactive demos, you'll gain a comprehensive understanding of operations โ€” from monitoring and alerting to troubleshooting, from capacity planning to automated operations, mastering all the skills needed to run production systems.

0. Introduction: Deployment Is Just the Beginning0. Introduction: Deployment Is Just the Beginning

Many beginners think: "Once the code is deployed, the job is done."Many beginners think: "Once the code is deployed, the job is done."

That couldn't be more wrong!That couldn't be more wrong!

Deployment is merely the starting point of operations work. It's like buying a new car โ€” the real work of maintenance, repairs, and refueling is what follows.Deployment is merely the starting point of operations work. It's like buying a new car โ€” the real work of maintenance, repairs, and refueling is what follows.

Operations has three goals:Operations has three goals:

  1. Stability: The system stays up and services remain availableStability: The system stays up and services remain available
  2. Performance: Fast responses and a great user experiencePerformance: Fast responses and a great user experience
  3. Security: No data leaks and protection against attacksSecurity: No data leaks and protection against attacks
  4. ------

    1. Monitoring1. Monitoring

    Monitoring is the "eyes" of operations. A system without monitoring is like driving blind โ€” you won't even know when something goes wrong.Monitoring is the "eyes" of operations. A system without monitoring is like driving blind โ€” you won't even know when something goes wrong.

    1.1 The Three Layers of Monitoring1.1 The Three Layers of Monitoring

    Infrastructure Monitoring: Tracking server hardware resourcesInfrastructure Monitoring: Tracking server hardware resources

    • CPU usageCPU usage
    • Memory usageMemory usage
    • Disk space and I/ODisk space and I/O
    • Network bandwidthNetwork bandwidth

    Application Monitoring: Tracking software runtime stateApplication Monitoring: Tracking software runtime state

    • QPS (Queries Per Second)QPS (Queries Per Second)
    • Response time (latency)Response time (latency)
    • Error rateError rate
    • Dependency service call statusDependency service call status

    Business Monitoring: Tracking business healthBusiness Monitoring: Tracking business health

    • DAU/MAU (Daily/Monthly Active Users)DAU/MAU (Daily/Monthly Active Users)
    • Order volumeOrder volume
    • Payment success ratePayment success rate
    • User retention rateUser retention rate

    1.2 Monitoring Tool Stack1.2 Monitoring Tool Stack

    ToolPurposeCharacteristics
    PrometheusMetric collection & storageTime-series database, ideal for monitoring data
    GrafanaVisualization dashboardsPowerful charts and dashboards
    ZabbixComprehensive monitoringVeteran tool with full-featured capabilities
    DatadogSaaS monitoring platformOne-stop solution, paid

    Key Point: Monitoring must be layered, covering everything from infrastructure to business to avoid blind spots.Key Point: Monitoring must be layered, covering everything from infrastructure to business to avoid blind spots.

    ------

    2. Alerting2. Alerting

    Once monitoring detects an issue, operations staff need to be notified promptly โ€” that's alerting.Once monitoring detects an issue, operations staff need to be notified promptly โ€” that's alerting.

    2.1 Alerting Flow2.1 Alerting Flow

    2.2 Alert Severity Levels2.2 Alert Severity Levels

    Proper alert classification helps prevent "alert fatigue":Proper alert classification helps prevent "alert fatigue":

    LevelResponse TimeTypical ScenarioNotification Channels
    P0Immediate (within 5 min)Core service down, payment failuresPhone + SMS + IM
    P1Within 30 minutesPartial feature outage, severe performance degradationSMS + IM + Email
    P2Same dayHigh resource usage, occasional errorsIM + Email
    P3Within the weekNon-critical issues, optimization suggestionsEmail

    2.3 Alert Deduplication & Noise Reduction2.3 Alert Deduplication & Noise Reduction

    Pain Point: A single small issue can trigger hundreds or thousands of alerts, numbing on-call staff.Pain Point: A single small issue can trigger hundreds or thousands of alerts, numbing on-call staff.

    Solutions:Solutions:

    1. Alert Grouping: Merge similar alerts (e.g., multiple issues on the same server combined into one)Alert Grouping: Merge similar alerts (e.g., multiple issues on the same server combined into one)
    2. Alert Suppression: If a parent issue has already fired, don't repeat alerts for child issuesAlert Suppression: If a parent issue has already fired, don't repeat alerts for child issues
    3. Silence Rules: Automatically suppress alerts during maintenance windowsSilence Rules: Automatically suppress alerts during maintenance windows
    4. Rate Limiting: Don't repeat the same alert notification within a short time windowRate Limiting: Don't repeat the same alert notification within a short time window
    5. Key Point: Alerts should be "few but meaningful" โ€” every alert must be worth acting on.Key Point: Alerts should be "few but meaningful" โ€” every alert must be worth acting on.

      ------

      3. Logging3. Logging

      Logs are the "black box" for troubleshooting.Logs are the "black box" for troubleshooting.

      3.1 Log Levels3.1 Log Levels

      javascript
      console.debug('Verbose debug info') // Used during development console.info('General information') // Normal flow logging console.warn('Warning') // Potential issues console.error('Error') // Errors that need attention
      

      3.2 Structured Logging3.2 Structured Logging

      Traditional logging (not ideal):Traditional logging (not ideal):

      CODE
      2024-01-15 10:23:45 ERROR User john failed to login, attempts=3, ip=192.168.1.100
      

      Structured logging (recommended):Structured logging (recommended):

      json
      { "timestamp": "2024-01-15T10:23:45Z", "level": "ERROR", "message": "User login failed", "user": "john", "attempts": 3, "ip": "192.168.1.100", "service": "auth-service" }
      

      3.3 The ELK Stack3.3 The ELK Stack

      ELK = Elasticsearch + Logstash + KibanaELK = Elasticsearch + Logstash + Kibana

      • Logstash: Log collection and filteringLogstash: Log collection and filtering
      • Elasticsearch: Log storage and searchElasticsearch: Log storage and search
      • Kibana: Log visualization and queryingKibana: Log visualization and querying

      Best Practices:Best Practices:

      • โœ… Don't log sensitive information (passwords, tokens)โœ… Don't log sensitive information (passwords, tokens)
      • โœ… Critical operations (login, payment, permission changes) must be loggedโœ… Critical operations (login, payment, permission changes) must be logged
      • โœ… Logs should include context (user ID, request ID, timestamp)โœ… Logs should include context (user ID, request ID, timestamp)
      • โœ… Regularly purge expired logs to avoid running out of disk spaceโœ… Regularly purge expired logs to avoid running out of disk space

      ------

      4. Distributed Tracing4. Distributed Tracing

      In a microservices architecture, a single request may pass through dozens of services โ€” how do you trace its complete path?In a microservices architecture, a single request may pass through dozens of services โ€” how do you trace its complete path?

      Trace ID and Span IDTrace ID and Span ID

      • Trace ID: The unique identifier for an entire request chain (like a package tracking number)Trace ID: The unique identifier for an entire request chain (like a package tracking number)
      • Span ID: The identifier for a single service call (like each transfer hub)Span ID: The identifier for a single service call (like each transfer hub)

      4.1 Distributed Tracing Demo4.1 Distributed Tracing Demo

      4.2 The OpenTelemetry Standard4.2 The OpenTelemetry Standard

      OpenTelemetry (OTel) is the industry standard for distributed tracing, providing a unified API and SDK.OpenTelemetry (OTel) is the industry standard for distributed tracing, providing a unified API and SDK.

      javascript
      // Example: Recording a Span with OpenTelemetry import { trace } from '@opentelemetry/api' const tracer = trace.getTracer('my-service') async function processOrder(orderId) { // Create a Span const span = tracer.startSpan('processOrder') try { // Set attributes span.setAttribute('order.id', orderId) // Business logic... await validateOrder(orderId) await saveToDatabase(orderId) span.setStatus({ code: SpanStatusCode.OK }) } catch (error) { span.recordException(error) span.setStatus({ code: SpanStatusCode.ERROR, message: error.message }) } finally { span.end() // End the Span } }
      

      Key Point: Distributed tracing quickly identifies performance bottlenecks and failure points โ€” an essential tool for microservices.Key Point: Distributed tracing quickly identifies performance bottlenecks and failure points โ€” an essential tool for microservices.

      ------

      5. Troubleshooting Process5. Troubleshooting Process

      Production incidents are inevitable. The key is fast response and fast recovery.Production incidents are inevitable. The key is fast response and fast recovery.

      5.1 Incident Response Process5.1 Incident Response Process

      5.2 Common Troubleshooting Tools5.2 Common Troubleshooting Tools

      ToolPurposeTypical Scenario
      tcpdumpPacket capture analysisNetwork issues, packet loss
      straceSystem call tracingProcess hanging, file permission issues
      ArthasJava diagnosticsCPU spikes, memory leaks, deadlocks
      top/htopSystem resource monitoringHigh CPU/memory usage
      netstatNetwork connection inspectionPort conflicts, abnormal connection counts
      lsofOpen file inspectionFile locks, disk full

      Arthas Example (Alibaba's open-source Java diagnostic tool):Arthas Example (Alibaba's open-source Java diagnostic tool):

      bash
      # View top 5 threads by CPU usage $ top -H -p 12345 # Trace the execution time of a method $ trace com.example.OrderService createOrder # View a class's static fields $ getstatic com.example.Config MAX_CONNECTIONS # Hot-reload code (no restart needed) $ mc /tmp/Test.java $ redefine /tmp/Test.class
      

      5.3 Post-mortem Analysis5.3 Post-mortem Analysis

      A post-mortem is not a blame session!A post-mortem is not a blame session!

      The purpose of a post-mortem is:The purpose of a post-mortem is:

      1. Reconstruct the incident timelineReconstruct the incident timeline
      2. Identify the root cause (Root Cause Analysis)Identify the root cause (Root Cause Analysis)
      3. Summarize lessons learnedSummarize lessons learned
      4. Define improvement actionsDefine improvement actions
      5. The 5 Whys Analysis:The 5 Whys Analysis:

        Ask "why" at least 5 times to find the root cause:Ask "why" at least 5 times to find the root cause:

        • Why did the service go down?Why did the service go down?
        • Because of an out-of-memory errorBecause of an out-of-memory error
        • Why did memory overflow?Why did memory overflow?
        • Because cached data grew too largeBecause cached data grew too large
        • Why was cached data too large?Why was cached data too large?
        • Because no expiration time was setBecause no expiration time was set
        • Why was no expiration time set?Why was no expiration time set?
        • Because it was overlooked during developmentBecause it was overlooked during development
        • Root cause: Lack of code review and test coverageRoot cause: Lack of code review and test coverage

        Key Point: Build a blameless culture โ€” focus on process improvement, not individual accountability.Key Point: Build a blameless culture โ€” focus on process improvement, not individual accountability.

        ------

        6. Performance Optimization6. Performance Optimization

        6.1 Performance Bottleneck Analysis6.1 Performance Bottleneck Analysis

        Top-down optimization approach:Top-down optimization approach:

        CODE
        User Experience โ†“ Frontend Optimization (reduce requests, CDN, lazy loading) โ†“ Network Optimization (HTTP/2, compression, persistent connections) โ†“ Backend Optimization (caching, async, batching) โ†“ Database Optimization (indexes, query tuning, sharding) โ†“ System Optimization (kernel parameters, JVM tuning)
        

        6.2 Database Optimization6.2 Database Optimization

        Index Optimization:Index Optimization:

        sql
        -- Slow query (no index) SELECT * FROM orders WHERE user_id = 12345; -- 100x faster after creating an index CREATE INDEX idx_user_id ON orders(user_id);
        

        Query Optimization:Query Optimization:

        sql
        -- โŒ Avoid SELECT * SELECT * FROM users WHERE id = 123; -- โœ… Only query needed fields SELECT id, name, email FROM users WHERE id = 123; -- โŒ Avoid overly large IN clauses SELECT * FROM orders WHERE user_id IN (1, 2, 3, ..., 10000); -- โœ… Use JOIN or batch queries SELECT * FROM orders o JOIN user_ids u ON o.user_id = u.id;
        

        6.3 Cache Optimization6.3 Cache Optimization

        Multi-level Cache Architecture:Multi-level Cache Architecture:

        CODE
        Browser Cache (CDN) โ†“ Local Cache (In-memory/Guava) โ†“ Distributed Cache (Redis/Memcached) โ†“ Database (MySQL/PostgreSQL)
        

        Cache Update Strategies:Cache Update Strategies:

        StrategyProsConsUse Case
        Cache-AsideSimple, reliableSlow on first queryRead-heavy, write-light
        Write-ThroughGood data consistencySlow writesBalanced read/write
        Write-BehindExtremely fast writesPotential data lossWrite-heavy, tolerates brief inconsistency

        Key Point: Caching is not a silver bullet โ€” consider consistency, avalanche, and penetration issues (refer to the "System Cache Design" chapter).Key Point: Caching is not a silver bullet โ€” consider consistency, avalanche, and penetration issues (refer to the "System Cache Design" chapter).

        ------

        7. Capacity Planning7. Capacity Planning

        7.1 Capacity Assessment7.1 Capacity Assessment

        7.2 Stress Testing7.2 Stress Testing

        Tool Selection:Tool Selection:

        ToolCharacteristicsUse Case
        JMeterFeature-rich, visualHTTP API stress testing
        wrk/abLightweight, command-lineQuick benchmarking
        LocustPython scripting, distributedComplex scenario testing
        K6Modern, JS scriptingCI/CD integration

        wrk Example:wrk Example:

        bash
        # Install wrk $ brew install wrk # macOS $ apt install wrk # Ubuntu # Stress test an HTTP endpoint (10 threads, 30 seconds) $ wrk -t10 -c100 -d30s http://example.com/api/users # Output: # Running 30s test @ http://example.com/api/users # 10 threads and 100 connections # Thread Stats Avg Stdev Max +/- Stdev # Latency 45.32ms 12.45ms 120.50ms 87.56% # Req/Sec 2.12k 123.45 3.45k 89.01% # 632450 requests in 30.00s, 1.23GB read # Requests/sec: 21081.67 

        7.3 Elastic Scaling7.3 Elastic Scaling

        Auto-scaling in the cloud-native era:Auto-scaling in the cloud-native era:

        yaml
        # Kubernetes HPA (Horizontal Pod Autoscaler) apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: my-app-hpa spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: my-app minReplicas: 2 maxReplicas: 10 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 70
        

        When CPU usage exceeds 70%, pods automatically scale up (up to 10)When CPU usage exceeds 70%, pods automatically scale up (up to 10)

        Key Point: Combine business forecasting (e.g., Black Friday sales) with proactive scaling to avoid last-minute scrambling.Key Point: Combine business forecasting (e.g., Black Friday sales) with proactive scaling to avoid last-minute scrambling.

        ------

        8. Security Operations8. Security Operations

        8.1 Access Control8.1 Access Control

        Principle of Least Privilege:Principle of Least Privilege:

        • Developers can only access the development environmentDevelopers can only access the development environment
        • Operations staff can only access production, and require approvalOperations staff can only access production, and require approval
        • Sensitive database operations require secondary confirmationSensitive database operations require secondary confirmation

        Jump Server (Bastion Host):Jump Server (Bastion Host):

        All operations tasks go through the bastion host, which records complete operation logs.All operations tasks go through the bastion host, which records complete operation logs.

        8.2 Data Backup8.2 Data Backup

        The 3-2-1 Backup Rule:The 3-2-1 Backup Rule:

        • 3 copies of data (1 original + 2 backups)3 copies of data (1 original + 2 backups)
        • 2 different storage media (local disk + cloud storage)2 different storage media (local disk + cloud storage)
        • 1 offsite backup (to prevent single-point disasters)1 offsite backup (to prevent single-point disasters)

        Backup Strategies:Backup Strategies:

        TypeFrequencyRetentionRTORPO
        Full BackupWeekly1 month4 hours24 hours
        Incremental BackupDaily1 week2 hours1 hour
        Real-time BackupPer second7 daysMinutesSeconds

        RTO (Recovery Time Objective): The maximum acceptable downtime durationRTO (Recovery Time Objective): The maximum acceptable downtime duration

        RPO (Recovery Point Objective): The maximum acceptable data lossRPO (Recovery Point Objective): The maximum acceptable data loss

        8.3 Vulnerability Scanning8.3 Vulnerability Scanning

        Regular Scanning:Regular Scanning:

        • Code Scanning: SonarQube, ESLint (detect potential vulnerabilities)Code Scanning: SonarQube, ESLint (detect potential vulnerabilities)
        • Dependency Scanning: npm audit, Snyk (detect third-party library vulnerabilities)Dependency Scanning: npm audit, Snyk (detect third-party library vulnerabilities)
        • Container Scanning: Trivy, Clair (detect image vulnerabilities)Container Scanning: Trivy, Clair (detect image vulnerabilities)
        bash
        # npm audit example $ npm audit found 3 vulnerabilities (1 moderate, 2 high) Package Severity Vulnerable versions lodash high <4.17.21 express moderate 4.0.0 - 4.18.2 # Auto-fix $ npm audit fix
        

        ------

        9. Automated Operations (DevOps)9. Automated Operations (DevOps)

        9.1 CI/CD Pipeline9.1 CI/CD Pipeline

        yaml
        # .gitlab-ci.yml example stages: - test - build - deploy test: stage: test script: - npm install - npm test tags: - docker build: stage: build script: - docker build -t myapp:$CI_COMMIT_SHA . - docker push registry.example.com/myapp:$CI_COMMIT_SHA only: - main deploy: stage: deploy script: - kubectl set image deployment/myapp myapp=registry.example.com/myapp:$CI_COMMIT_SHA environment: name: production when: manual # Manually triggered deployment
        

        9.2 Infrastructure as Code (IaC)9.2 Infrastructure as Code (IaC)

        Terraform Example (managing cloud resources):Terraform Example (managing cloud resources):

        hcl
        # main.tf resource "aws_instance" "web" { ami = "ami-0c55b159cbfafe1f0" instance_type = "t2.micro" tags = { Name = "WebServer" Env = "production" } } resource "aws_security_group" "web" { name = "web-sg" ingress { from_port = 80 to_port = 80 protocol = "tcp" cidr_blocks = ["0.0.0.0/0"] } }
        

        Advantages:Advantages:

        • โœ… Version Control: All configuration lives in Gitโœ… Version Control: All configuration lives in Git
        • โœ… Reproducibility: Environment consistencyโœ… Reproducibility: Environment consistency
        • โœ… Auditability: Clear change historyโœ… Auditability: Clear change history
        • โœ… Rollback: Quickly revert to previous versionsโœ… Rollback: Quickly revert to previous versions

        9.3 GitOps Practices9.3 GitOps Practices

        GitOps = Git + IaC + AutomationGitOps = Git + IaC + Automation

        Core principle: The Git repository is the single source of truth for infrastructureCore principle: The Git repository is the single source of truth for infrastructure

        Workflow:Workflow:

        CODE
        1. Modify config files (push to Git) โ†“ 2. Git repository changes trigger CI/CD โ†“ 3. Automatically run terraform apply / kubectl apply โ†“ 4. Infrastructure updates automatically โ†“ 5. Monitor and reconcile actual state vs. desired state
        

        Tools: ArgoCD, Flux (Kubernetes deployment)Tools: ArgoCD, Flux (Kubernetes deployment)

        ------

        10. Summary & Best Practices10. Summary & Best Practices

        Operations is a vast domain, but the core can be distilled into the following:Operations is a vast domain, but the core can be distilled into the following:

        10.1 Operations Maturity Model10.1 Operations Maturity Model

        LevelCharacteristicsPractices
        BeginnerReactive, manual operationsFix issues only when they arise, manual deploys
        IntermediateAutomated, standardizedCI/CD, monitoring & alerting, documentation
        AdvancedProactive, self-healingCapacity planning, chaos drills, auto-scaling
        ExpertIntelligent, unattendedAIOps, chaos engineering, serverless

        10.2 A Day in the Life of an SRE10.2 A Day in the Life of an SRE

        CODE
        09:00 - Review overnight alerts, confirm system status 10:00 - Handle user-reported issues 11:00 - Attend engineering weekly, assess operational risk of new proposals 14:00 - Optimize slow queries, improve performance 15:00 - Code review 16:00 - Write deployment docs, update monitoring rules 17:00 - Chaos engineering drills 18:00 - On-call handoff
        

        10.3 Learning Roadmap10.3 Learning Roadmap

        Beginner Stage (1โ€“3 months):Beginner Stage (1โ€“3 months):

        • Learn common Linux commandsLearn common Linux commands
        • Understand monitoring systems (Prometheus + Grafana)Understand monitoring systems (Prometheus + Grafana)
        • Master log querying (ELK)Master log querying (ELK)

        Intermediate Stage (3โ€“6 months):Intermediate Stage (3โ€“6 months):

        • Deep dive into container technology (Docker + K8s)Deep dive into container technology (Docker + K8s)
        • Master a diagnostic tool (Arthas, tcpdump)Master a diagnostic tool (Arthas, tcpdump)
        • Practice CI/CD pipelinesPractice CI/CD pipelines

        Advanced Stage (6โ€“12 months):Advanced Stage (6โ€“12 months):

        • Performance tuning (database, JVM, network)Performance tuning (database, JVM, network)
        • Capacity planning and cost optimizationCapacity planning and cost optimization
        • Post-mortems and process improvementPost-mortems and process improvement

        Expert Stage (1+ year):Expert Stage (1+ year):

        • Architecture design (high availability, disaster recovery)Architecture design (high availability, disaster recovery)
        • Chaos engineering (proactively inject failures)Chaos engineering (proactively inject failures)
        • AIOps (intelligent operations)AIOps (intelligent operations)

        ------

        11. Glossary11. Glossary

        TermFull NameExplanation
        Monitoring-Real-time observation of system health.
        Alerting-Notifying relevant personnel when anomalies occur.
        Logging-Recording events during system operation.
        Tracing-Tracking the full path of a request across a distributed system.
        QPSQueries Per SecondQueries per second, a measure of system throughput.
        Latency-The time from request initiation to response.
        RTORecovery Time ObjectiveMaximum acceptable downtime duration.
        RPORecovery Point ObjectiveMaximum acceptable data loss.
        Post-mortem-Incident review to analyze root causes and improvement actions.
        CI/CDContinuous Integration/DeliveryAutomated testing and deployment.
        IaCInfrastructure as CodeManaging servers, networks, and other resources via code.
        GitOps-Git-driven operations โ€” Git is the single source of truth.
        ELKElasticsearch + Logstash + KibanaThe log collection, storage, and visualization trifecta.
        SLAService Level AgreementCommitted service availability (e.g., 99.9%).
        Blameless-A no-blame culture where post-mortems focus on process over individuals.

        ------

        12. Further Reading12. Further Reading

        • [System Cache Design](/en/appendix/4-server-and-backend/caching) - Caching principles, patterns & best practices[System Cache Design](/en/appendix/4-server-and-backend/caching) - Caching principles, patterns & best practices
        • [Message Queue Design](/en/appendix/4-server-and-backend/message-queues) - Peak shaving, async decoupling[Message Queue Design](/en/appendix/4-server-and-backend/message-queues) - Peak shaving, async decoupling
        • [Authentication & Authorization in Practice](/en/appendix/4-server-and-backend/auth-authorization) - AuthN/AuthZ and security hardening[Authentication & Authorization in Practice](/en/appendix/4-server-and-backend/auth-authorization) - AuthN/AuthZ and security hardening
        • [Backend Evolution](/en/appendix/4-server-and-backend/backend-layered-architecture) - From monoliths to microservices to serverless[Backend Evolution](/en/appendix/4-server-and-backend/backend-layered-architecture) - From monoliths to microservices to serverless
        • [Deployment & Go-Live](/en/appendix/7-infrastructure-and-operations/ci-cd) - The last mile from development to production[Deployment & Go-Live](/en/appendix/7-infrastructure-and-operations/ci-cd) - The last mile from development to production