본문으로 건너뛰기
김신건의 로그

[AWS CAF] Operations Perspective

· 수정 · 📖 약 3분 · 858자/단어 #aws #cloud #caf #operations #sre #observability
AWS CAF Operations, CAF Operations Perspective, CAF 운영 관점, Cloud Operations, AWS SRE, AWS Observability

정의

Operations PerspectiveAWS CAF 6 관점 중 클라우드 서비스가 비즈니스가 요구하는 수준으로 딜리버리되도록 보장 하는 관점입니다. 자동화 + 최적화 로 신뢰성 있게 확장.

주요 stakeholder: 인프라/운영 리더, SRE (Site Reliability Engineer), IT service manager.

주요 Capabilities

1. Observability

시스템 상태 가시성. Metrics, Logs, Traces.

AWS 도구:

  • CloudWatch Metrics, Logs, Alarms, Dashboards
  • CloudWatch Application Signals (2024+)
  • X-Ray (tracing)
  • OpenTelemetry (표준)
  • Managed Prometheus, Managed Grafana

3 pillars of observability:

  • Metrics: 수치 시계열 (CPU%, requests/s)
  • Logs: 이벤트 (structured JSON)
  • Traces: 요청 흐름 (분산 트레이싱)

2. Event Management (AIOps)

이벤트 상관 + 자동 대응.

  • Alert de-duplication
  • Alert routing (PagerDuty, Opsgenie)
  • Auto-remediation (Lambda + Systems Manager)
  • ML 기반 이상 감지

3. Incident and Problem Management

사고 대응 프로세스.

  • Incident: 실시간 문제 (down)
  • Problem: 근본 원인 조사
  • PIR (Post-Incident Review): Blameless postmortem

단계:

  1. Detection
  2. Triage (severity 판단)
  3. Response (mitigation)
  4. Communication (상태 페이지, stakeholder)
  5. Resolution
  6. PIR

4. Change and Release Management

프로덕션 변경 관리.

  • Change Advisory Board (전통) -> Automated (클라우드)
  • Blue/green, canary
  • Rollback 자동화
  • Change failure rate 추적 (DORA metric)

5. Performance and Capacity Management

성능 목표 (SLO) 정의 + 용량 계획.

  • SLO (Service Level Objective): 예 99.9% 가용성
  • SLI (Service Level Indicator): 측정 지표
  • Error budget: SLO 대비 여유
  • Capacity planning: 미래 부하 예측

6. Configuration Management

리소스 설정 관리:

  • Config 로 리소스 상태 추적
  • Systems Manager Parameter Store
  • IaC (immutable, versioned)

7. Patch Management

OS/앱 패치:

  • Systems Manager Patch Manager
  • 자동 스케줄 (매주 등)
  • Compliance 상태 리포트

8. Availability and Continuity Management

가용성 + 재해복구.

  • RTO (Recovery Time Objective): 복구까지 시간
  • RPO (Recovery Point Objective): 데이터 손실 허용 시간
  • Multi-AZ, Multi-region
  • DR runbook

9. Application Performance Monitoring (APM)

앱 성능 관측.

  • Latency percentiles (p50, p95, p99)
  • Throughput
  • Error rate
  • Apdex score
  • User experience

SRE (Site Reliability Engineering)

Google 이 만든 클라우드 시대 운영 접근법. Operations perspective 의 실현.

핵심 원칙:

  • Automation: 반복 작업 자동화 (toil 제거)
  • Measurement: 모든 것 측정 (SLI/SLO/SLA)
  • Error budget: 100% 가용성 목표 X (혁신과 안정성 균형)
  • Blameless postmortem: 사람 아닌 시스템 문제로

Toil: 자동화 가능한 반복 수동 작업. SRE 는 toil 50% 이하 유지 목표.

SLO / SLI / SLA

SLI (Indicator): 측정 지표. 예: “요청의 99% 가 200ms 이하로 응답”

SLO (Objective): 내부 목표. 예: 99.9% availability/month

SLA (Agreement): 고객과의 계약. 위반 시 배상.

Error Budget = 100% - SLO. 예: SLO 99.9% -> error budget 0.1% = 43분/월.

  • Budget 남으면: 새 기능 배포 OK
  • Budget 소진: 안정성 개선 우선

DORA Metrics

Elite performing team 지표 (Google DevOps Research):

  1. Deployment Frequency: 몇 번 배포 (일/주)
  2. Lead Time for Changes: 코드 커밋 -> 프로덕션
  3. Change Failure Rate: 배포가 문제 유발 %
  4. Time to Restore Service: 사고 복구 시간

Elite: 하루 여러 번 배포, 리드타임 < 1시간, 실패율 < 15%, MTTR < 1시간.

Operational Runbook

표준 운영 절차 문서화. 특정 상황에 따라야 할 단계.

Runbook: RDS Failover
1. CloudWatch alarm 확인 (CPU / Connections)
2. Read replica lag 확인
3. Manual failover 필요 시:
   aws rds failover-db-cluster --db-cluster-identifier prod-db
4. Application 재시작 (필요 시)
5. PIR 준비

Systems Manager Automation runbook 으로 자동화 가능.

Chaos Engineering

의도적 실패 주입 으로 시스템 회복력 검증.

  • AWS Fault Injection Simulator (FIS): EC2 stop, latency 주입 등
  • Netflix Chaos Monkey (원조)
  • Game days (팀 훈련)

Reliability capability 검증.

Automation Priorities

  1. Deployment: CI/CD, IaC
  2. Testing: 자동 통합/E2E 테스트
  3. Monitoring: 알람 자동
  4. Remediation: 자동 복구 (예: EC2 replace)
  5. Scaling: Auto Scaling, HPA
  6. Backup / Restore
  7. Patching

Well-Architected Operational Excellence

Platform perspective 의 6 pillar 중 하나. Operations 는 이 pillar 의 실현.

설계 원칙:

  • Perform operations as code
  • Make frequent, small, reversible changes
  • Refine operations procedures frequently
  • Anticipate failure
  • Learn from all operational failures

실전 활동

Envision

  • 현재 운영 성숙도 진단 (DORA)
  • SLO 정의
  • Incident history 분석

Align

  • Observability 스택 설계
  • On-call 계획
  • Runbook 저장소

Launch

  • CloudWatch dashboard + alarm
  • On-call rotation 시작
  • 첫 GameDay

Scale

  • 자동 remediation 확대
  • SRE 팀 성숙
  • DORA metric 지속 개선

흔한 실패 패턴

WARNING

관측 없이 배포. 문제 발생 시 진단 불가. Metrics/logs/traces 첫날부터.

CAUTION

Alert fatigue. 노이즈 알람 많으면 진짜 사고 놓침. SLO 기반 알람만.

WARNING

Runbook 없이 사고. 매번 처음처럼. 반복되는 문제는 runbook + 자동화.

IMPORTANT

PIR 를 blame 세션으로. 사람 지목은 문화 파괴. 시스템/프로세스 개선 초점.

CAUTION

DR 안 훈련. 실제 사고 시 실패. 정기 DR drill.

관련 위키

이 글의 용어 (12개)
[AWS CAF] Business Perspectivecloud
정의 Business Perspective 는 AWS CAF 6 관점 중 비즈니스 성과 에 초점을 맞춘 관점입니다. 클라우드 투자가 디지털 전환 목표와 비즈니스 성과 (reven…
[AWS CAF] Governance Perspectivecloud
정의 Governance Perspective 는 AWS CAF 6 관점 중 클라우드 이니셔티브의 관리 (orchestration) 를 담당합니다. 조직의 이익을 극대화하면서 전…
[AWS CAF] People Perspectivecloud
정의 People Perspective 는 AWS CAF 6 관점 중 비즈니스 와 기술 사이의 다리 역할을 하는 관점입니다. 문화, 조직 구조, 리더십, 인력 (workforce…
[AWS CAF] Platform Perspectivecloud
정의 Platform Perspective 는 AWS CAF 6 관점 중 엔터프라이즈급 확장 가능한 하이브리드 클라우드 플랫폼 구축 을 담당합니다. 기존 워크로드 현대화 + 신규…
[AWS CAF] Security Perspectivecloud
정의 Security Perspective 는 AWS CAF 6 관점 중 데이터와 워크로드의 기밀성 (Confidentiality), 무결성 (Integrity), 가용성 (Av…
[AWS] Cloud Adoption Framework (CAF)cloud
정의 AWS Cloud Adoption Framework (AWS CAF) 는 AWS 가 수많은 고객의 클라우드 전환 경험을 정리해 만든 조직 수준의 클라우드 도입 방법론 입니다…
[AWS] CloudWatch: 메트릭, 로그, 알람cloud
정의 CloudWatch = AWS 의 모니터링 + 로그 + 알람 통합 서비스. 메트릭 수집, 로그 집계, 대시보드, 알람, 이상 감지를 하나의 서비스에서 제공. 사용 상황 | …
[AWS] Config (Resource Configuration Tracking)cloud
정의 AWS Config 는 AWS 계정 안 리소스의 구성 (configuration) 상태를 지속 기록 하고, 정책 규칙 (Config Rules) 위반을 자동 감지하는 서비스…
[Observability] OpenTelemetry: 표준화된 trace/metric/logdevops
정의 OpenTelemetry (OTel) = observability 의 vendor-neutral 표준. CNCF. trace + metric + log 의 SDK + pro…
[Observability] Prometheus: pull 기반 메트릭, PromQLdevops
정의 Prometheus = pull 기반 시계열 metric 시스템. PromQL 로 쿼리. CNCF graduated. 2026 클라우드 네이티브 메트릭 표준. 아키텍처 Pu…
[Observability] SLI / SLO / SLA / Error Budgetdevops
정의 | | 의미 | |---|---| | SLI (Indicator) | 측정 값 (예: 5xx 비율) | | SLO (Objective) | 목표 (예: 99.9% 가용성) …
[Voice AI] P50/P95/P99: latency 분포와 SLOdevops
정의 P50 / P95 / P99 = latency 분포의 분위수. 평균은 outlier 에 가려져 무의미. 음성 에이전트의 사용자 경험 은 p95/p99 가 결정. [!IMPO…

💬 댓글

사이트 검색 / 명령어

검색

스크롤 = 확대/축소 · 드래그 = 이동 · 0 = 원래 크기 · ESC = 닫기