[AWS CAF] Operations Perspective
정의
Operations Perspective 는 AWS CAF 6 관점 중 클라우드 서비스가 비즈니스가 요구하는 수준으로 딜리버리되도록 보장 하는 관점입니다. 자동화 + 최적화 로 신뢰성 있게 확장.
주요 stakeholder: 인프라/운영 리더, SRE (Site Reliability Engineer), IT service manager.
주요 Capabilities
1. Observability
시스템 상태 가시성. Metrics, Logs, Traces.
AWS 도구:
- CloudWatch Metrics, Logs, Alarms, Dashboards
- CloudWatch Application Signals (2024+)
- X-Ray (tracing)
- OpenTelemetry (표준)
- Managed Prometheus, Managed Grafana
3 pillars of observability:
- Metrics: 수치 시계열 (CPU%, requests/s)
- Logs: 이벤트 (structured JSON)
- Traces: 요청 흐름 (분산 트레이싱)
2. Event Management (AIOps)
이벤트 상관 + 자동 대응.
- Alert de-duplication
- Alert routing (PagerDuty, Opsgenie)
- Auto-remediation (Lambda + Systems Manager)
- ML 기반 이상 감지
3. Incident and Problem Management
사고 대응 프로세스.
- Incident: 실시간 문제 (down)
- Problem: 근본 원인 조사
- PIR (Post-Incident Review): Blameless postmortem
단계:
- Detection
- Triage (severity 판단)
- Response (mitigation)
- Communication (상태 페이지, stakeholder)
- Resolution
- PIR
4. Change and Release Management
프로덕션 변경 관리.
- Change Advisory Board (전통) -> Automated (클라우드)
- Blue/green, canary
- Rollback 자동화
- Change failure rate 추적 (DORA metric)
5. Performance and Capacity Management
성능 목표 (SLO) 정의 + 용량 계획.
- SLO (Service Level Objective): 예 99.9% 가용성
- SLI (Service Level Indicator): 측정 지표
- Error budget: SLO 대비 여유
- Capacity planning: 미래 부하 예측
6. Configuration Management
리소스 설정 관리:
- Config 로 리소스 상태 추적
- Systems Manager Parameter Store
- IaC (immutable, versioned)
7. Patch Management
OS/앱 패치:
- Systems Manager Patch Manager
- 자동 스케줄 (매주 등)
- Compliance 상태 리포트
8. Availability and Continuity Management
가용성 + 재해복구.
- RTO (Recovery Time Objective): 복구까지 시간
- RPO (Recovery Point Objective): 데이터 손실 허용 시간
- Multi-AZ, Multi-region
- DR runbook
9. Application Performance Monitoring (APM)
앱 성능 관측.
- Latency percentiles (p50, p95, p99)
- Throughput
- Error rate
- Apdex score
- User experience
SRE (Site Reliability Engineering)
Google 이 만든 클라우드 시대 운영 접근법. Operations perspective 의 실현.
핵심 원칙:
- Automation: 반복 작업 자동화 (toil 제거)
- Measurement: 모든 것 측정 (SLI/SLO/SLA)
- Error budget: 100% 가용성 목표 X (혁신과 안정성 균형)
- Blameless postmortem: 사람 아닌 시스템 문제로
Toil: 자동화 가능한 반복 수동 작업. SRE 는 toil 50% 이하 유지 목표.
SLO / SLI / SLA
SLI (Indicator): 측정 지표. 예: “요청의 99% 가 200ms 이하로 응답”
SLO (Objective): 내부 목표. 예: 99.9% availability/month
SLA (Agreement): 고객과의 계약. 위반 시 배상.
Error Budget = 100% - SLO. 예: SLO 99.9% -> error budget 0.1% = 43분/월.
- Budget 남으면: 새 기능 배포 OK
- Budget 소진: 안정성 개선 우선
DORA Metrics
Elite performing team 지표 (Google DevOps Research):
- Deployment Frequency: 몇 번 배포 (일/주)
- Lead Time for Changes: 코드 커밋 -> 프로덕션
- Change Failure Rate: 배포가 문제 유발 %
- Time to Restore Service: 사고 복구 시간
Elite: 하루 여러 번 배포, 리드타임 < 1시간, 실패율 < 15%, MTTR < 1시간.
Operational Runbook
표준 운영 절차 문서화. 특정 상황에 따라야 할 단계.
Runbook: RDS Failover
1. CloudWatch alarm 확인 (CPU / Connections)
2. Read replica lag 확인
3. Manual failover 필요 시:
aws rds failover-db-cluster --db-cluster-identifier prod-db
4. Application 재시작 (필요 시)
5. PIR 준비
Systems Manager Automation runbook 으로 자동화 가능.
Chaos Engineering
의도적 실패 주입 으로 시스템 회복력 검증.
- AWS Fault Injection Simulator (FIS): EC2 stop, latency 주입 등
- Netflix Chaos Monkey (원조)
- Game days (팀 훈련)
Reliability capability 검증.
Automation Priorities
- Deployment: CI/CD, IaC
- Testing: 자동 통합/E2E 테스트
- Monitoring: 알람 자동
- Remediation: 자동 복구 (예: EC2 replace)
- Scaling: Auto Scaling, HPA
- Backup / Restore
- Patching
Well-Architected Operational Excellence
Platform perspective 의 6 pillar 중 하나. Operations 는 이 pillar 의 실현.
설계 원칙:
- Perform operations as code
- Make frequent, small, reversible changes
- Refine operations procedures frequently
- Anticipate failure
- Learn from all operational failures
실전 활동
Envision
- 현재 운영 성숙도 진단 (DORA)
- SLO 정의
- Incident history 분석
Align
- Observability 스택 설계
- On-call 계획
- Runbook 저장소
Launch
- CloudWatch dashboard + alarm
- On-call rotation 시작
- 첫 GameDay
Scale
- 자동 remediation 확대
- SRE 팀 성숙
- DORA metric 지속 개선
흔한 실패 패턴
WARNING
관측 없이 배포. 문제 발생 시 진단 불가. Metrics/logs/traces 첫날부터.
CAUTION
Alert fatigue. 노이즈 알람 많으면 진짜 사고 놓침. SLO 기반 알람만.
WARNING
Runbook 없이 사고. 매번 처음처럼. 반복되는 문제는 runbook + 자동화.
IMPORTANT
PIR 를 blame 세션으로. 사람 지목은 문화 파괴. 시스템/프로세스 개선 초점.
CAUTION
DR 안 훈련. 실제 사고 시 실패. 정기 DR drill.
관련 위키
이 글의 용어 (12개)
- [AWS CAF] Business Perspectivecloud
- 정의 Business Perspective 는 AWS CAF 6 관점 중 비즈니스 성과 에 초점을 맞춘 관점입니다. 클라우드 투자가 디지털 전환 목표와 비즈니스 성과 (reven…
- [AWS CAF] Governance Perspectivecloud
- 정의 Governance Perspective 는 AWS CAF 6 관점 중 클라우드 이니셔티브의 관리 (orchestration) 를 담당합니다. 조직의 이익을 극대화하면서 전…
- [AWS CAF] People Perspectivecloud
- 정의 People Perspective 는 AWS CAF 6 관점 중 비즈니스 와 기술 사이의 다리 역할을 하는 관점입니다. 문화, 조직 구조, 리더십, 인력 (workforce…
- [AWS CAF] Platform Perspectivecloud
- 정의 Platform Perspective 는 AWS CAF 6 관점 중 엔터프라이즈급 확장 가능한 하이브리드 클라우드 플랫폼 구축 을 담당합니다. 기존 워크로드 현대화 + 신규…
- [AWS CAF] Security Perspectivecloud
- 정의 Security Perspective 는 AWS CAF 6 관점 중 데이터와 워크로드의 기밀성 (Confidentiality), 무결성 (Integrity), 가용성 (Av…
- [AWS] Cloud Adoption Framework (CAF)cloud
- 정의 AWS Cloud Adoption Framework (AWS CAF) 는 AWS 가 수많은 고객의 클라우드 전환 경험을 정리해 만든 조직 수준의 클라우드 도입 방법론 입니다…
- [AWS] CloudWatch: 메트릭, 로그, 알람cloud
- 정의 CloudWatch = AWS 의 모니터링 + 로그 + 알람 통합 서비스. 메트릭 수집, 로그 집계, 대시보드, 알람, 이상 감지를 하나의 서비스에서 제공. 사용 상황 | …
- [AWS] Config (Resource Configuration Tracking)cloud
- 정의 AWS Config 는 AWS 계정 안 리소스의 구성 (configuration) 상태를 지속 기록 하고, 정책 규칙 (Config Rules) 위반을 자동 감지하는 서비스…
- [Observability] OpenTelemetry: 표준화된 trace/metric/logdevops
- 정의 OpenTelemetry (OTel) = observability 의 vendor-neutral 표준. CNCF. trace + metric + log 의 SDK + pro…
- [Observability] Prometheus: pull 기반 메트릭, PromQLdevops
- 정의 Prometheus = pull 기반 시계열 metric 시스템. PromQL 로 쿼리. CNCF graduated. 2026 클라우드 네이티브 메트릭 표준. 아키텍처 Pu…
- [Observability] SLI / SLO / SLA / Error Budgetdevops
- 정의 | | 의미 | |---|---| | SLI (Indicator) | 측정 값 (예: 5xx 비율) | | SLO (Objective) | 목표 (예: 99.9% 가용성) …
- [Voice AI] P50/P95/P99: latency 분포와 SLOdevops
- 정의 P50 / P95 / P99 = latency 분포의 분위수. 평균은 outlier 에 가려져 무의미. 음성 에이전트의 사용자 경험 은 p95/p99 가 결정. [!IMPO…
💬 댓글