Back to skills
SKILL.md
Cloud Architect
BSecurityUse when configuring, automating, deploying, and debugging cloud architect pipelines, containers, servers, and cloud infrastructure.
- 5 stars
- 0 votes
- 0 copies
- 0 views
- Added September 27, 2026
Works with
Security analysis
88/100- Exfiltrates credentials via HTTP β exact pattern from Snyk ToxicSkills study
npx -y skills add Harmitx7/tribunal-kit --skill cloud-architect --agent claude-codeAre you the author of Cloud Architect?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/harmitx7-cloud-architect)---
name: cloud-architect
description: "Use when configuring, automating, deploying, and debugging cloud architect pipelines, containers, servers, and cloud infrastructure."
version: 6.0.0
last-updated: 2026-09-29
skills:
- cicd-pro
- devops-engineer
- platform-engineer
tools: Read, Grep, Glob, Bash, Edit, Write
scripts-binding:
- .agent/scripts/lint_runner.js
- .agent/scripts/verify_all.js
---
# Cloud Architect β Production AWS Mastery
## Mandatory Pre-Flight Context Inspection
Before reading, generating, or refactoring code in the `cloud-architect` domain, inspect these 5 critical parameters:
1. **System Boundaries & Dependencies**: Verify that all required dependencies exist in target package manifests and environment paths.
2. **Runtime Context & Platform Invariants**: Confirm target platform constraints (Node.js, Browser, Mobile OS, Edge runtime) before applying APIs.
3. **Execution Guardrails**: Identify potential side-effects, state mutations, and unhandled asynchronous exceptions.
4. **Validation & Type Contracts**: Validate input data schemas and strict type constraints across all module interfaces.
5. **Observability & Proof of Execution**: Ensure execution produces tangible verification signals (terminal output, tests, metrics).
## Activation Boundaries
- **Activate when:** Use when configuring, automating, deploying, and debugging cloud architect pipelines, containers, servers, and cloud infrastructure.
- **DO NOT activate when:** The task falls outside the `cloud-architect` domain or is managed by a different dedicated specialist agent.
## π Multi-Pass Execution Protocol
| Pass | Phase | Core Action | Adaptive Depth |
|:---|:---|:---|:---|
| **Pass 1** | **Understand** | Deconstruct the user's explicit objective, implicit requirements, and platform constraints. | Fast / Standard / Deep |
| **Pass 2** | **Plan** | Decompose task into smallest logical steps; map dependencies, affected files, and tool calls. | Standard / Deep |
| **Pass 3** | **Execute** | Implement solution with production-grade craft, zero placeholders, and strict typing. | All Modes |
| **Pass 4** | **Verify** | Run linters, unit tests, or compiler checks to validate structural correctness. | All Modes |
| **Pass 5** | **Attack & Falsify** | Perform adversarial search for edge-case failures, counterexamples, race conditions, and traps. | Standard / Deep |
| **Pass 6** | **Harden** | Eliminate discovered friction, optimize performance, and harden error boundaries. | Standard / Deep |
| **Pass 7** | **Quality Gate** | Enforce Verification-Before-Completion (VBC) with concrete terminal proof before finalizing. | All Modes |
---
## π οΈ Technical Architecture & Reference Recipes
## Hallucination Traps (Read First)
- β Confusing availability zones with regions β β
`us-east-1` is a region. `us-east-1a` is an AZ within that region. Multi-AZ β multi-region.
- β `s3:*` or `*:*` IAM policies β β
IAM policies must follow least privilege. Enumerate exact actions required.
- β Hardcoding ARNs, account IDs, or region strings β β
Always use `data.aws_caller_identity.current.account_id`, `var.region`, Terraform variables.
- β Storing secrets in environment variables on EC2/ECS β β
Use AWS Secrets Manager. Reference ARN in task definition, not plaintext values.
- β Lambda for all compute β β
Lambda cold starts hurt latency-sensitive APIs. ECS Fargate for always-warm containers.
- β Public subnets for application servers β β
Application tier lives in private subnets. Only ALB and NAT Gateway in public subnets.
---
## 1. Service Selection Matrix
### Compute
```
ECS Fargate:
β
Long-running HTTP services (APIs, web apps)
β
Container-based workloads (0 server management)
β
Predictable traffic (no cold start concerns)
β
>100ms response time acceptable
Cost: ~$0.04048/vCPU/hr + $0.004445/GB/hr
Lambda:
β
Event-driven (S3 trigger, SQS consumer, API Gateway)
β
Infrequent, short-duration tasks (<15 min)
β
True scale-to-zero requirements (cost optimization)
β Cold starts (100msβ1s for first request)
β Stateful connections (DB connection pooling β use RDS Proxy)
Cost: $0.20/1M requests + $0.0000166667/GB-second
EC2:
β
Specialized hardware (GPU, high I/O)
β
Legacy apps that can't containerize
β
Reserved capacity for cost optimization at scale
β Requires patch management and OS maintenance
Cost: Varies, reserved instances 40-60% cheaper than on-demand
```
### Database
```
RDS PostgreSQL:
β
ACID transactions required
β
Complex JOIN queries
β
Team knows SQL
β
<10TB data
Multi-AZ: automatic failover in ~60s
Read replicas: up to 5 (async replication)
Aurora PostgreSQL:
β
RDS PostgreSQL but need: higher availability, faster failover (<30s)
β
>5 read replicas needed (up to 15 Aurora replicas)
β
Global database (multi-region reads)
β
Aurora Serverless v2 for variable/unpredictable load
Cost: ~2x RDS, worth it for high-availability requirements
DynamoDB:
β
Single-digit ms latency at any scale
β
Key-value or simple document access patterns
β
Global tables (multi-region active-active)
β
Event-driven (DynamoDB Streams β Lambda)
β Complex queries (no joins, limited filtering)
β Requires careful partition key design
On-Demand pricing: $1.25/1M writes, $0.25/1M reads
```
---
## 2. VPC Architecture (Golden Pattern)
```
ββββββββββββββββββββ VPC (10.0.0.0/16) βββββββββββββββββββββ
β ββββ Public Subnet A βββ ββββ Public Subnet B βββ β
β β 10.0.1.0/24 β β 10.0.2.0/24 β β
β β - ALB β β - ALB (multi-AZ) β β
β β - NAT Gateway β β - NAT Gateway β β
β ββββββββββββββββββββββββ ββββββββββββββββββββββββ β
β ββββ Private Subnet A ββ ββββ Private Subnet B ββ β
β β 10.0.11.0/24 β β 10.0.12.0/24 β β
β β - ECS Tasks β β - ECS Tasks β β
β β - Lambda (VPC) β β - Lambda (VPC) β β
β ββββββββββββββββββββββββ ββββββββββββββββββββββββ β
β ββββ Data Subnet A βββββ ββββ Data Subnet B βββββ β
β β 10.0.21.0/24 β β 10.0.22.0/24 β β
β β - RDS Primary β β - RDS Replica β β
β β - ElastiCache β β - ElastiCache β β
β ββββββββββββββββββββββββ ββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
```
### Terraform: VPC Module
```hcl
# modules/networking/main.tf
module "vpc" {
source = "terraform-aws-modules/vpc/aws"
version = "~> 5.0"
name = "${var.project}-${var.environment}"
cidr = var.vpc_cidr # e.g., "10.0.0.0/16"
azs = data.aws_availability_zones.available.names
public_subnets = var.public_subnet_cidrs # ["10.0.1.0/24", "10.0.2.0/24"]
private_subnets = var.private_subnet_cidrs # ["10.0.11.0/24", "10.0.12.0/24"]
database_subnets = var.database_subnet_cidrs # ["10.0.21.0/24", "10.0.22.0/24"]
enable_nat_gateway = true
single_nat_gateway = var.environment == "staging" # save cost in staging
enable_dns_hostnames = true
enable_dns_support = true
# Required for ECS/EKS
public_subnet_tags = {
"kubernetes.io/role/elb" = 1
}
private_subnet_tags = {
"kubernetes.io/role/internal-elb" = 1
}
tags = local.common_tags
}
```
---
## 3. ECS Fargate Service (Complete Terraform)
```hcl
# modules/ecs/main.tf
# Cluster
resource "aws_ecs_cluster" "main" {
name = "${var.project}-${var.environment}"
setting {
name = "containerInsights"
value = "enabled" # CloudWatch Container Insights
}
tags = local.common_tags
}
# Task Definition
resource "aws_ecs_task_definition" "api" {
family = "${var.project}-api"
requires_compatibilities = ["FARGATE"]
network_mode = "awsvpc"
cpu = var.task_cpu # e.g., 512 (0.5 vCPU)
memory = var.task_memory # e.g., 1024 (1 GB)
execution_role_arn = aws_iam_role.ecs_execution.arn
task_role_arn = aws_iam_role.ecs_task.arn
container_definitions = jsonencode([{
name = "api"
image = "${var.ecr_repository_url}:${var.image_tag}"
portMappings = [{
containerPort = var.container_port
protocol = "tcp"
}]
# β
Secrets from Secrets Manager (not plaintext env vars)
secrets = [
{ name = "DATABASE_URL", valueFrom = aws_secretsmanager_secret.db_url.arn },
{ name = "JWT_SECRET", valueFrom = aws_secretsmanager_secret.jwt.arn }
]
environment = [
{ name = "NODE_ENV", value = var.environment },
{ name = "PORT", value = tostring(var.container_port) }
]
logConfiguration = {
logDriver = "awslogs"
options = {
"awslogs-group" = aws_cloudwatch_log_group.api.name
"awslogs-region" = var.aws_region
"awslogs-stream-prefix" = "api"
}
}
healthCheck = {
command = ["CMD-SHELL", "wget --quiet --tries=1 --spider http://localhost:${var.container_port}/health || exit 1"]
interval = 30
timeout = 5
retries = 3
startPeriod = 10
}
}])
tags = local.common_tags
}
# ECS Service
resource "aws_ecs_service" "api" {
name = "${var.project}-api"
cluster = aws_ecs_cluster.main.id
task_definition = aws_ecs_task_definition.api.arn
desired_count = var.desired_count
launch_type = "FARGATE"
network_configuration {
subnets = module.vpc.private_subnets
security_groups = [aws_security_group.api.id]
assign_public_ip = false # β
private subnets never get public IPs
}
load_balancer {
target_group_arn = aws_lb_target_group.api.arn
container_name = "api"
container_port = var.container_port
}
deployment_configuration {
maximum_percent = 200
minimum_healthy_percent = 50
deployment_circuit_breaker {
enable = true
rollback = true # auto-rollback on failure
}
}
depends_on = [aws_lb_listener.https]
tags = local.common_tags
}
```
---
## 4. IAM Least Privilege
```hcl
# β
Execution Role β what ECS needs to START your container
resource "aws_iam_role" "ecs_execution" {
name = "${var.project}-ecs-execution"
assume_role_policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Principal = { Service = "ecs-tasks.amazonaws.com" }
Action = "sts:AssumeRole"
}]
})
}
resource "aws_iam_role_policy_attachment" "ecs_execution_managed" {
role = aws_iam_role.ecs_execution.name
policy_arn = "arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
}
# Allow pulling secrets from Secrets Manager
resource "aws_iam_role_policy" "ecs_execution_secrets" {
name = "secrets-access"
role = aws_iam_role.ecs_execution.id
policy = jsonencode({
Version = "2012-10-17"
Statement = [{
Effect = "Allow"
Action = ["secretsmanager:GetSecretValue"]
Resource = [
aws_secretsmanager_secret.db_url.arn,
aws_secretsmanager_secret.jwt.arn
]
# β
Scoped to exact secret ARNs, not secretsmanager:*
}]
})
}
# β
Task Role β what your APPLICATION code can do at runtime
resource "aws_iam_role_policy" "ecs_task_app" {
name = "app-permissions"
role = aws_iam_role.ecs_task.id
policy = jsonencode({
Version = "2012-10-17"
Statement = [
{
Effect = "Allow"
Action = ["s3:GetObject", "s3:PutObject"]
Resource = "${aws_s3_bucket.uploads.arn}/*"
# β
Specific bucket, specific actions β not s3:*
},
{
Effect = "Allow"
Action = ["ses:SendEmail"]
Resource = "arn:aws:ses:${var.aws_region}:${data.aws_caller_identity.current.account_id}:identity/*"
}
]
})
}
```
---
## 5. CloudWatch Observability
```hcl
# Structured JSON logging β filter patterns work
resource "aws_cloudwatch_log_group" "api" {
name = "/ecs/${var.project}-${var.environment}/api"
retention_in_days = 30
tags = local.common_tags
}
# Metric filter β count 5xx errors from structured logs
resource "aws_cloudwatch_log_metric_filter" "errors_5xx" {
name = "${var.project}-5xx-errors"
pattern = "{ $.status >= 500 }"
log_group_name = aws_cloudwatch_log_group.api.name
metric_transformation {
name = "5xxErrorCount"
namespace = "${var.project}/API"
value = "1"
unit = "Count"
}
}
# Alarm β SNS β PagerDuty
resource "aws_cloudwatch_metric_alarm" "high_error_rate" {
alarm_name = "${var.project}-high-5xx-${var.environment}"
comparison_operator = "GreaterThanThreshold"
evaluation_periods = "2"
metric_name = "5xxErrorCount"
namespace = "${var.project}/API"
period = "60"
statistic = "Sum"
threshold = "10" # more than 10 errors in 2 minutes
alarm_actions = [aws_sns_topic.alerts.arn]
ok_actions = [aws_sns_topic.alerts.arn]
treat_missing_data = "notBreaching"
tags = local.common_tags
}
```
---
## 6. Multi-Environment Strategy
```hcl
# environments/staging/main.tf
module "app" {
source = "../../modules/app"
environment = "staging"
desired_count = 1 # single task in staging
task_cpu = 256 # 0.25 vCPU
task_memory = 512
single_nat_gateway = true # save $32/month vs multi-NAT
}
# environments/production/main.tf
module "app" {
source = "../../modules/app"
environment = "production"
desired_count = 3 # 3 tasks across 2 AZs
task_cpu = 1024 # 1 vCPU
task_memory = 2048
single_nat_gateway = false # high availability
}
```
### Terraform State Backend
```hcl
# backend.tf
terraform {
backend "s3" {
bucket = "mycompany-terraform-state"
key = "production/app.tfstate"
region = "us-east-1"
encrypt = true # AES-256 at rest
dynamodb_table = "terraform-locks" # prevents concurrent state writes
}
}
```
## π¨ Edge-Case & Failure Mode Matrix
| Scenario | Risk | Production Mitigation |
|:---|:---|:---|
| **Empty or Null Inputs** | Unhandled exception or unexpected rendering collapse | Enforce fallback guards, optional chaining, and explicit empty state handlers |
| **Network Timeout / Latency** | Hanging operations or duplicate side-effects | Implement bounded abort controllers, exponential backoff, and idempotency keys |
| **Concurrency / Race Conditions** | Stale state overwrite or inconsistent data mutations | Use atomic transactions, mutex locking, or cancel-on-resubmit controls |
| **Invalid Schema / Malformed Payload** | Downstream runtime errors or security injection | Validate boundary payloads with Zod/Pydantic schemas prior to execution |
| **Resource / Memory Saturation** | OOM errors, frame drops, or memory leaks | Clean up listeners, cancel active timers, and enforce pagination/virtualization |
## ποΈ Tribunal Verification & Guardrails
**Active Reviewers:** `pipeline-reviewer` Β· `devops-engineer` Β· `resilience-reviewer`
**Slash Command:** `/review` or `/tribunal-full`
### π¬ Evidence Standard (Tri-State Verification)
Every finding, audit statement, or completion claim must classify its factual certainty:
- **`[OBSERVED]`**: Directly confirmed in the codebase or verified via executed terminal command.
- **`[INFERRED]`**: Logically deduced from code patterns, architectural data flow, or schema relations.
- **`[UNVERIFIED]`**: Speculative hypothesis or runtime possibility requiring active testing or measurement.
### β
Pre-Flight Self-Audit Checklist
```
β
Are strict execution modes (set -euo pipefail) active on all scripts?
β
Are container images pinned to immutable digest/SHA tags instead of "latest"?
β
Are deployment health checks, liveness probes, and rollback baselines configured?
β
Are CI secrets masked and unexposed to untrusted pull requests?
β
Did I verify environment compatibility across target runtimes?
```
### π Verification-Before-Completion (VBC) Protocol
**CRITICAL:** You must follow a strict "evidence-based closeout" state machine.
- β **Forbidden:** Declaring a task complete because the output "looks correct."
- β
**Required:** You are explicitly forbidden from finalizing any task without providing **concrete evidence** (terminal output, passing test suites, compiler success, or equivalent operational proof) that your output works as intended.
Attribution
Comments
Loading commentsβ¦