The control room
One dashboard watched the whole shop. Each panel answers one question: are we losing money, or is the site down?

Eight alarms, one mailbox
Each alarm mails us before buyers notice a problem. Three of them show “insufficient data” when healthy. That means zero errors, not a broken alarm.
| Alarm | Fires when | Why this number |
|---|---|---|
| alb-5xx | 5xx > 10 / 5 min | Flask 500s only on real crashes |
| alb-unhealthy | unhealthy > 0, 2 min | Dead counter before ASG replace finishes |
| rds-cpu | CPU > 80% ×3 | db.t3.micro is bursty; sustained 80% = leak |
| rds-storage | free < 5 GB | Early warning on 20 GB gp2 |
| sqs-age | oldest > 300 s | Order sits 5 min unserved |
| dlq-depth | visible ≥ 1 | Any poison order pages us |
| worker-errors | errors > 0 | Packer crashed: bad payload, SNS deny, DDB fail |
| ddb-throttle | throttles > 0 | Hot orderId spike |
Receipts, not claims
We broke each part by hand and fixed it by hand. Scroll sideways. Each frame is a real screenshot from the build.





Day logs live in docs/days/ (day-0 to day-7). Run git shortlog -sn --no-merges. Our two names take turns. That is the receipt.
Keeping it cheap
$3.01/day when up
Biggest costs: the database, the load balancer, and the network gateway. Rule: shut down the compute, data, and queue stacks every night; keep network and storage. A $20 monthly budget mails us at 60% spent and at 100% forecast. Total spent building this: about $20 to $25.

Try the order flow
The AWS behind this is deleted, so the demo below runs the same checks in your browser. Totals must add up. Reused ids get rejected. Packing follows receiving after a short wait.
202 {"orderId":"ord-000050","status":"RECEIVED"}
