Real-time cloud cost control for AWS and Kubernetes

Your bill tells you a day late. Cloudblame tells you in minutes.

Cloudblame turns CloudTrail, CloudWatch and Kubernetes signals into named findings: the deploy, alert rule or service account behind a cost spike, a reversible brake, and savings verified on your own bill.

Connects in minutes. Read-only role with your external ID, nothing to install.

FEED · FROM THE DOCUMENTED EKS AND ATHENA CASES

Reads from 14 sources. Writes to none.

AWS CloudTrail AWS CloudTrail CloudWatch CloudWatch Cost Explorer Cost Explorer Cost and Usage Report Cost and Usage Report EKS EKS Karpenter Karpenter Argo CD Argo CD GitHub GitHub Terraform Terraform Slack Slack Cloudflare Cloudflare Datadog Datadog Grafana Grafana Athena Athena

CLOUD COST ENGINEERING PLATFORM

One loop for detection, attribution, brakes and verification.

Cloudblame reads your usage signals as they happen, names the change behind a spike, hands you a reversible brake and verifies the result on your own bill. One finding id carries through all four steps.

01 / 04 · Detect

Detect the spike while it is still cheap.

Cloudblame watches CloudTrail, CloudWatch and Kubernetes events as they happen and raises an alarm minutes after the run rate moves, not the day after the billing export lands.

See the case file
Run rate · EC2 · 7 days +$34/h Spike · 14:07 14:07
Signals · 7 days3 signals
SignalRun rateStatus
EC2 RunInstances ×22karpenter-prod +$34/h Spike
Athena 1.9 TB / 5 mingrafana-sa $912/day Spike
Pending pods ×38game-services 320 vCPU Named 14:09

READ-ONLY BY DESIGN

The exact policy and the CloudFormation template are public: policy.json · readonly-role.yaml

  • IAM read-only role
  • External ID
  • ~55 actions, published policy
  • No reseller, no billing transfer
  • Queries run in your account
  • Delete the stack to revoke

PROBLEM

Cost spikes don't wait for the bill.

A merged pull request, an alert rule or a service account can move your AWS bill within minutes. Cost Explorer and anomaly alerts report it a day later as service, account and usage type, and leave the team to find the change by hand.

Cloudblame closes the loop between usage signals and a verified fix.

Run rate at risk $0/h

MANUAL INTERVENTION REQUIRED. CLICK PODS TO BRAKE.

THE CORRELATOR

The engine that joins your bill to your deploys.

Six links from a line on your bill to the pull request and the team behind it. AWS Cost Anomaly Detection stops at the account and the IAM principal; Cloudblame keeps walking the chain until it has a name.

6 LINKS · 1 FINDING ID

Evidence chain: CUR line, IAM principal, NodeClaim, pending pods, Argo CD revision, PR #1832 and team. AWS stops at the IAM principal; Cloudblame continues to the pull request. CUR LINE EC2 · $811/day IAM PRINCIPAL karpenter-prod NODECLAIM general-x7k2q PENDING PODS 38 · game-services ARGO CD REVISION 4f2e… PR #1832 + TEAM gameplay · 14:02 AWS STOPS HERE CLOUDBLAME CONTINUES Evidence chain: CUR line, IAM principal, NodeClaim, pending pods, Argo CD revision, PR #1832 and team. AWS stops at the IAM principal; Cloudblame continues to the pull request. CUR LINE EC2 · $811/day IAM PRINCIPAL karpenter-prod NODECLAIM general-x7k2q PENDING PODS 38 · game-services ARGO CD REVISION 4f2e… PR #1832 + TEAM gameplay · 14:02 AWS STOPS HERE CLOUDBLAME CONTINUES

Below the IAM principal

AWS tells you which role launched the nodes. Cloudblame follows the NodeClaim to the pending pods, the Argo CD revision and the pull request that merged at 14:02.

Frozen unit price

Savings are measured on your Cost and Usage Report against a unit price frozen at the finding. Later price changes and discounts do not move the number.

Reversible brakes

Every brake is a change you run in your own account and can undo: revert the PR, cap the node pool, widen the alert interval.

CASE FILE

22 nodes in four minutes. Named in seven.

A pull request raised CPU requests from 1 to 8 vCPU on 40 replicas; Karpenter added 22 nodes in four minutes. AWS reported it the next day as “EC2 usage increased”. Cloudblame named the PR, the team and a one-line fix at 14:09.

EKS · MATCHMAKING · PR #1832 · $24,300/MONTH VERIFIED

Red: Cloudblame. Grey: AWS, the next day. Green: verified on your Cost and Usage Report.

EKS · game-services/matchmaking
Finding #003 · node pool general · eu-central-1
Verified

  1. 14:02 PR #1832 merged: matchmaking CPU requests 1 → 8 vCPU, 40 replicas. Argo CD syncs.
  2. 14:03 320 vCPU requested, 38 pods pending. Karpenter starts launching nodes.
  3. 14:07 Cloudblame alarm: 22 × m6i.8xlarge launched by karpenter in 4 minutes, ≈ $34/h at list.
  4. 14:09 Slack card: who, what, why, and a brake. Revert the PR, or cap the node pool.
  5. 14:26 Requests reverted by the gameplay team. Run rate was $811 a day.
  6. 24 hours later
  7. 10:40 AWS Cost Anomaly Detection: "EC2 usage increased." Root cause: service, account, region, usage type.
  8. at month end
  9. Oct 15 Cost and Usage Report confirms $24,300 a month avoided. Fee: $2,001 a month under the published tiers.

HOW IT WORKS

From read-only role to verified savings.

  1. Connect

    Launch one CloudFormation stack: a single read-only role, your external ID. About ten minutes, nothing installed in your cluster.

    ~10 min · read-only
  2. Scan

    A free one-page answer: what AWS already explains, what it cannot, and a clear go or no-go before you spend another hour.

    free · one page
  3. Investigate

    Named findings with apply-ready fixes and a reversible brake. You approve every change; nothing runs without you.

    you approve every change
  4. Verify

    The delta is measured on your Cost and Usage Report against a frozen baseline. You pay only on that number.

    measured on your CUR
Run a free cost scan
CLOUDFORMATION · STACKS CREATE_COMPLETE
cloudblame-readonly eu-central-1 · created in 0:58
KeyValueDescription
RoleArn arn:aws:iam::123456789012:role/CloudCostControlReadOnly Paste into step 2 of the scan
ExternalId 7f3a19c4…e91b Generated in your browser
PolicyActions 55 ce, cloudtrail, cloudwatch, athena, ec2, eks · read-only
CurQueries false Switch on later for the measurement annex
DELETE THE STACK TO REVOKE · NOTHING INSTALLED IN THE CLUSTER

INTEGRATIONS

Reads the stack you already run. Changes nothing in it.

Every signal comes from services you already pay for. Findings land where your team already looks: a Slack card and a comment on the pull request that caused the spike.

READS · ONE READ-ONLY ROLE, ~55 ACTIONS

AWS · the published role

  • CloudTrail cloudtrail:LookupEvents
  • CloudWatch cloudwatch:GetMetricData · logs:StartQuery
  • Cost Explorer ce:GetCostAndUsage · ce:GetAnomalies
  • Cost and Usage Report athena:StartQueryExecution · s3:GetObject on the CUR bucket
  • EC2 and NAT ec2:DescribeInstances · ec2:DescribeNatGateways
  • EKS eks:DescribeCluster · eks:ListNodegroups
  • Cost Optimization Hub cost-optimization-hub:ListRecommendations

Around AWS · read-only tokens

  • Kubernetes API get, list on pods, deployments, nodeclaims · view ClusterRole
  • Argo CD applications: get
  • GitHub contents: read · pull requests: read
  • Grafana · Datadog alert rules: read · usage: read
  • Cloudflare Analytics: Read · Cache Rules: Read

The exact list is public: policy.json · readonly-role.yaml

WRITES BACK · TWO PLACES, NOTHING ELSE

Cloudblame
14:09 · #cloud-cost · finding #003
Spike

22 × m6i.8xlarge launched for game-services/matchmaking in 4 minutes

Who
gameplay team · PR #1832 · Argo CD app matchmaking
Run rate
$34/h at list · $811/day
chat:write · one channel you choose
cloudblame commented on PR #1832 · 14:09 · finding #003

This PR raised matchmaking CPU requests 1 → 8 vCPU on 40 replicas. Karpenter added 22 nodes (704 vCPU) to node pool general at 14:03; p95 usage is 0.4 vCPU per pod.

Suggested fix: requests back to 1 vCPU, HPA at 70% → PR #1840

pull requests: write · a comment, never a merge

READ-ONLY ROLE · ~55 ACTIONS, PUBLISHED POLICY · NOTHING INSTALLED IN THE CLUSTER

PRICING

Pay only when we create verified savings.

Net means after any cost the fix introduced, like a VPC endpoint. Verified means measured on your Cost and Usage Report against a frozen baseline. Recommendations are never invoiced.

Read the measurement annex
  • 0% on commitments
  • 12 months, then $0
  • No fixed fees
Fee schedule
LineFee
Commitments (RI / Savings Plans)0%
First $10k / month10%
$10–30k / month7%
$30–60k / month5%
Above $60k / month4%
Per finding12 months, then $0
Under $500 / monthnot billed
Diagnostic, platform, setup feesnone
$20,000
Fee / month
$1,700
Effective rate
8.5%
12-month total
$20,400

A 20%-of-annual consultancy would take $48,000 for the same savings.

Statement · November
Statement · November
Layout of a month-end statement · baselines frozen at the fix date
EXAMPLE STATEMENT · ILLUSTRATIVE FIGURES

FindingBaselineActualNet savingFee
#003 EKS matchmaking requests 8 → 1 vCPU (PR #1832) Verified 704 vCPU 64 vCPU $24,300
#005 NAT → S3 gateway endpoint (payments) Illustrative 2.1 TB/day 0.3 TB/day $4,860
#007 DynamoDB orders table, provisioned → on-demand Illustrative 3,000 WCU 640 WCU $2,180
Net savings in this example $31,340
Fee · 10% of first $10k + 7% of next $20k + 5% of $1,340 $2,467
#003 is the documented EKS case, verified on the Cost and Usage Report. #005 and #007 are illustrative rows that show the layout. In a real statement every row links to the Athena query that produced it, saved in your account.

PLAYBOOKS

Eleven playbooks. One finding id.

Each card is a playbook: the signal we watch, how fast it fires, where AWS stops and what we add. The same finding id follows it from the Slack card to the month-end statement.

11 PLAYBOOKS · AWS SEES / WE ADD

Query storms (Athena, Logs Insights)

Signal in minutes

AWS seesService, usage type, principal

We addThe alert rule or dashboard, its owner, the fix

Log bloat (CloudWatch Logs)

Signal in minutes

AWS seesUsage type

We addThe log group, the deploy that raised the level, who queries it

NAT and data transfer

Signal in minutes

AWS seesAn idle NAT

We addThe pod and the missing VPC endpoint

Over-provisioned and forgotten resources

Signal the same day

AWS seesThe list (free)

We addWho created it, when, and the PR to remove it

Kubernetes idle requests

Signal the same day

AWS seesNode level

We addNamespace, workload, owner, the requests to change

Lambda retry storms

Signal in minutes

AWS seesThe principal

We addThe retry setting and the loop

FREE SCAN · 10 MINUTES · READ-ONLY

Go from surprised by the bill to warned in minutes.