CostEQ  ·  Engagement teardown  ·  No. 01

$78k/mo $46k/mo

Series B fintech, ~200 engineers. Anthropic, OpenAI and AWS. Forty-one days from first invoice to verified reduction. No change to model output quality, no change to system performance.

What follows is how that number was found. Same spend, same models, two ways of looking at it — only one of which can be acted on.

01 / Before  ·  Vendor telemetry

The provider console bills correctly and explains nothing.

This is the complete picture available from the platform's own usage and cost view. Totals are accurate and real-time. Attribution stops at the API key, and four of the seven keys were shared across teams.

Claude Platform · Usage & Cost · last 30 days
Claude Platform · Usage & Cost · last 30 days
What is present: spend, input and output tokens, cache hit rate, per-key and per-model breakdown, CSV export. What is absent: any owner, any team, any workflow, and any measure of what the spend produced. The owner column reads unknown for every row because the console has no concept of one.
Question asked

Which team spent the $31,880?
Key shared by platform, QE and data science.

Question asked

What did we get for it?
No link to tickets, PRs or tests.

Question asked

How many agents are running?
Agent traffic is indistinguishable from human.

The gap is not a defect in the console. It is a billing system, and it bills well. Cost data lives in one platform and delivery data lives in another, and nobody had joined them. That join is the entire engagement.

02 / After  ·  Jira + eazyBI

The same tokens, resolved to people, teams and output.

Rebuilt inside the client's own Jira, using eazyBI REST source cubes against APIs they already owned. Every figure below traces to a person, a squad, an epic and a delivery artifact — and every panel carries a recommendation rather than only an observation.

Jira · Dashboards · Claude API Adoption, Efficiency & Spend
Jira · Dashboards · Claude API Adoption, Efficiency & Spend
Filter strip scopes by squad, model, epic and period, so finance and engineering read the same cube at different altitudes. Recommended cost savings is the panel that changed the conversation: each row is a named, sized, owned action — right-size opus workloads to sonnet, raise cache hit rate, move async work to the Batches API, trim output tokens, route defect triage to haiku, retire idle keys. Token-to-output ratio is measured per merged PR for developers and per ticket closed for everyone else, because a single denominator would have flattered one group and punished another.

Vendor console reports

  • Spend per API key
  • Tokens in, tokens out
  • Model mix, org-wide
  • Thirty days of history
  • An observation

The rebuild reports

  • Spend per person, squad and epic
  • Cost per merged PR, ticket closed, test authored
  • Model mix against task class, with a routing recommendation
  • Baselined against each team's own pre-AI delivery rate
  • A ranked action with an owner and an expected saving

03 / Quality as the control variable

A saving that degrades delivery is not a saving.

Cost measures alone will reward a team for producing less. Every efficiency figure is therefore read alongside the quality signal for the same sprint, drawn from Xray, SonarQube, Snyk and Jenkins. A cost reduction that moves any of them the wrong way is reported as a regression.

Jira · Dashboards · QE & Quality Outcomes — AI-assisted delivery
Jira · Dashboards · QE & Quality Outcomes — AI-assisted delivery
Authored-by filter separates human, Claude-assisted and Rovo agent work, resolved through the Bitbucket branch key rather than self-reporting. Squad detail is the row that matters: output, coverage, escaped defects, net Snyk findings, spend and cost-per-test on one line. Ledger shows high authoring volume with a failing gate and net-positive vulnerabilities — volume without quality reads as cost here, which is the behaviour the dashboard was built to produce.

Every dashboard in a cost engagement should be capable of embarrassing the engagement. If nothing on it can ever look bad, it is marketing rather than measurement.

04 / How the join works

Fourteen systems, read-only, joined on two keys.

Nothing sits in the delivery path. Every connector is a scheduled read against an API the client already owned, normalized to the Jira issue key and the commit SHA — the two identifiers every one of these tools already carries.

Connector architecture · source → ingest → model → surface
Connector architecture · source → ingest → model → surface
Systems of record (Jira, Bitbucket, Confluence, Xray, Outlook, Slack) supply the delivery signal. Build and quality tooling (Jenkins, SonarQube, Snyk via the Bitbucket extension, JFrog Artifactory, AWS Cost and Usage Report) supplies the control variables. AI telemetry arrives from the Claude Platform usage and cost endpoints and from Rovo agent runs. eazyBI models all of it into one cube and publishes into Jira dashboards the client already knows how to read.
Join 01

Attribute the spend. Usage rows carry an api_key_id; a managed reference table in Jira maps each key to a person, squad and cost centre.

Join 02

Attach it to output. Branch names, build stamps and Xray executions all carry the Jira key, so cost lands on a ticket rather than on a token count.

Join 03

Divide, then compare. Calculated measures produce cost per unit of delivery, compared against that same team's pre-AI baseline.

05 / What moved the number

Configuration first, because it is reversible.

06 / What kept it down

Then the half that lives in habits.

Once every dollar resolved to a person, the shape of the org appeared. Nine people accounted for 44% of spend; four of them sat in the bottom quartile for merged PRs, closed tickets and authored tests.

Trained the nine, not the two hundred

Four ninety-minute working sessions, run against each person's own transcripts with their own spend chart open. Habits changed: scoped context instead of pasted repositories, one well-formed request instead of thirty conversational turns, task class chosen before model.

in range at 90d

Named the champions

The same analysis surfaced seven engineers producing above-median output at below-median cost. Given protected time, one per squad, ownership of the CLAUDE.md hierarchy and internal skills after v1, and a fortnightly public transcript clinic in Confluence.

7 named

Returned deterministic work to scripts

Release notes to shell, dependency and licence reporting to Python against the Artifactory and Snyk APIs, log triage prefiltered by rule before anything reached a model, test fixtures seeded and committed, and a nightly cost-anomaly watch that opens its own Jira ticket.

$2.9k/mo
# the rule the teams now apply
Is the output exact and verifiable?   → script it
Does it run more than once a week?    → script it
Does it need judgement or language?   → model it
Both?                                 → script the filter, model the residue

Configuration produced roughly two-thirds of the reduction. This third is the part that holds after the engagement ends, because it lives in habits and repositories rather than in a settings page. Everything above runs inside the client's infrastructure, on their accounts, owned by them outright — no hosted middle layer, no key held, nothing to disconnect.

Send one month of your bill.

After my initial analysis, I give you my estimate on the potential amount I can save you — with 1 free actionable advice you can implement immediately and see the result. Read-only access, scoped to the analysis. Anything we pull is deleted within 30 days.

Book a demo